Working manuscript

Strategic Representations for Multiagent Coordination

Sandy Tanwisuth

10 October 2026

Abstract

What must an agent learn about a partner to choose a useful response? Predicting every detail of the partner’s behavior can be unnecessary, yet discarding a distinction can change which response is optimal. We study representations that preserve complete response policies in finite-horizon partner-induced Markov decision processes. We characterize exact equivalence of entropy-regularized responses and bound the regularized return loss from substituting an approximate response. Successor features give a separate preservation result for a linear reward class. These results motivate a learning method with response, consequence, and transition predictors, trained from response samples and interaction histories. A Bayesian model specifies the information gained about a fixed partner type through further interaction. The learning method remains to be evaluated: its central question is whether a representation can preserve useful responses at a lower information and computational cost than a detailed model of the partner.

Contents

Working manuscript

1 Which differences between partners matter?

A collaborator’s route can matter because it blocks a narrow corridor. Its exact motor sequence may not. Conversely, two collaborators that elicit the same immediate move can create different future opportunities. Strategic representation learning asks which distinctions should be retained, for which response objective, and at which time scale.

Strategic equivalence formalizes the minimal distinctions needed to preserve full best-response sets [1]. Computing those distinctions may itself require extensive information about a partner. A learned representation is useful only if it preserves what the decision requires at an acceptable cost, including the information used to train it. The mathematical target, learning objective, and evaluation metric must therefore be specified separately.

Clear incentives can sometimes help an observer choose a response, but neither competition nor cooperation guarantees predictability. Shared goals alone do not identify a useful convention. A helpful partner can leave several incompatible responses plausible; an obstructive partner can be easy to model. The relevant object here is the partner’s effect on the ego agent’s decision problem, with desirability of the resulting interaction assessed separately.

We consider four aspects of an interaction. A partner’s response contribution measures how much it changes the response relative to a reference population. Contextual response distributions describe the action to take at a particular decision point. Successor features describe longer-term consequences for a specified reward class, while partner-induced transition models describe the dynamics available to the ego agent. Response contribution, contextual compatibility, and successor features provide the three targets for the proposed layered learner. Their relationships depend on the reference population, reward class, and continuation policies; nesting these targets requires additional assumptions.

We first characterize the partition of partners that preserves complete soft-optimal policies, then bound the loss from approximate response substitution. We also give a preservation result for feature-based rewards. These results specify targets for the learning method. A Bayesian model describes how further interaction can resolve uncertainty about a partner. Together they lead to a practical question: when does learning less about behavior preserve enough information to act well?

Opponent modeling spans behavioral prediction, latent-state inference, and decision-focused representations [6]. Bayesian accounts of theory of mind infer beliefs or goals [7], [14]; learned models can infer aspects of another agent from interaction [15], [16]. Our target is narrower: predict the response-relevant effect of a specified partner policy.

Learning about human partners and selecting conventions are related but distinct problems [9], [13], [12]. Opponent-learning awareness models how an opponent may update [10], whereas the formal guarantees here fix the partner during the response calculation. Fictitious play provides historical context for best-response learning [8], but a contrastive encoder alone does not inherit a convergence guarantee. The multi-agent safety motivation includes coordination failures and strategic exploitation [11].

2 Which policy must a representation preserve?

2.1 A finite-horizon model

Let t∈{0,…,H−1}t\in\{0,\ldots,H-1\}, let SS and AA be finite, and assume every action in the nonempty set AA is available at every state. A partner policy is indexed by z∈𝒵z\in\mathcal Z. Fixing zz induces an MDP for the ego agent with transition law Pz,t(s′∣s,a)P_{z,t}(s'\mid s,a) and expected reward rz,t(s,a)∈[0,Rmax]r_{z,t}(s,a)\in[0,R_{\max}]. Terminal value is zero. All partners are compared on the same state/action spaces and horizon. A known finite memory state of a partner can be included in SS; an arbitrary partially observed or history-dependent partner is not covered without constructing an adequate state.

An ego policy π=(πt)t=0H−1\pi=(\pi_t)_{t=0}^{H-1} maps each time/state context to a distribution over actions. We compare policies at every context, including states that have zero probability under a particular initial distribution. This convention gives a policy equivalence independent of which starting states happen to be sampled. A weaker on-distribution relation would require its own support and coverage assumptions.

For τ>0\tau>0, define the entropy-regularized objective from context (t,s)(t,s) by Jz,tτ(π∣s)=𝔼z,π[∑u=tH−1(rz,u(Su,Au)+τℋ(πu(⋅∣Su)))|St=s].J^\tau_{z,t}(\pi\mid s)=\mathbb{E}_{z,\pi}\!\left[ \sum_{u=t}^{H-1}\bigl(r_{z,u}(S_u,A_u)+ \tau\mathcal H(\pi_u(\cdot\mid S_u))\bigr)\,\middle|\,S_t=s\right]. The conditional notation means starting the controlled process at (t,s)(t,s); it does not condition on a potentially zero-probability event under a different initial distribution. The reward term is task value; the entropy term is an explicit modeling choice. It is not uncertainty about whether the Q-function is correct. Throughout, ℋ\mathcal H uses natural logarithms, and all compared policies use the same finite positive temperature.

2.2 A full soft-optimal response

The backward recursion is Equations (1), (2), (3)Vz,Hτ(s)=0,Qz,tτ(s,a)=rz,t(s,a)+∑s′Pz,t(s′∣s,a)Vz,t+1τ(s′),Vz,tτ(s)=τlog⁡∑a∈Aexp⁡(Qz,tτ(s,a)/τ),πz,tτ,*(a∣s)=exp⁡((Qz,tτ(s,a)−Vz,tτ(s))/τ).\begin{align} V^\tau_{z,H}(s)&=0,\nonumber\\ Q^\tau_{z,t}(s,a)&=r_{z,t}(s,a)+ \sum_{s'}P_{z,t}(s'\mid s,a)V^\tau_{z,t+1}(s'),\label{eq:softQ}\\ V^\tau_{z,t}(s)&=\tau\log\sum_{a\in A} \exp(Q^\tau_{z,t}(s,a)/\tau),\label{eq:softV}\\ \pi^{\tau,*}_{z,t}(a\mid s)&= \exp\bigl((Q^\tau_{z,t}(s,a)-V^\tau_{z,t}(s))/\tau\bigr). \label{eq:softpolicy} \end{align} This gives a complete time/state policy, rather than a current action under an unspecified continuation. Finite rewards and horizon ensure finite values and a positive probability for every action.

Lemma 1 (Soft Bellman variational identity). For q∈ℝ|A|q\in\mathbb{R}^{|A|} and p∈Δ(A)p\in\Delta(A), let p*=softmax⁡(q/τ)p^*=\operatorname{softmax}(q/\tau). Then τlog⁡∑aeqa/τ=∑apaqa+τℋ(p)+τDKL(p‖p*).\tau\log\sum_a e^{q_a/\tau} =\sum_a p_aq_a+\tau\mathcal H(p)+\tau D_{\mathrm{KL}}(p\Vert p^*). Consequently p*p^* is the unique maximizer of the first two terms.

Proof. Substitute log⁡pa*=qa/τ−log⁡∑beqb/τ\log p_a^*=q_a/\tau-\log\sum_b e^{q_b/\tau} into the KL definition, with 0log⁡0=00\log0=0. Rearranging gives the identity. KL is nonnegative and is zero precisely at p=p*p=p^*. ◻

Backward induction using Lemma 1 proves that equation (3) is the unique entropy-regularized optimal policy at every context. The construction permits general-sum environments after the partner is fixed. It is a best-response construction, not a theorem that simultaneous learning converges to an equilibrium.

2.3 Exact equivalence of full policies

Definition 1 (Soft-policy equivalence). Write z∼τz′z\sim_\tau z' if πz,tτ,*(⋅∣s)=πz′,tτ,*(⋅∣s)\pi^{\tau,*}_{z,t}(\cdot\mid s)=\pi^{\tau,*}_{z',t}(\cdot\mid s) for every t,st,s.

Theorem 1 (Complete characterization). The following statements are equivalent:

  1. z∼τz′z\sim_\tau z';

  2. at each context there is a real ct,sc_{t,s} such that Qz,tτ(s,a)−Qz′,tτ(s,a)=ct,sQ^\tau_{z,t}(s,a)-Q^\tau_{z',t}(s,a)=c_{t,s} for all aa;

  3. max⁡t,sDKL(πz,tτ,*(⋅∣s)‖πz′,tτ,*(⋅∣s))=0\max_{t,s}D_{\mathrm{KL}}(\pi^{\tau,*}_{z,t}(\cdot\mid s)\Vert \pi^{\tau,*}_{z',t}(\cdot\mid s))=0.

Proof. At a fixed context, equality of softmax distributions implies equality of the ratios for actions a,ba,b, hence equality of (Qa−Qb)/τ(Q_a-Q_b)/\tau. Subtracting a fixed reference action gives the action-independent difference. Conversely, an additive constant cancels in softmax. Zero KL is equivalent to equality of the finite-support distributions. Repeating the argument at every context proves the result. ◻

The preimages of z↦πzτ,*z\mapsto\pi^{\tau,*}_z are therefore the coarsest exact partition preserving this full response. An abstraction hh preserves the response exactly when h(z)=h(z′)h(z)=h(z') implies z∼τz′z\sim_\tau z'. It may retain more distinctions. This condition constrains which partners can share an embedding; it does not force equivalent partners to have equal or nearby embeddings.

2.4 Hard best responses are a different relation

For the unregularized objective, define Vz,H0=0V^0_{z,H}=0, Qz,t0=rz,t+Pz,tVz,t+10Q^0_{z,t}=r_{z,t}+P_{z,t}V^0_{z,t+1}, and Vz,t0(s)=max⁡aQz,t0(s,a)V^0_{z,t}(s)=\max_a Q^0_{z,t}(s,a). Let Bz,t(s)=arg⁡max⁡aQz,t0(s,a)B_{z,t}(s)=\arg\max_aQ^0_{z,t}(s,a).

Proposition 1 (Full hard-policy best-response sets). A Markov policy is optimal from every time/state context iff it puts probability only on Bz,t(s)B_{z,t}(s) at every context. Consequently two partners have the same sets of such optimal policies iff all their Bz,t(s)B_{z,t}(s) agree.

Proof. Backward induction shows sufficiency. For necessity, a policy optimal at every context must have optimal continuation values. Assigning positive mass to a strictly suboptimal action then strictly lowers its expected value at that context. Equality of the local maximizing sets gives equality of the policy sets. If a maximizing action belongs to only one set at some context, a policy selecting it there and other maximizing actions elsewhere witnesses a difference. ◻

In particular, action-independent shifts of Q0Q^0 at every context are sufficient for full hard-policy equivalence. They are not necessary. Nor does Theorem 1 by itself identify hard and soft equivalence: its action values include entropy in future decisions. Matching greedy sets of those soft values is not a proof about unregularized optimal policies. This separates the preserved target from the full-policy SER motivation [1].

3 What does approximate response preservation guarantee?

Define the directed discrepancy dτ(z,z′)=maxt,sDKL(πz,tτ,*(⋅∣s)‖πz′,tτ,*(⋅∣s)).d_\tau(z,z')=\max_{t,s}D_{\mathrm{KL}}\bigl( \pi^{\tau,*}_{z,t}(\cdot\mid s)\Vert \pi^{\tau,*}_{z',t}(\cdot\mid s)\bigr). Its zero set is the exact equivalence relation above. For positive tolerance, the condition dτ(z,z′)≤εd_\tau(z,z')\leq\varepsilon is a similarity condition, not an equivalence relation. KL is asymmetric, and thresholding even a symmetric distance generally fails transitivity. For binary probabilities 0.4,0.5,0.60.4,0.5,0.6, adjacent directed divergences are below 0.030.03, whereas the first-to-last divergence exceeds it. Any clustering procedure must specify how it enforces its desired within-cluster guarantee.

Theorem 2 (Regularized response-substitution loss). Fix a partner zz, an initial state distribution μ\mu, and any Markov policy π\pi. Write Jzτ(π;μ)=𝔼S0∼μJz,0τ(π∣S0)J^\tau_z(\pi;\mu)=\mathbb{E}_{S_0\sim\mu}J^\tau_{z,0}(\pi\mid S_0). Then 𝔼μVz,0τ(S0)−Jzτ(π;μ)=τ𝔼z,π,μ∑t=0H−1DKL(πt(⋅∣St)‖πz,tτ,*(⋅∣St)).\mathbb{E}_\mu V^\tau_{z,0}(S_0)-J^\tau_z(\pi;\mu) =\tau\mathbb{E}_{z,\pi,\mu}\sum_{t=0}^{H-1} D_{\mathrm{KL}}\bigl(\pi_t(\cdot\mid S_t)\Vert \pi^{\tau,*}_{z,t}(\cdot\mid S_t)\bigr). If every contextual divergence on the right is at most ε\varepsilon, the loss is at most τHε\tau H\varepsilon.

Proof. Apply Lemma 1 at (t,St)(t,S_t) with q=Qz,tτ(St,⋅)q=Q^\tau_{z,t}(S_t,\cdot) and p=πt(⋅∣St)p=\pi_t(\cdot\mid S_t). Substitute equation (1), take the expectation under the actual partner zz and policy π\pi, and sum over tt. The next-state value terms cancel with successive current-state values, and Vz,Hτ=0V^\tau_{z,H}=0. The remaining reward and entropy terms form Jzτ(π;μ)J^\tau_z(\pi;\mu). Bounding each of the HH divergences gives the inequality. ◻

To substitute the optimal response to partner z′z' when the actual partner is zz, the sufficient condition is dτ(z′,z)≤εd_\tau(z',z)\leq\varepsilon. The direction matters. A bound measured only under an unrelated training-state distribution does not supply this uniform guarantee. The result concerns regularized return, not raw reward alone, human welfare, or equilibrium selection.

4 Which response-relevant objects should be represented?

4.1 Per-partner response contribution

Fix a context c=(t,s)c=(t,s) and a reference distribution ν\nu over partner types. Define the per-partner influence contribution Equation (4)ICc(z)=DKL(πz,cτ,*‖p‾c),p‾c=𝔼Z∼νπZ,cτ,*.\begin{equation} \label{eq:ic} \mathrm{IC}_c(z)=D_{\mathrm{KL}}\bigl(\pi^{\tau,*}_{z,c}\Vert\bar p_c\bigr), \qquad \bar p_c=\mathbb{E}_{Z\sim\nu}\pi^{\tau,*}_{Z,c}. \end{equation} For a specified context distribution ρ\rho, use ICρ(z)=𝔼X∼ρICX(z)\mathrm{IC}_\rho(z)=\mathbb{E}_{X\sim\rho}\mathrm{IC}_X(z). The reference population and context weights are part of the definition, so comparisons across datasets must report them. To define the population quantity, draw X∼ρX\sim\rho and Z∼νZ\sim\nu independently, then draw A∣X=c,Z=z∼πz,cτ,*A\mid X=c,Z=z\sim\pi^{\tau,*}_{z,c}. Under this law, I(A;Z∣X)=𝔼X∼ρ,Z∼νICX(Z).I(A;Z\mid X)=\mathbb{E}_{X\sim\rho,Z\sim\nu}\mathrm{IC}_X(Z). With a context-dependent partner prior, replace ν\nu by ν(⋅∣c)\nu(\cdot\mid c) consistently.

A proposed filter retains {z:ICρ(z)≥θ}\{z:\mathrm{IC}_\rho(z)\geq\theta\}. This selects partners that change the response relative to the reference mixture. Response entropy ℋ(πz,cτ,*)\mathcal H(\pi^{\tau,*}_{z,c}) measures a different property. If every partner induces the same deterministic limiting response, conditional entropy and mutual information are both zero. A rare partner can have a high KL contribution even when the population average is small. Neither quantity identifies cooperation or benevolence.

We use Intent Certainty (IC) to denote this response influence relative to a reference population. A consistently adversarial partner may be easy to model and highly influential. An IC filter allocates modeling effort according to response influence, independently of whether the partner’s goals are desirable. Evaluation must account for the cost of excluding partners whose responses matter outside the sampled contexts.

4.2 Contextual compatibility

At a context (t,s)(t,s), the directed divergence between two soft-optimal response distributions measures agreement about the current decision. Those distributions already incorporate future interaction through their Bellman values. Thus “short-term” identifies the location of the response, not a myopic objective. Checking every context yields the full-policy discrepancy; checking sampled contexts gives a distribution-dependent empirical measurement with a separate coverage obligation.

4.3 Successor features and occupancy

Fix an ego continuation policy π\pi and a bounded feature map φ\varphi on the relevant states or transitions. A finite-horizon successor feature vector is ψtπ,z(s,a)=𝔼π,z[∑u=tH−1φ(u,Su,Au,Su+1)|St=s,At=a].\psi^{\pi,z}_t(s,a)=\mathbb{E}_{\pi,z}\!\left[ \sum_{u=t}^{H-1}\varphi(u,S_u,A_u,S_{u+1}) \,\middle|\, S_t=s,A_t=a\right]. Time is included so that a feature can distinguish when an event occurs. For a finite feature-label set FF, an alternative is the occupancy measure ξtπ,z(s,a,f)=∑u=tH−1ℙπ,z(Φu=f∣St=s,At=a).\xi^{\pi,z}_t(s,a,f)=\sum_{u=t}^{H-1} \mathbb{P}_{\pi,z}(\Phi_u=f\mid S_t=s,A_t=a). It has total mass H−tH-t; ξ/(H−t)\xi/(H-t) is a probability distribution. The discounted infinite-horizon analogue, ξ(s,a,f)=∑k≥0γkℙ(Φt+k=f∣s,a)\xi(s,a,f)=\sum_{k\geq0}\gamma^k\mathbb{P}(\Phi_{t+k}=f\mid s,a), requires 0≤γ<10\leq\gamma<1 and has mass 1/(1−γ)1/(1-\gamma). Ordinary KL applies to normalized distributions, not arbitrary signed feature vectors. A norm can instead be used directly on ψ\psi.

Successor-feature representations offer a way to compare consequences independently of one chosen reward parameter [3]. They preserve expected feature visitation, not the complete joint distribution of trajectories. Two processes can have the same marginal occupancy and different temporal correlations. Statements about “all goals” must specify the reward class represented by the features.

4.4 A reward-class guarantee

For this subsection, use a separate common linear reward model rw(u,s,a,s′)=w⊤φ(u,s,a,s′)r_w(u,s,a,s')=w^\top\varphi(u,s,a,s'), with bounded features and finite ww. These rewards are bounded but may have either sign; they do not inherit the earlier nonnegative reward convention. Let Jz,wJ_{z,w} denote the unregularized return and let best responses maximize that return. For a given initial context, let ψπ,z\psi^{\pi,z} denote the expected sum without conditioning on a forced first action. Then Jz,w(π)=w⊤ψπ,zJ_{z,w}(\pi)=w^\top\psi^{\pi,z} by linearity of expectation.

Proposition 2 (Full-policy preservation for a feature reward class). If ψπ,z=ψπ,z′\psi^{\pi,z}=\psi^{\pi,z'} for every ego policy and every initial context, then all policies have the same return against zz and z′z' for every ww, and their full hard best-response sets agree. More generally, if sup⁡π∥ψπ,z−ψπ,z′∥≤δ\sup_\pi\|\psi^{\pi,z}-\psi^{\pi,z'}\|\leq\delta and ∥w∥*≤M\|w\|_*\leq M, a hard best response to z′z' loses at most 2Mδ2M\delta against zz, for the specified initial context.

Proof. Exact equality gives equality of each policy’s return. For the approximate case, dual-norm inequality bounds each policy’s return difference by MδM\delta. Insert the z′z' returns between the return of a best response to zz and the return of a best response to z′z'. Optimality under z′z' makes the middle difference nonpositive, leaving two errors of at most MδM\delta. ◻

The quantifier over all ego policies is substantive. Equality under a single fixed continuation establishes that continuation’s feature agreement, not equality of full best-response sets. Estimating the uniform condition is a separate computational problem.

4.5 Learning and examples

A successor-feature encoder may be designed to place similar outcomes nearby. This is a design constraint, not a consequence of forward continuity alone. If similar outcomes imply similar embeddings, the reverse implication need not hold: a constant embedding satisfies the first implication while collapsing every distinction. A decoder or inverse error bound is needed before nearby embeddings certify similar outcomes.

In Overcooked, clockwise and counterclockwise delivery routes might yield similar counts of onions delivered or dishes served under a particular partner policy. Whether they are interchangeable depends on timing, blocked paths, the ego continuation, and the selected features. In Cleanup, cleaning before or after collecting an apple can yield similar aggregate outcomes while changing early opportunities. These examples explain why the feature map and time horizon must be reported rather than assumed to preserve every strategically relevant consequence.

4.6 Partner-induced world models

A partner also changes the effective transition kernel seen by the ego agent. In a Markov game with joint transition law Pt(s′∣s,a,a−i)P_t(s'\mid s,a,a_{-i}), a fixed Markov partner induces Pz,t(s′∣s,a)=∑a−iPt(s′∣s,a,a−i)zt(a−i∣s).P_{z,t}(s'\mid s,a)=\sum_{a_{-i}}P_t(s'\mid s,a,a_{-i})z_t(a_{-i}\mid s). Model-aware equality means these kernels agree for every t,s,a,s′t,s,a,s'. If rewards also agree, backward induction gives equal hard and soft response objects. Transition equality alone does not establish equal incentives when the partner changes rewards. Conversely, similar long-term feature occupancy under one ego policy need not imply the same kernel.

The model-aware target is useful for planning and for identifying changed affordances. A partner blocking a corridor alters which next states are likely after movement. A partner delaying river cleaning may alter resource availability. Similar eventual task completion does not prove identical one-step dynamics. Approximate kernel comparison needs a context-weighting measure, support conventions for KL, and a bound relating model error to the relevant value or response error.

The four targets can therefore disagree without contradiction. A partner can change its motion model while preserving a best response, or preserve a short-term action preference while changing future feature occupancy. Such disagreement should motivate model revision or a task-specific decision, not an automatic claim that the partner is unsafe.

5 Learning representations from response samples

What can be learned without storing a complete response table for every partner? Fix a context c=(t,s)c=(t,s) and a reference distribution ν\nu over partner types. Draw Z∼νZ\sim\nu and A∣Z=z,c∼πz,cτ,*A\mid Z=z,c\sim\pi^{\tau,*}_{z,c}. This defines a response-sampling law. It can be generated by an exact dynamic program in the finite model or approximated by estimated values; those two sources should be evaluated separately.

Let hθ(z,c)h_\theta(z,c) encode a partner at the context and let a positive score be fθ(a,h)=exp⁡(gθ(a)⊤h/η)f_\theta(a,h)=\exp(g_\theta(a)^\top h/\eta), with contrastive temperature η>0\eta>0. The contrastive temperature is distinct from the decision temperature τ\tau. For an integer N≥2N\geq2, with one conditional positive A+A^+ and N−1N-1 independent negatives from the marginal response law p‾c(a)=𝔼νπZ,cτ,*(a)\bar p_c(a)=\mathbb{E}_\nu\pi^{\tau,*}_{Z,c}(a), use Equation (5)ℒN(c)=−𝔼log⁡fθ(A+,hθ(Z,c))fθ(A+,hθ(Z,c))+∑j=1N−1fθ(Aj−,hθ(Z,c)).\begin{equation} \label{eq:nce} \mathcal L_N(c)=-\mathbb{E}\log \frac{f_\theta(A^+,h_\theta(Z,c))} {f_\theta(A^+,h_\theta(Z,c))+ \sum_{j=1}^{N-1}f_\theta(A_j^-,h_\theta(Z,c))}. \end{equation} The positive belongs in the denominator. Positivity of the score is necessary for the log ratio; an untransformed dot product or cosine similarity does not ensure it. Sampling each negative partner independently from ν\nu and then sampling its response produces the required marginal. Conditioning negatives to exclude the anchor, selecting only dissimilar partners, or using a correlated replay batch changes the sampling law and requires another bound.

Let XX denote the sampled decision context, or the constant context cc for this conditional calculation. Under the stated sampling scheme, the InfoNCE bound gives log⁡N−ℒN(c)≤I(A;hθ(Z,c)∣X=c)≤I(A;Z∣X=c)\log N-\mathcal L_N(c)\leq I(A;h_\theta(Z,c)\mid X=c)\leq I(A;Z\mid X=c) [2]. For a fixed response-generating law, optimizing the encoder changes a lower bound and the information retained by its representation, not the underlying I(A;Z∣X=c)I(A;Z\mid X=c) itself.

5.1 Prediction and geometry

A low contrastive loss can support response prediction without making every pair of equivalent partners nearby in Euclidean space. For example, append an arbitrary coordinate to an embedding and let the scorer ignore it. The loss is unchanged while distances can become arbitrarily large. Even an optimal loss therefore does not prove an embedding-proximity claim.

A representation can instead be evaluated through a response decoder π̂θ(⋅∣c,h)\widehat\pi_\theta(\cdot\mid c,h). Held-out KL estimates diagnose prediction error; a certified uniform contextual bound, in the required direction, supplies the uniform response-substitution guarantee. A geometric clustering claim needs additional constraints—for example, a specified metric-learning objective, decoder regularity, or a canonical representation—and its own proof. The current construction treats geometric consistency as an empirical objective, not a theorem derived from InfoNCE alone.

In Overcooked, partners using the same circulation convention may induce similar useful responses despite different motor behavior. Which direction the ego should travel depends on the layout and joint routes; it cannot be inferred from a partner’s clockwise motion alone. In Cleanup, standing near a river may precede cleaning or apple collection. A representation should be updated from observed, context-dependent effects rather than treating a single action as proof of intention. These are motivating examples, not experiment results.

6 Which uncertainty is worth resolving?

6.1 A Bayes-adaptive interpretation

Bayesian decision problems augment the physical state with a posterior over unknown parameters [19], [20]. The framework draws on the distinction between value obtainable under current knowledge and value gained by obtaining information [21]. This motivates asking whether a partner changes the response, what response is useful now, and what long-term effects remain uncertain.

The model below specifies a latent partner type, a prior, an observation likelihood, and a posterior update. The learner’s predictors summarize aspects of this belief; their agreement alone is insufficient to recover the posterior. We retain the full belief in the formal analysis and leave the sufficiency of compressed predictors to future work. Adding an information-gain reward can change the optimal policy, so its effect on the original task objective must be assessed separately.

6.2 Information gain about a fixed partner type

Can an agent choose interactions that resolve the response distinctions it still does not understand? Let CC be a fixed latent partner type during an evaluation episode. For example, CC may be a class under a fixed full-response equivalence relation, provided a generative model over the members of each class is specified. Let b0(c)b_0(c) be a prior on a finite type set and let the observation alphabet be finite. The action-selection rule depends only on observed history and independent randomization. The observation model is Ot+1∼Lt(⋅∣C,Dt,ut),O_{t+1}\sim L_t(\cdot\mid C,D_t,u_t), where DtD_t is the observed interaction history and utu_t is the ego’s selected action or experiment. The likelihood must include any state transition, observed partner action, partial observation, and nuisance variables needed by the environment. It is not determined merely by naming an equivalence class.

Write bt(c)=ℙ(C=c∣Dt)b_t(c)=\mathbb{P}(C=c\mid D_t) for the current type posterior. The predictive law and posterior update are mt(o∣Dt,u)=∑cbt(c)Lt(o∣c,Dt,u),bt+1(c)=bt(c)Lt(Ot+1∣c,Dt,ut)mt(Ot+1∣Dt,ut).\begin{align*} m_t(o\mid D_t,u)&=\sum_c b_t(c)L_t(o\mid c,D_t,u),\\ b_{t+1}(c)&=\frac{b_t(c)L_t(O_{t+1}\mid c,D_t,u_t)} {m_t(O_{t+1}\mid D_t,u_t)}. \end{align*} Updates are defined for observations of positive predictive probability. The proposed expected learning reward is Equation (6)IGt(u∣Dt)=𝔼O∼mt(⋅∣Dt,u)DKL(bt(⋅∣O,u)‖bt).\begin{equation} \label{eq:ig} \mathrm{IG}_t(u\mid D_t)= \mathbb{E}_{O\sim m_t(\cdot\mid D_t,u)} D_{\mathrm{KL}}\bigl(b_t(\cdot\mid O,u)\Vert b_t\bigr). \end{equation}

Proposition 3 (Expected uncertainty reduction). Under the stated finite Bayesian model, IGt(u∣Dt)=ℋ(bt)−𝔼Oℋ(bt(⋅∣O,u))=I(C;Ot+1∣Dt,u)≥0.\mathrm{IG}_t(u\mid D_t)= \mathcal H(b_t)-\mathbb{E}_O\mathcal H(b_t(\cdot\mid O,u)) = I(C;O_{t+1}\mid D_t,u)\geq0.

Proof. Expand the expected KL as a sum over c,oc,o using bt(c)Lt(o∣c,Dt,u)=mt(o∣Dt,u)bt(c∣o,u)b_t(c)L_t(o\mid c,D_t,u)=m_t(o\mid D_t,u)b_t(c\mid o,u). The posterior log term gives negative expected posterior entropy, and summing the prior log term over observations gives prior entropy. Nonnegativity follows from KL nonnegativity. ◻

A realized entropy difference can be negative even though its expectation is nonnegative. The realized posterior KL is a different nonnegative random reward with the same conditional expectation as equation (6). An implementation should choose and report which estimator it uses.

Information gain is about the agent’s knowledge of a latent type. It is distinct from changing the partner’s behavior and from receiving task reward. One may optimize expected information gain subject to a task or safety constraint, or combine it with explicitly weighted environment reward; neither choice makes the objectives identical. A causal-influence claim requires interventions or identification assumptions in addition to an observation model.

For a fixed finite latent type and a correct model, cumulative expected information gain is at most ℋ(C)\mathcal H(C). Repeated observation therefore cannot supply an unbounded learning signal about that same type. Continual open-ended learning needs new tasks, types, or model structure and a defined mechanism for introducing them. Learned changes in a partition are tracked separately from posterior learning within a fixed partition; comparing entropy across different class spaces is not automatically information gain.

6.3 Using the layers during inference

IC selects response-changing partners relative to the reference mixture; contextual response comparison measures local policy compatibility; successor features measure a declared class of longer-term consequences. These targets can disagree without a logical contradiction. A changing transition kernel can preserve a response, and a shared current action can precede different later outcomes.

A proposed learner maintains these predictors alongside a type posterior. It updates them from trajectories, compares their errors on held-out contexts, and decides whether an additional observation has expected decision or information value. It may defer a decision, collect a targeted observation, or take a task-optimal action. A rule for deferral needs an explicit cost and deadline. No theorem here states that arbitrary layer agreement is necessary or sufficient for action.

7 Computation and learning procedure

7.1 Training a response representation

The proposed training procedure separates response-target construction from representation learning.

  1. Fix the partner training law ν\nu, contexts, horizon, reward, and temperatures τ\tau and η\eta. Keep a held-out partner set for evaluation.

  2. For a sampled partner zz, compute the soft-optimal response by backward induction, or obtain a separately evaluated estimate of it. State whether the experiment uses an oracle model or learned values.

  3. Sample a context and a positive action from that partner’s response. Generate N−1N-1 independent marginal negatives using independently drawn reference partners.

  4. Evaluate equation (5), including the positive in the denominator, and update the encoder and score model. If a response decoder is trained, report its own objective.

  5. Evaluate contextual response error and task/regularized return on held-out partners. Measure the cost of generating response targets as well as training the encoder.

When only trajectory prefixes are observed, the operational input is an observation encoder or a posterior over partner hypotheses, not an inaccessible complete policy. A policy-input oracle and a trajectory-input implementation should not be reported as the same experiment.

7.2 Response neighborhoods and uncertainty

For a finite reference pool 𝒫\mathcal P, an estimated partner ẑt\widehat z_t, and a chosen direction, one can define 𝒩ε(ẑt)={z′∈𝒫:dτ(z′,ẑt)≤ε}.\mathcal N_\varepsilon(\widehat z_t) =\{z'\in\mathcal P:d_\tau(z',\widehat z_t)\leq\varepsilon\}. If a uniform diagnostic is desired, its entropy is log⁡|𝒩ε|\log|\mathcal N_\varepsilon| for a nonempty set. This reports neighborhood size, not a calibrated posterior over the actual partner. A conditioned prior generally gives different weights. An empty set needs an explicit out-of-model outcome, and successive neighborhoods need not be nested.

This is an oracle response neighborhood: membership uses response distributions rather than embedding distance. It serves as an evaluation diagnostic. Replacing it with embedding distance requires a validated relationship between embedding geometry and response error; it cannot be assumed from the InfoNCE objective. Reversing the KL direction changes the diagnostic and must be reported.

7.3 Joint prediction and refinement

The response encoder may be trained alongside successor-feature and transition predictors. Their targets should be kept explicit: response examples train the normalized contrastive objective, collected trajectories train consequence and dynamics predictors, and an observation likelihood updates the type posterior. If the likelihood is learned, its calibration must be evaluated independently. IC is estimated against the same declared reference partner population throughout a comparison.

The encoder, environment policy, and Bayesian model need not update on the same time scale. Slower model updates or a fixed evaluation snapshot can keep the target interpretable while behavior changes. Evaluation would track response error, regularized loss, task return, and performance on held-out partner families alongside training loss and posterior concentration. Testing the full-policy guarantee requires coverage of the time/state contexts in its bound.

A useful evaluation separates three sources of error: approximating partner-induced dynamics or values; learning a compressed response representation; and selecting actions with an uncertain partner model. Ablations should remove one component at a time while keeping the partner population, context distribution, and budget fixed.

8 Proposed evaluation

8.1 Environments and comparisons

We propose evaluating the method in Overcooked: Counter Circuit, where coordination depends on shared paths and workstations, and Coin Game, where partners can have mixed incentives. Existing code provides two starting points. The value-of-intent repository contains Overcooked PPO and temporal-CPC implementations with saved comparison outputs [17]. The strategic-representation-learning notebook illustrates response and KL calculations using hand-defined Q-vectors [18]. Neither implements the complete response-conditioned learner and Bayesian information-gain model studied here. The evaluation described below has not yet been carried out.

The comparison would use a common partner pool and interaction budget for the response-based encoder, behavioral contrastive embeddings, and state–action co-occurrence representations. An oracle response representation would isolate the cost of estimating the training targets. Varying partner stochasticity, convention, and observation length would test which distinctions the encoders retain. Evaluation would include both familiar-distribution and held-out partner families.

8.2 Measurements

The main measurements would be contextual response KL, regularized response-substitution loss, task return, posterior calibration, and adaptation speed. Comparing predicted information gain with realized posterior updates and type-prediction accuracy would assess the observation model. Embedding plots and neighborhood sizes would provide additional diagnostics; response and return measurements are needed to assess decision quality.

Reproducibility requires recording the seeds, partner and context sampling laws, train/evaluation split, horizon, temperatures, negative-sampling distribution, and computation used to generate targets. Confidence intervals would use independent runs as the sampling unit. A trajectory-input implementation would also need an explicit observation model and belief or recurrent-state construction. The full-state guarantees apply only when that state satisfies the assumptions of the decision model.

8.3 Hypotheses that remain open

Response-based learning may resolve some partner differences earlier than behavioral reconstruction, particularly where many behaviors support the same useful response. Noisy partners may need longer interaction to identify their response-relevant type. We also hypothesize that lower response error improves held-out coordination under the relevant reward model.

A plateau in a learning or entropy curve can indicate successful identification, an uninformative experiment, model misspecification, or optimization failure. It does not by itself establish enough information to act. Zero-shot transfer and resilience to exploitation require direct held-out tests; neither follows merely from a small training loss or stable embedding.

9 What would sustained learning require?

The proposed intrinsic motivation is interaction-centered: an agent seeks observations that resolve how others affect its choices. Predictability can assist adaptation, but a reward for learning about a partner must not be confused with endorsing its goals. Information gain concerns knowledge; a social-influence reward concerns an intervention that changes another agent’s behavior [22]. A causal interpretation needs an intervention or identification model and a counterfactual baseline, beyond a learned association.

9.1 Continual discovery and state marginal matching

Observer-relative open-endedness asks for continuing novelty and learnability [4]. The fixed finite type model above eventually exhausts its information. Continued discovery therefore requires a growing hypothesis space, a curriculum, or changing tasks, together with a mechanism that introduces them and distinguishes learning from model collapse or arbitrary relabeling. None of the finite-model results proves indefinitely improving intelligence.

State marginal matching with mixtures of policies offers a possible outer objective [5]. One could select interaction settings or policy mixtures that visit a specified distribution of strategic situations. This target must be defined separately from the information-gain reward, and its effect on response learning must be tested. No implemented outer loop or convergence result for agreement among abstraction layers is established here.

9.2 Internal components and self-models

Response, successor-feature, and model checks could also be applied to an agent’s own components. Disagreement might reveal an error or a mismatch between stated and enacted objectives. Interpreting this as self-awareness or self-alignment remains an application hypothesis requiring an operational evaluation; agreement among predictors does not by itself establish correctness.

9.3 Interpretation, autonomy, and broader impact

In this framework a “type” is a task-dependent model of response-relevant behavior, not a judgment about a person’s identity or worth. Such a model should be revised when observations conflict with its assumptions. A partner can be predictable without being cooperative, and agreement across representations can coexist with undesirable goals. The formal results cover the specified finite model; extensions to biological, artificial, or mixed teams require evidence about the actual observation and decision setting.

Decision-focused representations may support human–robot coordination, teaching systems, and decentralized planning while avoiding unnecessary behavioral reconstruction. They may also assist manipulation, preemption, or inference about sensitive roles. Representing another agent’s influence does not establish consent, permission to use inferred preferences, or the legitimacy of an intervention. Preserving autonomy therefore concerns how a model is used as well as what it predicts. Response predictability does not solve value conflict or establish universally safe coordination.

10 Conclusion

Strategic representations can be specified by the decisions they preserve. In the finite model studied here, complete soft-optimal responses admit an exact characterization, and uniform directed response error controls entropy-regularized loss. Successor features provide a separate guarantee for a declared linear reward class. These results identify what a learned representation would need to preserve; contrastive training alone does not establish that it does so.

The next question is whether these response, consequence, and belief models can meet useful error targets with less interaction or computation than modeling a partner in full. An answer requires held-out evaluation, calibrated inference, and an accounting of the information used to construct training targets. The proposed experiments would test this tradeoff.

Acknowledgments

I thank Niklas Lauffer for his advice; the CHAI community, including Cassidy Laidlaw, Michelle Li, Karim Abdelsadek, Eli Bronstein, and Cam Allen; and the MATS community, including Alex Cloud and Kola Ayonrinde.

Dedicated to Meridian, Cassiopeia, and Chris.

References

[1]

Niklas Lauffer, Ameesh Shah, Micah Carroll, Michael Dennis, and Stuart Russell. Who Needs to Know? Minimal Knowledge for Optimal Coordination. 2023. https://arxiv.org/abs/2306.09309.

[2]

Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation Learning with Contrastive Predictive Coding. 2018, revised 2019. https://arxiv.org/abs/1807.03748.

[3]

Chris Reinke and Xavier Alameda-Pineda. Successor Feature Representations. Transactions on Machine Learning Research, 2022. https://openreview.net/forum?id=MTFf1rDDEI.

[4]

Edward Hughes et al. Open-Endedness is Essential for Artificial Superhuman Intelligence. 2024. https://arxiv.org/abs/2406.04268.

[5]

Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Ruslan Salakhutdinov, and Sergey Levine. State Marginal Matching with Mixtures of Policies. 2019.

[6]

Stefano V. Albrecht and Peter Stone. Autonomous agents modelling other agents: A comprehensive survey and open problems. Artificial Intelligence, 258:66–95, 2018.

[7]

Chris Baker, Rebecca Saxe, and Joshua Tenenbaum. Bayesian theory of mind: Modeling joint belief-desire attribution. Proceedings of the Cognitive Science Society, 33, 2011.

[8]

George W. Brown. Iterative solution of games by fictitious play. 1951.

[9]

Micah Carroll et al. On the utility of learning about humans for human–AI coordination. 2020. https://arxiv.org/abs/1910.05789.

[10]

Jakob N. Foerster et al. Learning with Opponent-Learning Awareness. 2018. https://arxiv.org/abs/1709.04326.

[11]

Lewis Hammond et al. Multi-Agent Risks from Advanced AI. 2025. https://arxiv.org/abs/2502.14143.

[12]

Hengyuan Hu et al. Off-Belief Learning. 2021. https://arxiv.org/abs/2103.04000.

[13]

Hengyuan Hu, Adam Lerer, Alex Peysakhovich, and Jakob Foerster. Other-Play for Zero-Shot Coordination. 2021. https://arxiv.org/abs/2003.02979.

[14]

Julian Jara-Ettinger. Theory of mind as inverse reinforcement learning. Current Opinion in Behavioral Sciences, 29:105–110, 2019.

[15]

Neil C. Rabinowitz et al. Machine Theory of Mind. 2018. https://arxiv.org/abs/1802.07740.

[16]

Roberta Raileanu, Emily Denton, Arthur Szlam, and Rob Fergus. Modeling Others using Oneself in Multi-Agent Reinforcement Learning. 2018. https://arxiv.org/abs/1802.09640.

[17]

Sandy Tanwisuth. value-of-intent. Code and notebook outputs, commit 686557b2f99d419dd5c07f142817060234207203. https://github.com/sandguine/value-of-intent/tree/686557b2f99d419dd5c07f142817060234207203.

[18]

Sandy Tanwisuth. strategic-representation-learning demonstration, commit 2ab81418a97c2410d05a3e53dcce7c92858a2cd8. https://github.com/sandguine/strategic-representation-learning/tree/2ab81418a97c2410d05a3e53dcce7c92858a2cd8.

[19]

Richard Bellman and Robert Kalaba. A mathematical theory of adaptive control processes. Proceedings of the National Academy of Sciences, 45(8):1288–1290, 1959.

[20]

James John Martin. Bayesian Decision Problems and Markov Chains. 1967.

[21]

Aly Lidayan, Michael Dennis, and Stuart Russell. BAMDP Shaping: A Unified Framework for Intrinsic Motivation and Reward Shaping. 2025 revision. https://arxiv.org/abs/2409.05358.

[22]

Natasha Jaques et al. Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning. 2019. https://arxiv.org/abs/1810.08647.