Working manuscript
Strategic Representations for Multiagent Coordination
10 October 2026
Abstract
What must an agent learn about a partner to choose a useful response? Predicting every detail of the partner’s behavior can be unnecessary, yet discarding a distinction can change which response is optimal. We study representations that preserve complete response policies in finite-horizon partner-induced Markov decision processes. We characterize exact equivalence of entropy-regularized responses and bound the regularized return loss from substituting an approximate response. Successor features give a separate preservation result for a linear reward class. These results motivate a learning method with response, consequence, and transition predictors, trained from response samples and interaction histories. A Bayesian model specifies the information gained about a fixed partner type through further interaction. The learning method remains to be evaluated: its central question is whether a representation can preserve useful responses at a lower information and computational cost than a detailed model of the partner.
Contents
Working manuscript
1 Which differences between partners matter?
A collaborator’s route can matter because it blocks a narrow corridor. Its exact motor sequence may not. Conversely, two collaborators that elicit the same immediate move can create different future opportunities. Strategic representation learning asks which distinctions should be retained, for which response objective, and at which time scale.
Strategic equivalence formalizes the minimal distinctions needed to preserve full best-response sets [1]. Computing those distinctions may itself require extensive information about a partner. A learned representation is useful only if it preserves what the decision requires at an acceptable cost, including the information used to train it. The mathematical target, learning objective, and evaluation metric must therefore be specified separately.
Clear incentives can sometimes help an observer choose a response, but neither competition nor cooperation guarantees predictability. Shared goals alone do not identify a useful convention. A helpful partner can leave several incompatible responses plausible; an obstructive partner can be easy to model. The relevant object here is the partner’s effect on the ego agent’s decision problem, with desirability of the resulting interaction assessed separately.
We consider four aspects of an interaction. A partner’s response contribution measures how much it changes the response relative to a reference population. Contextual response distributions describe the action to take at a particular decision point. Successor features describe longer-term consequences for a specified reward class, while partner-induced transition models describe the dynamics available to the ego agent. Response contribution, contextual compatibility, and successor features provide the three targets for the proposed layered learner. Their relationships depend on the reference population, reward class, and continuation policies; nesting these targets requires additional assumptions.
We first characterize the partition of partners that preserves complete soft-optimal policies, then bound the loss from approximate response substitution. We also give a preservation result for feature-based rewards. These results specify targets for the learning method. A Bayesian model describes how further interaction can resolve uncertainty about a partner. Together they lead to a practical question: when does learning less about behavior preserve enough information to act well?
1.1 Related work and scope
Opponent modeling spans behavioral prediction, latent-state inference, and decision-focused representations [6]. Bayesian accounts of theory of mind infer beliefs or goals [7], [14]; learned models can infer aspects of another agent from interaction [15], [16]. Our target is narrower: predict the response-relevant effect of a specified partner policy.
Learning about human partners and selecting conventions are related but distinct problems [9], [13], [12]. Opponent-learning awareness models how an opponent may update [10], whereas the formal guarantees here fix the partner during the response calculation. Fictitious play provides historical context for best-response learning [8], but a contrastive encoder alone does not inherit a convergence guarantee. The multi-agent safety motivation includes coordination failures and strategic exploitation [11].
2 Which policy must a representation preserve?
2.1 A finite-horizon model
Let , let and be finite, and assume every action in the nonempty set is available at every state. A partner policy is indexed by . Fixing induces an MDP for the ego agent with transition law and expected reward . Terminal value is zero. All partners are compared on the same state/action spaces and horizon. A known finite memory state of a partner can be included in ; an arbitrary partially observed or history-dependent partner is not covered without constructing an adequate state.
An ego policy maps each time/state context to a distribution over actions. We compare policies at every context, including states that have zero probability under a particular initial distribution. This convention gives a policy equivalence independent of which starting states happen to be sampled. A weaker on-distribution relation would require its own support and coverage assumptions.
For , define the entropy-regularized objective from context by The conditional notation means starting the controlled process at ; it does not condition on a potentially zero-probability event under a different initial distribution. The reward term is task value; the entropy term is an explicit modeling choice. It is not uncertainty about whether the Q-function is correct. Throughout, uses natural logarithms, and all compared policies use the same finite positive temperature.
2.2 A full soft-optimal response
The backward recursion is Equations (1), (2), (3) This gives a complete time/state policy, rather than a current action under an unspecified continuation. Finite rewards and horizon ensure finite values and a positive probability for every action.
Lemma 1 (Soft Bellman variational identity). For and , let . Then Consequently is the unique maximizer of the first two terms.
Proof. Substitute into the KL definition, with . Rearranging gives the identity. KL is nonnegative and is zero precisely at . ◻
Backward induction using Lemma 1 proves that equation (3) is the unique entropy-regularized optimal policy at every context. The construction permits general-sum environments after the partner is fixed. It is a best-response construction, not a theorem that simultaneous learning converges to an equilibrium.
2.3 Exact equivalence of full policies
Definition 1 (Soft-policy equivalence). Write if for every .
Theorem 1 (Complete characterization). The following statements are equivalent:
;
at each context there is a real such that for all ;
.
Proof. At a fixed context, equality of softmax distributions implies equality of the ratios for actions , hence equality of . Subtracting a fixed reference action gives the action-independent difference. Conversely, an additive constant cancels in softmax. Zero KL is equivalent to equality of the finite-support distributions. Repeating the argument at every context proves the result. ◻
The preimages of are therefore the coarsest exact partition preserving this full response. An abstraction preserves the response exactly when implies . It may retain more distinctions. This condition constrains which partners can share an embedding; it does not force equivalent partners to have equal or nearby embeddings.
2.4 Hard best responses are a different relation
For the unregularized objective, define , , and . Let .
Proposition 1 (Full hard-policy best-response sets). A Markov policy is optimal from every time/state context iff it puts probability only on at every context. Consequently two partners have the same sets of such optimal policies iff all their agree.
Proof. Backward induction shows sufficiency. For necessity, a policy optimal at every context must have optimal continuation values. Assigning positive mass to a strictly suboptimal action then strictly lowers its expected value at that context. Equality of the local maximizing sets gives equality of the policy sets. If a maximizing action belongs to only one set at some context, a policy selecting it there and other maximizing actions elsewhere witnesses a difference. ◻
In particular, action-independent shifts of at every context are sufficient for full hard-policy equivalence. They are not necessary. Nor does Theorem 1 by itself identify hard and soft equivalence: its action values include entropy in future decisions. Matching greedy sets of those soft values is not a proof about unregularized optimal policies. This separates the preserved target from the full-policy SER motivation [1].
3 What does approximate response preservation guarantee?
Define the directed discrepancy Its zero set is the exact equivalence relation above. For positive tolerance, the condition is a similarity condition, not an equivalence relation. KL is asymmetric, and thresholding even a symmetric distance generally fails transitivity. For binary probabilities , adjacent directed divergences are below , whereas the first-to-last divergence exceeds it. Any clustering procedure must specify how it enforces its desired within-cluster guarantee.
Theorem 2 (Regularized response-substitution loss). Fix a partner , an initial state distribution , and any Markov policy . Write . Then If every contextual divergence on the right is at most , the loss is at most .
Proof. Apply Lemma 1 at with and . Substitute equation (1), take the expectation under the actual partner and policy , and sum over . The next-state value terms cancel with successive current-state values, and . The remaining reward and entropy terms form . Bounding each of the divergences gives the inequality. ◻
To substitute the optimal response to partner when the actual partner is , the sufficient condition is . The direction matters. A bound measured only under an unrelated training-state distribution does not supply this uniform guarantee. The result concerns regularized return, not raw reward alone, human welfare, or equilibrium selection.
4 Which response-relevant objects should be represented?
4.1 Per-partner response contribution
Fix a context and a reference distribution over partner types. Define the per-partner influence contribution Equation (4) For a specified context distribution , use . The reference population and context weights are part of the definition, so comparisons across datasets must report them. To define the population quantity, draw and independently, then draw . Under this law, With a context-dependent partner prior, replace by consistently.
A proposed filter retains . This selects partners that change the response relative to the reference mixture. Response entropy measures a different property. If every partner induces the same deterministic limiting response, conditional entropy and mutual information are both zero. A rare partner can have a high KL contribution even when the population average is small. Neither quantity identifies cooperation or benevolence.
We use Intent Certainty (IC) to denote this response influence relative to a reference population. A consistently adversarial partner may be easy to model and highly influential. An IC filter allocates modeling effort according to response influence, independently of whether the partner’s goals are desirable. Evaluation must account for the cost of excluding partners whose responses matter outside the sampled contexts.
4.2 Contextual compatibility
At a context , the directed divergence between two soft-optimal response distributions measures agreement about the current decision. Those distributions already incorporate future interaction through their Bellman values. Thus “short-term” identifies the location of the response, not a myopic objective. Checking every context yields the full-policy discrepancy; checking sampled contexts gives a distribution-dependent empirical measurement with a separate coverage obligation.
4.3 Successor features and occupancy
Fix an ego continuation policy and a bounded feature map on the relevant states or transitions. A finite-horizon successor feature vector is Time is included so that a feature can distinguish when an event occurs. For a finite feature-label set , an alternative is the occupancy measure It has total mass ; is a probability distribution. The discounted infinite-horizon analogue, , requires and has mass . Ordinary KL applies to normalized distributions, not arbitrary signed feature vectors. A norm can instead be used directly on .
Successor-feature representations offer a way to compare consequences independently of one chosen reward parameter [3]. They preserve expected feature visitation, not the complete joint distribution of trajectories. Two processes can have the same marginal occupancy and different temporal correlations. Statements about “all goals” must specify the reward class represented by the features.
4.4 A reward-class guarantee
For this subsection, use a separate common linear reward model , with bounded features and finite . These rewards are bounded but may have either sign; they do not inherit the earlier nonnegative reward convention. Let denote the unregularized return and let best responses maximize that return. For a given initial context, let denote the expected sum without conditioning on a forced first action. Then by linearity of expectation.
Proposition 2 (Full-policy preservation for a feature reward class). If for every ego policy and every initial context, then all policies have the same return against and for every , and their full hard best-response sets agree. More generally, if and , a hard best response to loses at most against , for the specified initial context.
Proof. Exact equality gives equality of each policy’s return. For the approximate case, dual-norm inequality bounds each policy’s return difference by . Insert the returns between the return of a best response to and the return of a best response to . Optimality under makes the middle difference nonpositive, leaving two errors of at most . ◻
The quantifier over all ego policies is substantive. Equality under a single fixed continuation establishes that continuation’s feature agreement, not equality of full best-response sets. Estimating the uniform condition is a separate computational problem.
4.5 Learning and examples
A successor-feature encoder may be designed to place similar outcomes nearby. This is a design constraint, not a consequence of forward continuity alone. If similar outcomes imply similar embeddings, the reverse implication need not hold: a constant embedding satisfies the first implication while collapsing every distinction. A decoder or inverse error bound is needed before nearby embeddings certify similar outcomes.
In Overcooked, clockwise and counterclockwise delivery routes might yield similar counts of onions delivered or dishes served under a particular partner policy. Whether they are interchangeable depends on timing, blocked paths, the ego continuation, and the selected features. In Cleanup, cleaning before or after collecting an apple can yield similar aggregate outcomes while changing early opportunities. These examples explain why the feature map and time horizon must be reported rather than assumed to preserve every strategically relevant consequence.
4.6 Partner-induced world models
A partner also changes the effective transition kernel seen by the ego agent. In a Markov game with joint transition law , a fixed Markov partner induces Model-aware equality means these kernels agree for every . If rewards also agree, backward induction gives equal hard and soft response objects. Transition equality alone does not establish equal incentives when the partner changes rewards. Conversely, similar long-term feature occupancy under one ego policy need not imply the same kernel.
The model-aware target is useful for planning and for identifying changed affordances. A partner blocking a corridor alters which next states are likely after movement. A partner delaying river cleaning may alter resource availability. Similar eventual task completion does not prove identical one-step dynamics. Approximate kernel comparison needs a context-weighting measure, support conventions for KL, and a bound relating model error to the relevant value or response error.
The four targets can therefore disagree without contradiction. A partner can change its motion model while preserving a best response, or preserve a short-term action preference while changing future feature occupancy. Such disagreement should motivate model revision or a task-specific decision, not an automatic claim that the partner is unsafe.
5 Learning representations from response samples
What can be learned without storing a complete response table for every partner? Fix a context and a reference distribution over partner types. Draw and . This defines a response-sampling law. It can be generated by an exact dynamic program in the finite model or approximated by estimated values; those two sources should be evaluated separately.
Let encode a partner at the context and let a positive score be , with contrastive temperature . The contrastive temperature is distinct from the decision temperature . For an integer , with one conditional positive and independent negatives from the marginal response law , use Equation (5) The positive belongs in the denominator. Positivity of the score is necessary for the log ratio; an untransformed dot product or cosine similarity does not ensure it. Sampling each negative partner independently from and then sampling its response produces the required marginal. Conditioning negatives to exclude the anchor, selecting only dissimilar partners, or using a correlated replay batch changes the sampling law and requires another bound.
Let denote the sampled decision context, or the constant context for this conditional calculation. Under the stated sampling scheme, the InfoNCE bound gives [2]. For a fixed response-generating law, optimizing the encoder changes a lower bound and the information retained by its representation, not the underlying itself.
5.1 Prediction and geometry
A low contrastive loss can support response prediction without making every pair of equivalent partners nearby in Euclidean space. For example, append an arbitrary coordinate to an embedding and let the scorer ignore it. The loss is unchanged while distances can become arbitrarily large. Even an optimal loss therefore does not prove an embedding-proximity claim.
A representation can instead be evaluated through a response decoder . Held-out KL estimates diagnose prediction error; a certified uniform contextual bound, in the required direction, supplies the uniform response-substitution guarantee. A geometric clustering claim needs additional constraints—for example, a specified metric-learning objective, decoder regularity, or a canonical representation—and its own proof. The current construction treats geometric consistency as an empirical objective, not a theorem derived from InfoNCE alone.
In Overcooked, partners using the same circulation convention may induce similar useful responses despite different motor behavior. Which direction the ego should travel depends on the layout and joint routes; it cannot be inferred from a partner’s clockwise motion alone. In Cleanup, standing near a river may precede cleaning or apple collection. A representation should be updated from observed, context-dependent effects rather than treating a single action as proof of intention. These are motivating examples, not experiment results.
6 Which uncertainty is worth resolving?
6.1 A Bayes-adaptive interpretation
Bayesian decision problems augment the physical state with a posterior over unknown parameters [19], [20]. The framework draws on the distinction between value obtainable under current knowledge and value gained by obtaining information [21]. This motivates asking whether a partner changes the response, what response is useful now, and what long-term effects remain uncertain.
The model below specifies a latent partner type, a prior, an observation likelihood, and a posterior update. The learner’s predictors summarize aspects of this belief; their agreement alone is insufficient to recover the posterior. We retain the full belief in the formal analysis and leave the sufficiency of compressed predictors to future work. Adding an information-gain reward can change the optimal policy, so its effect on the original task objective must be assessed separately.
6.2 Information gain about a fixed partner type
Can an agent choose interactions that resolve the response distinctions it still does not understand? Let be a fixed latent partner type during an evaluation episode. For example, may be a class under a fixed full-response equivalence relation, provided a generative model over the members of each class is specified. Let be a prior on a finite type set and let the observation alphabet be finite. The action-selection rule depends only on observed history and independent randomization. The observation model is where is the observed interaction history and is the ego’s selected action or experiment. The likelihood must include any state transition, observed partner action, partial observation, and nuisance variables needed by the environment. It is not determined merely by naming an equivalence class.
Write for the current type posterior. The predictive law and posterior update are Updates are defined for observations of positive predictive probability. The proposed expected learning reward is Equation (6)
Proposition 3 (Expected uncertainty reduction). Under the stated finite Bayesian model,
Proof. Expand the expected KL as a sum over using . The posterior log term gives negative expected posterior entropy, and summing the prior log term over observations gives prior entropy. Nonnegativity follows from KL nonnegativity. ◻
A realized entropy difference can be negative even though its expectation is nonnegative. The realized posterior KL is a different nonnegative random reward with the same conditional expectation as equation (6). An implementation should choose and report which estimator it uses.
Information gain is about the agent’s knowledge of a latent type. It is distinct from changing the partner’s behavior and from receiving task reward. One may optimize expected information gain subject to a task or safety constraint, or combine it with explicitly weighted environment reward; neither choice makes the objectives identical. A causal-influence claim requires interventions or identification assumptions in addition to an observation model.
For a fixed finite latent type and a correct model, cumulative expected information gain is at most . Repeated observation therefore cannot supply an unbounded learning signal about that same type. Continual open-ended learning needs new tasks, types, or model structure and a defined mechanism for introducing them. Learned changes in a partition are tracked separately from posterior learning within a fixed partition; comparing entropy across different class spaces is not automatically information gain.
6.3 Using the layers during inference
IC selects response-changing partners relative to the reference mixture; contextual response comparison measures local policy compatibility; successor features measure a declared class of longer-term consequences. These targets can disagree without a logical contradiction. A changing transition kernel can preserve a response, and a shared current action can precede different later outcomes.
A proposed learner maintains these predictors alongside a type posterior. It updates them from trajectories, compares their errors on held-out contexts, and decides whether an additional observation has expected decision or information value. It may defer a decision, collect a targeted observation, or take a task-optimal action. A rule for deferral needs an explicit cost and deadline. No theorem here states that arbitrary layer agreement is necessary or sufficient for action.
7 Computation and learning procedure
7.1 Training a response representation
The proposed training procedure separates response-target construction from representation learning.
Fix the partner training law , contexts, horizon, reward, and temperatures and . Keep a held-out partner set for evaluation.
For a sampled partner , compute the soft-optimal response by backward induction, or obtain a separately evaluated estimate of it. State whether the experiment uses an oracle model or learned values.
Sample a context and a positive action from that partner’s response. Generate independent marginal negatives using independently drawn reference partners.
Evaluate equation (5), including the positive in the denominator, and update the encoder and score model. If a response decoder is trained, report its own objective.
Evaluate contextual response error and task/regularized return on held-out partners. Measure the cost of generating response targets as well as training the encoder.
When only trajectory prefixes are observed, the operational input is an observation encoder or a posterior over partner hypotheses, not an inaccessible complete policy. A policy-input oracle and a trajectory-input implementation should not be reported as the same experiment.
7.2 Response neighborhoods and uncertainty
For a finite reference pool , an estimated partner , and a chosen direction, one can define If a uniform diagnostic is desired, its entropy is for a nonempty set. This reports neighborhood size, not a calibrated posterior over the actual partner. A conditioned prior generally gives different weights. An empty set needs an explicit out-of-model outcome, and successive neighborhoods need not be nested.
This is an oracle response neighborhood: membership uses response distributions rather than embedding distance. It serves as an evaluation diagnostic. Replacing it with embedding distance requires a validated relationship between embedding geometry and response error; it cannot be assumed from the InfoNCE objective. Reversing the KL direction changes the diagnostic and must be reported.
7.3 Joint prediction and refinement
The response encoder may be trained alongside successor-feature and transition predictors. Their targets should be kept explicit: response examples train the normalized contrastive objective, collected trajectories train consequence and dynamics predictors, and an observation likelihood updates the type posterior. If the likelihood is learned, its calibration must be evaluated independently. IC is estimated against the same declared reference partner population throughout a comparison.
The encoder, environment policy, and Bayesian model need not update on the same time scale. Slower model updates or a fixed evaluation snapshot can keep the target interpretable while behavior changes. Evaluation would track response error, regularized loss, task return, and performance on held-out partner families alongside training loss and posterior concentration. Testing the full-policy guarantee requires coverage of the time/state contexts in its bound.
A useful evaluation separates three sources of error: approximating partner-induced dynamics or values; learning a compressed response representation; and selecting actions with an uncertain partner model. Ablations should remove one component at a time while keeping the partner population, context distribution, and budget fixed.
8 Proposed evaluation
8.1 Environments and comparisons
We propose evaluating the method in Overcooked: Counter Circuit,
where coordination depends on shared paths and workstations, and Coin
Game, where partners can have mixed incentives. Existing code provides
two starting points. The value-of-intent repository
contains Overcooked PPO and temporal-CPC implementations with saved
comparison outputs [17]. The
strategic-representation-learning notebook illustrates
response and KL calculations using hand-defined Q-vectors [18]. Neither implements the complete
response-conditioned learner and Bayesian information-gain model studied
here. The evaluation described below has not yet been carried out.
The comparison would use a common partner pool and interaction budget for the response-based encoder, behavioral contrastive embeddings, and state–action co-occurrence representations. An oracle response representation would isolate the cost of estimating the training targets. Varying partner stochasticity, convention, and observation length would test which distinctions the encoders retain. Evaluation would include both familiar-distribution and held-out partner families.
8.2 Measurements
The main measurements would be contextual response KL, regularized response-substitution loss, task return, posterior calibration, and adaptation speed. Comparing predicted information gain with realized posterior updates and type-prediction accuracy would assess the observation model. Embedding plots and neighborhood sizes would provide additional diagnostics; response and return measurements are needed to assess decision quality.
Reproducibility requires recording the seeds, partner and context sampling laws, train/evaluation split, horizon, temperatures, negative-sampling distribution, and computation used to generate targets. Confidence intervals would use independent runs as the sampling unit. A trajectory-input implementation would also need an explicit observation model and belief or recurrent-state construction. The full-state guarantees apply only when that state satisfies the assumptions of the decision model.
8.3 Hypotheses that remain open
Response-based learning may resolve some partner differences earlier than behavioral reconstruction, particularly where many behaviors support the same useful response. Noisy partners may need longer interaction to identify their response-relevant type. We also hypothesize that lower response error improves held-out coordination under the relevant reward model.
A plateau in a learning or entropy curve can indicate successful identification, an uninformative experiment, model misspecification, or optimization failure. It does not by itself establish enough information to act. Zero-shot transfer and resilience to exploitation require direct held-out tests; neither follows merely from a small training loss or stable embedding.
9 What would sustained learning require?
The proposed intrinsic motivation is interaction-centered: an agent seeks observations that resolve how others affect its choices. Predictability can assist adaptation, but a reward for learning about a partner must not be confused with endorsing its goals. Information gain concerns knowledge; a social-influence reward concerns an intervention that changes another agent’s behavior [22]. A causal interpretation needs an intervention or identification model and a counterfactual baseline, beyond a learned association.
9.1 Continual discovery and state marginal matching
Observer-relative open-endedness asks for continuing novelty and learnability [4]. The fixed finite type model above eventually exhausts its information. Continued discovery therefore requires a growing hypothesis space, a curriculum, or changing tasks, together with a mechanism that introduces them and distinguishes learning from model collapse or arbitrary relabeling. None of the finite-model results proves indefinitely improving intelligence.
State marginal matching with mixtures of policies offers a possible outer objective [5]. One could select interaction settings or policy mixtures that visit a specified distribution of strategic situations. This target must be defined separately from the information-gain reward, and its effect on response learning must be tested. No implemented outer loop or convergence result for agreement among abstraction layers is established here.
9.2 Internal components and self-models
Response, successor-feature, and model checks could also be applied to an agent’s own components. Disagreement might reveal an error or a mismatch between stated and enacted objectives. Interpreting this as self-awareness or self-alignment remains an application hypothesis requiring an operational evaluation; agreement among predictors does not by itself establish correctness.
9.3 Interpretation, autonomy, and broader impact
In this framework a “type” is a task-dependent model of response-relevant behavior, not a judgment about a person’s identity or worth. Such a model should be revised when observations conflict with its assumptions. A partner can be predictable without being cooperative, and agreement across representations can coexist with undesirable goals. The formal results cover the specified finite model; extensions to biological, artificial, or mixed teams require evidence about the actual observation and decision setting.
Decision-focused representations may support human–robot coordination, teaching systems, and decentralized planning while avoiding unnecessary behavioral reconstruction. They may also assist manipulation, preemption, or inference about sensitive roles. Representing another agent’s influence does not establish consent, permission to use inferred preferences, or the legitimacy of an intervention. Preserving autonomy therefore concerns how a model is used as well as what it predicts. Response predictability does not solve value conflict or establish universally safe coordination.
10 Conclusion
Strategic representations can be specified by the decisions they preserve. In the finite model studied here, complete soft-optimal responses admit an exact characterization, and uniform directed response error controls entropy-regularized loss. Successor features provide a separate guarantee for a declared linear reward class. These results identify what a learned representation would need to preserve; contrastive training alone does not establish that it does so.
The next question is whether these response, consequence, and belief models can meet useful error targets with less interaction or computation than modeling a partner in full. An answer requires held-out evaluation, calibrated inference, and an accounting of the information used to construct training targets. The proposed experiments would test this tradeoff.
Acknowledgments
I thank Niklas Lauffer for his advice; the CHAI community, including Cassidy Laidlaw, Michelle Li, Karim Abdelsadek, Eli Bronstein, and Cam Allen; and the MATS community, including Alex Cloud and Kola Ayonrinde.
Dedicated to Meridian, Cassiopeia, and Chris.
References
- [1]
-
Niklas Lauffer, Ameesh Shah, Micah Carroll, Michael Dennis, and Stuart Russell. Who Needs to Know? Minimal Knowledge for Optimal Coordination. 2023. https://arxiv.org/abs/2306.09309.
- [2]
-
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation Learning with Contrastive Predictive Coding. 2018, revised 2019. https://arxiv.org/abs/1807.03748.
- [3]
-
Chris Reinke and Xavier Alameda-Pineda. Successor Feature Representations. Transactions on Machine Learning Research, 2022. https://openreview.net/forum?id=MTFf1rDDEI.
- [4]
-
Edward Hughes et al. Open-Endedness is Essential for Artificial Superhuman Intelligence. 2024. https://arxiv.org/abs/2406.04268.
- [5]
-
Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Ruslan Salakhutdinov, and Sergey Levine. State Marginal Matching with Mixtures of Policies. 2019.
- [6]
-
Stefano V. Albrecht and Peter Stone. Autonomous agents modelling other agents: A comprehensive survey and open problems. Artificial Intelligence, 258:66–95, 2018.
- [7]
-
Chris Baker, Rebecca Saxe, and Joshua Tenenbaum. Bayesian theory of mind: Modeling joint belief-desire attribution. Proceedings of the Cognitive Science Society, 33, 2011.
- [8]
-
George W. Brown. Iterative solution of games by fictitious play. 1951.
- [9]
-
Micah Carroll et al. On the utility of learning about humans for human–AI coordination. 2020. https://arxiv.org/abs/1910.05789.
- [10]
-
Jakob N. Foerster et al. Learning with Opponent-Learning Awareness. 2018. https://arxiv.org/abs/1709.04326.
- [11]
-
Lewis Hammond et al. Multi-Agent Risks from Advanced AI. 2025. https://arxiv.org/abs/2502.14143.
- [12]
-
Hengyuan Hu et al. Off-Belief Learning. 2021. https://arxiv.org/abs/2103.04000.
- [13]
-
Hengyuan Hu, Adam Lerer, Alex Peysakhovich, and Jakob Foerster. Other-Play for Zero-Shot Coordination. 2021. https://arxiv.org/abs/2003.02979.
- [14]
-
Julian Jara-Ettinger. Theory of mind as inverse reinforcement learning. Current Opinion in Behavioral Sciences, 29:105–110, 2019.
- [15]
-
Neil C. Rabinowitz et al. Machine Theory of Mind. 2018. https://arxiv.org/abs/1802.07740.
- [16]
-
Roberta Raileanu, Emily Denton, Arthur Szlam, and Rob Fergus. Modeling Others using Oneself in Multi-Agent Reinforcement Learning. 2018. https://arxiv.org/abs/1802.09640.
- [17]
-
Sandy Tanwisuth.
value-of-intent. Code and notebook outputs, commit 686557b2f99d419dd5c07f142817060234207203. https://github.com/sandguine/value-of-intent/tree/686557b2f99d419dd5c07f142817060234207203. - [18]
-
Sandy Tanwisuth.
strategic-representation-learningdemonstration, commit 2ab81418a97c2410d05a3e53dcce7c92858a2cd8. https://github.com/sandguine/strategic-representation-learning/tree/2ab81418a97c2410d05a3e53dcce7c92858a2cd8. - [19]
-
Richard Bellman and Robert Kalaba. A mathematical theory of adaptive control processes. Proceedings of the National Academy of Sciences, 45(8):1288–1290, 1959.
- [20]
-
James John Martin. Bayesian Decision Problems and Markov Chains. 1967.
- [21]
-
Aly Lidayan, Michael Dennis, and Stuart Russell. BAMDP Shaping: A Unified Framework for Intrinsic Motivation and Reward Shaping. 2025 revision. https://arxiv.org/abs/2409.05358.
- [22]
-
Natasha Jaques et al. Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning. 2019. https://arxiv.org/abs/1810.08647.