Revised working edition
Strategic Openendedness
Original version: April 23, 2025
Abstract
What should an agent learn when the environment offers no single stable task objective? This proposal studies strategic information acquisition: learning which features of a partner change the ego agent’s useful responses. It combines response-based representation learning with a Bayesian information-gain objective over fixed response-relevant partner types. The mathematical scope is a finite-horizon model with explicit priors, likelihoods, and full-policy response objects. A proposed MettaGrid evaluation separates information acquisition, controllable social influence, and task performance, and tests four representation layers through ablations and held-out partners. The empirical record remains under reconciliation with experiment repositories; no unverified result is reported as a measurement. The open question is whether curricula that reveal new strategic distinctions can support sustained, useful learning beyond a fixed finite type model.
Contents
Working draft for author review. Reconstructed October 9, 2026.
Reconstruction note. This standalone source reconstructs
Strategic_Openendedness.pdf. Original PDFs are preserved. Editorial cleanup, documented corrections, and author-approved changes in mathematical scope are identified in the accompanying review. It is not an accepted final revision or a claim of experimental reproduction.
1 What is learned under strategic pressure?
Two-player zero-sum tasks provide a clear objective and an adversarial test of behavior. Their successes motivate a question: which aspects of that training pressure can be retained when partners, goals, and environmental structure change? The proposal does not assume that open-ended settings inherit the same guarantees. Instead, it asks whether learning to predict response-relevant influence supplies a useful objective when a single extrinsic task is insufficient.
The source describes influence as the capacity to shape and be shaped by others. This revision separates three quantities that this phrase can obscure. Response influence measures how a partner changes a specified ego response relative to a reference population. Information gain measures what the ego learns about a latent partner type. Controllable social influence concerns the effects of interventions on another agent’s behavior. None is identical to environment reward or a guarantee of desirable coordination.
The Unified Strategic Representation Learning framework [5] motivates four learning targets: IC, contrastive strategic coding (CSC), successor-feature representations (SFR), and model-aware strategic representations (MASR). These range from contextual response differences to longer-term consequences and induced dynamics. The proposal is to use them to ask better questions about partners and select informative interactions. It does not prove that agreement among the layers constitutes alignment or that predictable influence is inherently benign.
1.1 Contributions and evidence status
The present contribution is a specified learning problem and
evaluation design. It includes a full-policy response target, a
conditional return-loss guarantee, an explicit Bayesian data-generating
model, and a curriculum/ablation protocol. The source’s references to
completed MettaGrid tests conflict with its expected-results and
limitation sections. The public value-of-intent repository
contains Overcooked training code and saved notebook comparison outputs
[17]. Its inspected CPC variants predict future
observation features through an auxiliary loss; they do not implement
the information-gain objective defined here. A separate repository
supplies hand-defined response/KL demonstrations [18]. Until runs and outputs are matched to the
MettaGrid claims, that material is presented as a proposed evaluation.
This preserves the research direction without supplying invented
measurements.
2 Which policy must a representation preserve?
2.1 A finite-horizon model
Let , let and be finite, and assume every action in the nonempty set is available at every state. A partner policy is indexed by . Fixing induces an MDP for the ego agent with transition law and expected reward . Terminal value is zero. All partners are compared on the same state/action spaces and horizon. A known finite memory state of a partner can be included in ; an arbitrary partially observed or history-dependent partner is not covered without constructing an adequate state.
An ego policy maps each time/state context to a distribution over actions. We compare policies at every context, including states that have zero probability under a particular initial distribution. This convention gives a policy equivalence independent of which starting states happen to be sampled. A weaker on-distribution relation would require its own support and coverage assumptions.
For , define the entropy-regularized objective from context by The reward term is task value; the entropy term is an explicit modeling choice. It is not uncertainty about whether the Q-function is correct. Throughout, uses natural logarithms, and all compared policies use the same finite positive temperature.
2.2 A full soft-optimal response
The backward recursion is Equations (1), (2), (3) This gives a complete time/state policy, rather than a current action under an unspecified continuation. Finite rewards and horizon ensure finite values and a positive probability for every action.
Lemma 1 (Soft Bellman variational identity). For and , let . Then Consequently is the unique maximizer of the first two terms.
Proof. Substitute into the KL definition, with . Rearranging gives the identity. KL is nonnegative and is zero precisely at . ◻
Backward induction using Lemma 1 proves that equation (3) is the unique entropy-regularized optimal policy at every context. The construction permits general-sum environments after the partner is fixed. It is a best-response construction, not a theorem that simultaneous learning converges to an equilibrium.
2.3 Exact equivalence of full policies
Definition 1 (Soft-policy equivalence). Write if for every .
Theorem 1 (Complete characterization). The following statements are equivalent:
;
at each context there is a real such that for all ;
.
Proof. At a fixed context, equality of softmax distributions implies equality of the ratios for actions , hence equality of . Subtracting a fixed reference action gives the action-independent difference. Conversely, an additive constant cancels in softmax. Zero KL is equivalent to equality of the finite-support distributions. Repeating the argument at every context proves the result. ◻
The preimages of are therefore the coarsest exact partition preserving this full response. An abstraction preserves the response exactly when implies . It may retain more distinctions. This condition constrains which partners can share an embedding; it does not force equivalent partners to have equal or nearby embeddings.
2.4 Hard best responses are a different relation
For the unregularized objective, define , , and . Let .
Proposition 1 (Full hard-policy best-response sets). A Markov policy is optimal from every time/state context iff it puts probability only on at every context. Consequently two partners have the same sets of such optimal policies iff all their agree.
Proof. Backward induction shows sufficiency. For necessity, a policy optimal at every context must have optimal continuation values. Assigning positive mass to a strictly suboptimal action then strictly lowers its expected value at that context. Equality of the local maximizing sets gives equality of the policy sets. If a maximizing action belongs to only one set at some context, a policy selecting it there and other maximizing actions elsewhere witnesses a difference. ◻
In particular, action-independent shifts of at every context are sufficient for full hard-policy equivalence. They are not necessary. Nor does Theorem 1 by itself identify hard and soft equivalence: its action values include entropy in future decisions. Matching greedy sets of those soft values is not a proof about unregularized optimal policies. This separates the preserved target from the full-policy SER motivation [1].
3 What does approximate response preservation guarantee?
Define the directed discrepancy Its zero set is the exact equivalence relation above. For positive tolerance, the condition is a similarity condition, not an equivalence relation. KL is asymmetric, and thresholding even a symmetric distance generally fails transitivity. For binary probabilities , adjacent directed divergences are below , whereas the first-to-last divergence exceeds it. Any clustering procedure must specify how it enforces its desired within-cluster guarantee.
Theorem 2 (Regularized response-substitution loss). Fix a partner , an initial state distribution , and any Markov policy . Write . Then If every contextual divergence on the right is at most , the loss is at most .
Proof. Apply Lemma 1 at with and . Substitute equation (1), take the expectation under the actual partner and policy , and sum over . The next-state value terms cancel with successive current-state values, and . The remaining reward and entropy terms form . Bounding each of the divergences gives the inequality. ◻
To substitute the optimal response to partner when the actual partner is , the sufficient condition is . The direction matters. A bound measured only under an unrelated training-state distribution does not supply this uniform guarantee. The result concerns regularized return, not raw reward alone, human welfare, or equilibrium selection.
4 An explicit data-generating model for strategic types
4.1 Finite candidate partners and response-relevant classes
Fix a finite candidate set of partner policies, a prior , a finite-horizon game, and the reward/temperature used to define a full response. Draw one partner at the beginning of an episode and keep it fixed. Draw independently of and observe it; a type-dependent initial state would instead require an initial likelihood update. Let be a fixed map to response-relevant types, such as the exact full-policy equivalence classes. The latent target is , with prior .
For a fully observed baseline, let the ego choose , let the partner take action , and let the next state follow a known game kernel . Observe . Conditional on the known current state, The posterior over candidate policies is The action-selection rule is a function of the observed history and declared randomization. Conditioning on its chosen action therefore does not add an unmodeled partner-dependent selection likelihood. A different data-collection policy must be accounted for explicitly.
4.2 Why a class still needs an observation model
Different partners in the same response class may produce different observations. Their class likelihood is the history-dependent mixture when . Retaining the within-class posterior prevents a fixed representative from silently replacing a heterogeneous class. The information-gain objective below concerns , not necessarily the exact identity .
For partially observed MettaGrid variants, observations may hide partner actions or state. Then the model must marginalize hidden actions and maintain a posterior over hidden state and partner together. A recurrent embedding is a possible approximation but is not automatically the exact Bayesian filter or a sufficient state for the full-policy theorem. Unknown dynamics, nonstationary partners, or learned likelihoods introduce model error and require held-out calibration checks.
4.3 A finite informative example
Suppose two response types have equal prior probability and a binary observation satisfies and . The predictive observation is uniform. After observing either outcome, the posterior assigns probability to one type and to the other. Expected information gain is An action whose likelihood is the same for both types has zero information gain, even if it changes the environment substantially. This calculation illustrates the stated model; it is not a MettaGrid experimental result.
5 Learning progress as information gain
Can an agent choose interactions that resolve the response distinctions it still does not understand? Let be a fixed latent partner type during an evaluation episode. For example, may be a class under a fixed full-response equivalence relation, provided a generative model over the members of each class is specified. Let be a prior on a finite type set and let the observation alphabet be finite. The action-selection rule depends only on observed history and independent randomization. The observation model is where is the observed interaction history and is the ego’s selected action or experiment. The likelihood must include any state transition, observed partner action, partial observation, and nuisance variables needed by the environment. It is not determined merely by naming an equivalence class.
The predictive law and posterior are Updates are defined for observations of positive predictive probability. The proposed expected learning reward is Equation (4)
Proposition 2 (Expected uncertainty reduction). Under the stated finite Bayesian model,
Proof. Expand the expected KL as a sum over using . The posterior log term gives negative expected posterior entropy, and summing the prior log term over observations gives prior entropy. Nonnegativity follows from KL nonnegativity. ◻
A realized entropy difference can be negative even though its expectation is nonnegative. The realized posterior KL is a different nonnegative random reward with the same conditional expectation as equation (4). An implementation should choose and report which estimator it uses.
Information gain is about the agent’s knowledge of a latent type. It is distinct from changing the partner’s behavior and from receiving task reward. One may optimize expected information gain subject to a task or safety constraint, or combine it with explicitly weighted environment reward; neither choice makes the objectives identical. A causal-influence claim requires interventions or identification assumptions in addition to an observation model.
For a fixed finite latent type and a correct model, cumulative expected information gain is at most . Repeated observation therefore cannot supply an unbounded learning signal about that same type. Continual open-ended learning needs new tasks, types, or model structure and a defined mechanism for introducing them. Learned changes in a partition are tracked separately from posterior learning within a fixed partition; comparing entropy across different class spaces is not automatically information gain.
6 When does a partner change the response?
Fix a context and a reference distribution over partner types. Define the per-partner influence contribution Equation (5) For a specified context distribution , use . The reference population and context weights are part of the definition, so comparisons across datasets must report them. The population quantity is when context and partner are sampled independently according to the stated design. With a context-dependent partner prior, replace by consistently.
A proposed filter retains . This selects partners that change the response relative to the reference mixture. It does not identify cooperation, benevolence, or low uncertainty. Report conditional response entropy separately. If every partner induces the same deterministic limiting response, conditional entropy is zero while mutual information is zero too. A rare partner can have a high KL contribution even when the population average is small.
The term Intent Certainty is retained for continuity with the source, but the mathematical target is response influence relative to a reference population. A consistently adversarial partner may be easy to model and highly influential. That does not make its outcomes desirable or justify trusting it with unrelated tasks. Filtering is a modeling-budget choice whose effect on excluded but strategically important partners needs evaluation.
7 Representation layers and their roles
7.1 Contrastive strategic coding
For a context and a partner drawn from the declared reference law, sample a positive action from the partner-induced soft-optimal response. Draw negative actions independently from the marginal response law. A positive score such as defines a normalized InfoNCE loss with the positive and all negatives in its denominator [2]. The objective can retain response-predictive information. It does not prove that all equivalent partners occupy a small Euclidean neighborhood: scorer-irrelevant coordinates can preserve loss while changing geometry.
A practical model takes observed trajectories as input and predicts a contextual response distribution. Its response error should be measured independently of its contrastive loss. A true-policy input, an oracle Q-function, and a learned trajectory encoder provide different levels of information and should be reported as separate conditions.
7.2 Successor features
A successor vector under ego continuation is It measures expected feature consequences, in the spirit of successor features and successor-feature representations [6], [3]. A norm can compare such vectors. If KL is used, define a probability law over feature labels and normalize occupancy; an arbitrary signed feature vector is not a probability distribution. Full-policy equivalence for a feature reward class requires agreement across ego policies, not just one current continuation.
7.3 Model-aware strategic representations
A partner induces a transition kernel . A model-aware layer learns that kernel or a declared abstraction of it. Comparing it with a baseline requires a state/action weighting law and support convention. Equality of a scalar KL distance to a baseline does not establish equality between two kernels; opposite deviations can have the same distance. Evaluate pairwise prediction error and consequences for planning directly.
Model-aware changes can reveal blocked routes, delayed effects, or altered resource access. They are not automatically causal effects of a chosen ego intervention. Social-influence rewards and causal game models motivate that separate question [12], [13], [14], [7]. An influence evaluation must specify what is intervened on, what is held fixed, and what observed or simulated counterfactual supports the comparison.
7.4 A modular learning loop
Select a partner or interaction setting using a declared curriculum and keep the evaluation distribution fixed for comparisons.
Collect trajectories, record observations and actions, and update the candidate-policy/type posterior under the likelihood model.
Compute predicted information gain for available experiments. Use it as a declared auxiliary objective or subject it to task and safety constraints.
Update response encoders, successor predictors, and transition models using their separate losses.
Evaluate posterior calibration, contextual response error, task return, and held-out transfer before adapting the curriculum.
If separate layer types and rewards are used, a weighted sum may double-count correlated information. A joint target supplies a different objective. The selected factorization and weights must be part of the experiment configuration, rather than inferred from the number of layers.
8 Proposed experimental design in MettaGrid
8.1 Environments and partner populations
The original design places two or three agents in a grid environment with shared resources, dynamic tiles, object triggers, and causal gates. Interaction zones may include doors, shared levers, or tiles that change access to later states. Controlled perturbations include object swaps, delay channels, and restricted observations. Each layout and intervention should be versioned so that the policy’s observation and action spaces are reconstructible.
Use fixed-policy partners with known behavior, such as resource collection or deliberate delay, to establish identifiable baseline tasks. A second condition may sample partners from a meta-trained family. Training, validation, and held-out partners must be disjoint as specified by the experiment, and the prior used by the inference model must be reported. A partner should remain fixed within each Bayesian evaluation episode; adaptation within an episode requires a different latent-transition model.
8.2 Four evaluation blocks
- Influence clarification.
-
Pair the learner with partners whose response effects differ under known contexts. Test whether the IC estimate recovers the intended reference-relative contribution, without interpreting it as goodness or certainty.
- Strategic disambiguation.
-
Provide multiple affordances whose useful response depends on the latent partner type. Measure calibrated inference and the regret of acting under an uncertain or misidentified type.
- Successor-feature prediction.
-
Test longer-term feature consequences under specified ego continuations, including delayed resource use and access to later states. Report the feature map and whether the reward lies in its represented class.
- Model refinement.
-
Change transition effects associated with partner behavior, such as a trigger opening a route. Evaluate conditional next-state predictions and planning consequences under the altered model.
Solo or fixed-partner controls can test whether improvements depend on interaction rather than additional training time. Shared-space task completion should be evaluated separately from the learning objective.
8.3 Metrics and sampling
Record posterior log loss and calibration, expected information gain, realized posterior KL, type-disambiguation time, contextual response divergence, regularized response loss, and raw task return. For long-term and model-aware layers, report held-out feature or transition prediction error. Cross-layer mutual information or agreement is descriptive; its magnitude depends on class balance and does not by itself prove validity.
For zero-shot transfer, define the held-out partner families and whether the encoder, policy, likelihood, or posterior are allowed to adapt. For controllable influence, use a separate intervention protocol with a specified counterfactual comparison. Trajectory analysis, attention inspection, and embedding plots can help interpret failures but do not replace quantitative evaluation.
Every result should include repository/commit provenance, environment versions, partner checkpoints, training and evaluation seeds, interaction budgets, hyperparameters, uncertainty estimates, and the rule for selecting checkpoints. The proposed data-generating process is explicit above; no implementation is claimed to reproduce it until code and configurations are matched to those assumptions.
8.4 An uncertainty-based curriculum
The source proposes selecting partners near the learner’s abstraction frontier: interactions that may reveal a new useful distinction. With the fixed-type model, select by predicted information gain or a declared approximation. Evaluate this selection rule against uniform partner sampling at a common budget. Curriculum selection changes which data are seen; maintaining a fixed test distribution avoids confusing that change with improved generalization.
Early interactions can use identifiable fixed partners, followed by more ambiguous observation or richer dynamics. The source also proposes difficult ambiguous partners early in training. These schedules are alternative experimental designs, not an established ordering theorem. Record the schedule and measure both learning progress and failures caused by model misspecification.
9 Ablations and hypotheses
9.1 Layer removal
Remove IC, CSC, successor features, or the model-aware predictor separately while holding the task, partner distribution, and compute budget fixed. Removing IC tests whether reference-relative response screening helps; removing CSC tests response supervision; removing successor features tests long-term consequence prediction; removing MASR tests the value of explicit transition modeling. These ablations should not be described as proving a layer necessary outside the tested setting.
9.2 Reward and curriculum controls
Compare the information-gain objective with the source’s proposed alternatives: random-network-distillation-style novelty, curiosity or prediction error, and empowerment. Define each implementation and match training budgets. Include no-curriculum and randomized-curriculum controls. These comparisons distinguish learning about response-relevant types from generic exploration and from the effect of collecting different data.
| Ablation | Question to measure |
|---|---|
| Remove IC | Does screening by response contribution improve use of interaction? |
| Remove CSC | Does direct response prediction matter beyond behavioral encoding? |
| Remove SFR | Are delayed feature consequences predicted less accurately? |
| Remove MASR | Does changed world dynamics cause more planning error? |
| Replace information gain | Is any benefit specific to inference about partner type? |
| Remove or randomize curriculum | Is the selection rule useful at a fixed budget? |
The source’s qualitative checkmark table is replaced by evaluation questions because the accompanying text labels outcomes as expected. No measured superiority is asserted.
9.3 Expected patterns and falsification
The hypothesis is that type inference becomes more accurate as informative interaction accumulates, and that response error falls when the inferred type is useful for the task. A fixed finite type model predicts eventual exhaustion of its information, not perpetual positive reward. Cross-layer alignment may emerge in some environments, but disagreement can also expose different valid prediction targets.
Held-out coordination gains, improved sample efficiency, and robust adaptation remain empirical questions. A baseline may match or exceed the proposed method. Failure to calibrate the posterior, loss of task performance, or no benefit from response supervision would count against the intended advantage. The evaluation should report these outcomes rather than treating every entropy decrease as progress.
10 Limitations and future directions
10.1 Empirical validation and computation
The original proposal identifies missing empirical validation as a limitation. Matrix games, small gridworlds, or simple cooperative tasks can isolate the likelihood and response targets before complex MettaGrid studies. Introduce layers incrementally and compare against appropriate baselines. Approximate successor and transition models, amortized inference, sparse updates, and slower update schedules are engineering options; their errors and costs must be measured.
10.2 Continual discovery
A finite latent model cannot be indefinitely open-ended. A growing partner population, new layouts, or revised hypotheses may create new learning problems, but changing the target also changes the information-gain interpretation. Diversity regularization and adversarial partners are candidate mechanisms to prevent collapse, not established convergence guarantees. Observer-relative novelty and learnability [4] provide motivation; no theorem here guarantees unbounded novelty or superhuman intelligence.
10.3 Humans, groups, and interpretation
Human feedback may help identify which response distinctions matter, but inferred preferences are not a substitute for consent or autonomy. Proposed tools include interpretable type assignments, influence trajectories, and explanations of uncertainty. Extending to larger groups requires a model of coalitions and interdependent types rather than assuming pairwise terms suffice. Markets, ecosystems, or other processes might be modeled as virtual partners only after defining the relevant generative model.
The source also proposes combining strategic abstraction with planning and episodic memory, studying emergent conventions, and making an agent’s behavior easier for others to interpret. These directions ask how agents learn together rather than merely optimize a fixed task. Which new distinctions are worth discovering, and how can that discovery remain useful to the other agents involved?
11 Annotated connections
The following connections preserve the source’s literature map while separating cited ideas from this proposal’s results.
- Strategic abstraction.
-
The related Unified manuscript [5] organizes the four representation targets. Strategic equivalence [1] motivates preserving response distinctions. Neither supplies a free computational guarantee for a learned approximation.
- Representation learning.
-
Contrastive predictive coding [2] supplies the normalized loss and its sampling-dependent information bound. Successor features and richer feature representations [6], [3] motivate reward-class transfer and long-term prediction.
- Theory of mind and causality.
-
Incomplete-information influence diagrams [7], causal games [13], and causal reinforcement learning [14] motivate explicit belief and intervention models. Association, posterior learning, and intervention effects remain different quantities.
- Social learning.
-
The social-path perspective [8] motivates curricula produced through interaction. It does not establish that a particular intrinsic reward yields desirable group outcomes.
- Goals and recursion.
-
Maximum-entropy goal-directedness [9] provides a distinct behavioral diagnostic. General metagames [10] supply historical context for recursive strategic reasoning, without implying that the present encoder implements that recursion.
- Exploration and shaping.
-
Bayes-adaptive shaping [11] and social-influence rewards [12] motivate related objectives with their own assumptions. Control-as-inference [15] connects probabilistic response models and control. Unsupervised environment design [16] motivates a curriculum over tasks; its guarantees do not automatically extend to a curriculum over learned partner types.
- Open-endedness.
-
Novelty and learnability [4] motivate the question of continued discovery. The finite information budget proved here makes the need for a changing task or hypothesis space explicit.
Acknowledgments and revision provenance
The original source thanks members of Softmax, particularly Emmett Shear, Jack Heart, and David Bloomin, for conversations, and describes the work as independently authored and unaffiliated. Those credits are preserved without implying current affiliation or endorsement. Personal contact addresses are omitted.
This revision preserves the motivation, four-layer method, MettaGrid protocol, ablation questions, expected outcomes, limitations, future directions, and literature map. It replaces changing-cluster entropy with author-approved information gain under an explicit model, separates response influence from controllable influence, and labels unverified empirical claims accordingly. The accompanying mathematical review records the source discrepancies and assumptions.
References
- [1]
-
Niklas Lauffer, Ameesh Shah, Micah Carroll, Michael Dennis, and Stuart Russell. Who Needs to Know? Minimal Knowledge for Optimal Coordination. 2023. https://arxiv.org/abs/2306.09309.
- [2]
-
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation Learning with Contrastive Predictive Coding. 2018, revised 2019. https://arxiv.org/abs/1807.03748.
- [3]
-
Chris Reinke and Xavier Alameda-Pineda. Successor Feature Representations. Transactions on Machine Learning Research, 2022. https://openreview.net/forum?id=MTFf1rDDEI.
- [4]
-
Edward Hughes et al. Open-Endedness is Essential for Artificial Superhuman Intelligence. 2024. https://arxiv.org/abs/2406.04268.
- [5]
-
Sandy Tanwisuth. A Unified Theory of Strategic Representations. Independent working draft, 2025.
- [6]
-
Andre Barreto et al. Successor Features for Transfer in Reinforcement Learning. 2018 version. https://arxiv.org/abs/1606.05312.
- [7]
-
Jack Foxabbott et al. A Causal Model of Theory-of-Mind in AI Agents. 2024, as attributed in the original manuscript. Matching anonymous submission: https://openreview.net/pdf?id=ASA2jdKtf3. Author/venue metadata remains to be verified against the final record.
- [8]
-
Edgar A. Duenez-Guzman et al. A social path to human-like artificial intelligence. Nature Machine Intelligence, 5(11):1181–1188, 2023.
- [9]
-
Matt MacDermott, James Fox, Francesco Belardinelli, and Tom Everitt. Measuring Goal-Directedness. NeurIPS, 2024. https://arxiv.org/abs/2412.04758.
- [10]
-
Nigel Howard. “General metagames: An extension of the metagame concept. In Game Theory as a Theory of Conflict Resolution, pp. 261–283. Springer, 1974.
- [11]
-
Aly Lidayan, Michael Dennis, and Stuart Russell. BAMDP Shaping: A Unified Framework for Intrinsic Motivation and Reward Shaping. 2024. https://arxiv.org/abs/2409.05358.
- [12]
-
Natasha Jaques et al. Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning. 2019. https://arxiv.org/abs/1810.08647.
- [13]
-
Lewis Hammond et al. Reasoning about causality in games. Artificial Intelligence, 320:103919, 2023.
- [14]
-
Elias Bareinboim, Junzhe Zhang, and Sanghack Lee. An Introduction to Causal Reinforcement Learning. Technical report R-65, first version 2024. https://www.causalai.net/r65.pdf.
- [15]
-
Sergey Levine. Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review. 2018. https://arxiv.org/abs/1805.00909.
- [16]
-
Michael Dennis et al. Emergent complexity and zero-shot transfer via unsupervised environment design. NeurIPS, 33:13049–13061, 2020.
- [17]
-
Sandy Tanwisuth.
value-of-intentrepository, commit 686557b2f99d419dd5c07f142817060234207203. Code and saved notebook outputs inspected October 9, 2026; experiments not reproduced. https://github.com/sandguine/value-of-intent/tree/686557b2f99d419dd5c07f142817060234207203. - [18]
-
Sandy Tanwisuth.
strategic-representation-learningdemonstration, commit 2ab81418a97c2410d05a3e53dcce7c92858a2cd8. https://github.com/sandguine/strategic-representation-learning/tree/2ab81418a97c2410d05a3e53dcce7c92858a2cd8.