Revised working edition

The Effects of Uncertainty on Behavioral Strategic-Policy Adaptation and the Neural Representations of Human Strategic Decision-making in Sequential Social Dilemmas

Sandy Tanwisuth

Original research proposal: August 2020
Editorial revision for author review: 9 October 2026

Abstract

How do people adapt their behavior when both the environment and other agents are uncertain? Existing computational studies of strategic social learning motivate a closer examination of decisions that unfold across space and time. This proposal studies how social uncertainty, environmental uncertainty, and their interaction affect changes between cooperative and defecting policies, and how the brain represents the information needed for those changes.

We propose online behavioral experiments followed by functional magnetic resonance imaging (fMRI). Human participants would interact with several types of pretrained opponents in the sequential social dilemmas Cleanup and Harvest. The central behavioral hypothesis is that greater volatility, and hence lower reliability of current predictions, increases the frequency of policy adaptation. Model-based fMRI and connectivity analyses would test hypotheses about reliability tracking in ventrolateral prefrontal cortex, temporoparietal junction, and rostral cingulate cortex; state-value representation in ventromedial prefrontal cortex; action-value representation in premotor regions; and mentalizing in medial prefrontal cortex, posterior superior temporal sulcus, and temporoparietal junction. Models in the partially observable Markov decision process family would provide candidate computational accounts. A further aim is to study whether learned representations of the task help explain how complex social observations are reduced to information relevant to action. The intended contribution is a quantitative bridge between multiagent learning, human strategic decision-making, and research on value alignment.

Contents

Status of this document

This is an editorial revision of the August 2020 research proposal. It preserves the proposed behavioral and neuroimaging studies, all three aims, the original hypotheses, and the cited literature. The proposed studies and predicted neural correlates are not reported as completed experiments or established results. Historical reference details are retained; incomplete entries are identified at the end.

Keywords: reinforcement learning; social decision-making; fMRI; sequential social dilemmas; partially observable Markov decision processes.

1 Specific aims

Successful interaction requires adapting to uncertainty about both other agents and the environment [19], [14], [16], [26]. Studies have identified neural signals related to reinforcement learning during strategic social interaction [10], [33], [13]. Many experimental tasks use matrix games, where the relevant choices and payoffs can be specified in a compact form [30]. Such settings are useful, but they do not by themselves capture the spatial and temporal structure of more extended social decisions. Matrix games need not be zero-sum or involve strictly opposing interests; the original proposal’s description should be read as referring to its motivating examples rather than the definition of a matrix game.

This proposal asks three connected questions. How do social uncertainty, environmental uncertainty, and their interaction affect policy adaptation? Which neural systems track those changes, and how do they interact? How does the brain represent a complex social environment in a form that supports action selection?

The questions connect game-theoretic multiagent reinforcement learning, social cognitive neuroscience, and biologically inspired artificial intelligence. Better accounts of these mechanisms may eventually inform studies of social-function difficulties and possible interventions, but the proposed experiments would not alone establish a clinical application. They may also inform value-alignment research by clarifying behavioral algorithms and neural representations involved in social decisions [1], [24].

The principal hypothesis is that switching between cooperative and defecting policies depends on the reliability of social and environmental predictions and on their interaction. Greater volatility is hypothesized to reduce the reliability of a current policy and increase the rate of adaptation. This is a proposed relationship to test, not an assumption that all forms of uncertainty are equivalent.

The study would begin with online behavioral tasks in which participants interact with different pretrained agents. A subsequent fMRI experiment would use the same task family. Model-based fMRI [23] and connectivity analyses would examine candidate neural computations and relationships between regions. Complementary analyses not tied to a particular behavioral model would investigate representations of the environment and other agents.

1.1 Aim 1: uncertainty and strategic-policy adaptation

How do the behavior of other agents and the availability of resources change the way a person adapts a strategy? Prior work on mentalizing and social learning motivates the hypothesis that uncertainty in either source, as well as their interaction, affects the frequency of policy changes [26], [3], [7].

The proposed social manipulation varies the types of opponents a participant encounters: always-cooperating agents, always-defecting agents, random agents, and agents with stable pretrained reinforcement-learning policies. Environmental conditions vary the availability of shared resources within sequential social dilemmas. These are temporally extended settings in which individual and collective incentives may conflict [25], [6], [20].

Two task families provide different incentive structures [17]. In a public-goods game, an individual can incur a cost to produce a shared benefit [9]. In a tragedy-of-the-commons setting, an individual can gain by consuming a shared resource, while aggregate consumption can deplete it [12]. The proposal uses Cleanup as a public-goods task [14], [16] and Harvest as a commons task [19], [14], [16]. Comparing these tasks would help distinguish effects that depend on a particular incentive structure from effects shared across them.

The proposed outcome is the rate of adaptation between cooperative and defecting policies. Its operational definition, the schedules that manipulate uncertainty, and the statistical model for the interaction of social and environmental factors must be fixed before data collection. The original proposal does not specify those details or a sample size.

1.2 Aim 2: neural computations and network relationships

Which neural signals track the reliability of a social strategy, and how do they relate to value computations? Work on observational learning implicates ventrolateral prefrontal cortex (vlPFC), temporoparietal junction (TPJ), and rostral cingulate cortex (RCC) in computations associated with imitation and goal emulation [3]. Related accounts of arbitration and social decision-making motivate testing whether these regions track the reliability of competing policy predictions [22], [4], [18], [29].

The proposal also predicts state-value signals in ventromedial prefrontal cortex (vmPFC) and action-value signals in premotor regions [5]. It includes mentalizing hypotheses involving medial prefrontal cortex (mPFC), posterior superior temporal sulcus (pSTS), and TPJ. These region-level expectations guide tests; they do not identify a neural computation uniquely from an observed correlation.

Candidate behavioral and neural accounts would come from the partially observable Markov decision process (POMDP) family, including models that represent uncertainty about another agent [26], [2], [8], [31], [11]. Model-based fMRI would relate fitted computational variables to neural responses, while connectivity analyses would examine relationships among the implicated regions. The model family, fitting procedure, model-comparison criteria, and specific connectivity analysis remain design choices that the original proposal leaves open.

1.3 Aim 3: representations supporting strategic action

How can a complex observation of the social world be represented by a smaller set of distinctions relevant to action? Work relating biological and artificial neural networks suggests that nonlinear transformations can separate features useful for behavior [5], [32], [15]. The proposal asks whether a similar approach can clarify representations of the environment, other adaptive agents, and their combined consequences for a decision.

The working hypothesis is that social decision-making draws on compressed representations of information such as state value, action value, and the behavior of self and others. Research on learned task-state representations and predictive maps provides a basis for investigating this possibility [21], [28]. The proposed method uses Contrastive Unsupervised Representations for Reinforcement Learning (CURL) [27] as a candidate way to learn task representations and compare them with neural data.

Such a comparison would require an explicit representation-comparison analysis and controls for simpler accounts of the observed signals. Similarity between model and neural representations would be evidence to interpret, rather than proof that the brain implements the learning algorithm. The open question is which learned distinctions explain neural and behavioral variation beyond those already captured by value and task variables.

Design and bibliographic questions retained for author review

The original document is a proposal, not a methods report. It leaves participant counts, recruitment, exclusions, power analysis, uncertainty schedules, preregistration, imaging parameters, multiple-comparison control, and exact model-fitting procedures unspecified. These must not be filled in as if experiments had occurred. The meaning of policy adaptation also needs an operational definition that separates a change in an inferred strategy from a single exploratory action.

Several reference entries were incomplete in the original: Baker et al. [2] and Woodward and Wood [31] lack venue and year; Cross [5] is a dated talk; the Iigaya et al. preprint [15] and CURL [27] have limited metadata. The Neumann–Morgenstern entry [30] combines an original publication year with a later commemorative-edition designation. These are preserved as historical source details awaiting bibliographic confirmation, not silently supplemented with invented metadata.

References

[1]

D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mane. Concrete Problems in AI Safety. arXiv:1606.06565, July 2016.

[2]

C. L. Baker, R. R. Saxe, and J. B. Tenenbaum. Bayesian Theory of Mind: Modeling Joint Belief-Desire Attribution. Original reference records six pages; venue and year unspecified.

[3]

C. J. Charpentier, K. Iigaya, and J. P. O’Doherty. A Neuro-computational Account of Arbitration between Choice Imitation and Goal Emulation during Human Observational Learning. Neuron, 106(4):687–699.e7, May 2020.

[4]

C. J. Charpentier and J. P. O’Doherty. The application of computational models to social neuroscience: promises and pitfalls. Social Neuroscience, 13(6):637–647, November 2018. doi:10.1080/17470919.2018.1518834.

[5]

L. Cross. Using Deep Reinforcement Learning to Uncover the Decision-Making Mechanisms. Talk, 25 October 2019.

[6]

R. M. Dawes. Social Dilemmas. Annual Review of Psychology, 31(1):169–193, 1980. doi:10.1146/annurev.ps.31.020180.001125.

[7]

M. Devaine, G. Hollard, and J. Daunizeau. The Social Bayesian Brain: Does Mentalizing Make a Difference When We Learn? PLOS Computational Biology, 10(12):e1003992, December 2014.

[8]

P. Doshi and P. J. Gmytrasiewicz. A Framework for Sequential Planning in Multi-Agent Settings. Journal of Artificial Intelligence Research, 24:49–79, July 2005. arXiv:1109.2135.

[9]

E. Fehr and S. Gachter. Cooperation and Punishment in Public Goods Experiments. American Economic Review, 90(4):980–994, September 2000.

[10]

A. N. Hampton, P. Bossaerts, and J. P. O’Doherty. Neural correlates of mentalizing-related computations during strategic interactions in humans. Proceedings of the National Academy of Sciences, 105(18):6741–6746, May 2008.

[11]

Y. Han and P. Gmytrasiewicz. IPOMDP-Net: A Deep Neural Network for Partially Observable Multi-Agent Planning Using Interactive POMDPs. Proceedings of the AAAI Conference on Artificial Intelligence, 33(1):6062–6069, July 2019.

[12]

G. Hardin. The Tragedy of the Commons. Journal of Natural Resources Policy Research, 1(3):243–253, July 2009. doi:10.1080/19390450903037302. Edition cited by the original proposal.

[13]

C. A. Hill, S. Suzuki, R. Polania, M. Moisa, J. P. O’Doherty, and C. C. Ruff. A causal account of the brain network computations underlying strategic social behavior. Nature Neuroscience, 20(8):1142–1149, August 2017.

[14]

E. Hughes, J. Z. Leibo, M. Phillips, K. Tuyls, E. Dueñez-Guzman, A. G. Castañeda, I. Dunning, T. Zhu, K. McKee, R. Koster, H. Roff, and T. Graepel. Inequity aversion improves cooperation in intertemporal social dilemmas. Advances in Neural Information Processing Systems, pp. 3330–3340, 2018.

[15]

K. Iigaya, S. Yi, I. Wahle, S. Tanwisuth, and J. P. O’Doherty. Aesthetic preference for art emerges from a weighted integration over hierarchically structured visual features in the brain. bioRxiv preprint. The author’s current name is used in this editorial source.

[16]

N. Jaques, A. Lazaridou, E. Hughes, C. Gulcehre, P. Ortega, D. Strouse, J. Z. Leibo, and N. de Freitas. Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning. International Conference on Machine Learning, pp. 3040–3049, 2019.

[17]

P. Kollock. Social Dilemmas: The Anatomy of Cooperation. Annual Review of Sociology, 24(1):183–214, 1998. doi:10.1146/annurev.soc.24.1.183.

[18]

S. W. Lee, S. Shimojo, and J. P. O’Doherty. Neural computations underlying arbitration between model-based and model-free learning. Neuron, 81(3):687–699, February 2014.

[19]

J. Z. Leibo, V. Zambaldi, M. Lanctot, J. Marecki, and T. Graepel. Multi-agent Reinforcement Learning in Sequential Social Dilemmas. Proceedings of AAMAS, pp. 464–473, May 2017.

[20]

M. W. Macy and A. Flache. Learning dynamics in social dilemmas. Proceedings of the National Academy of Sciences, 99(suppl. 3):7229–7236, May 2002.

[21]

Y. Niv. Learning task-state representations. Nature Neuroscience, 22(10):1544–1553, October 2019.

[22]

J. O’Doherty, S. Lee, R. Tadayonnejad, J. Cockburn, K. Iigaya, and C. J. Charpentier. Why and how the brain weights contributions from a mixture of experts. PsyArXiv preprint, June 2020.

[23]

J. P. O’Doherty, A. Hampton, and H. Kim. Model-based fMRI and its application to reward learning and decision making. Annals of the New York Academy of Sciences, 1104:35–53, May 2007.

[24]

M. Peterson. The value alignment problem: a geometric approach. Ethics and Information Technology, 21(1):19–28, March 2019.

[25]

A. Rapoport. Prisoner’s Dilemma—Recollections and Observations. In Game Theory as a Theory of a Conflict Resolution, pp. 17–34. Springer Netherlands, 1974.

[26]

T. Rusch, S. Steixner-Kumar, P. Doshi, M. Spezio, and J. Glascher. Theory of mind and decision science: Towards a typology of tasks and computational models. Neuropsychologia, 146:107488, September 2020.

[27]

A. Srinivas, M. Laskin, and P. Abbeel. CURL: Contrastive Unsupervised Representations for Reinforcement Learning. arXiv, 2020.

[28]

K. L. Stachenfeld, M. M. Botvinick, and S. J. Gershman. The hippocampus as a predictive map. Nature Neuroscience, 20(11):1643–1653, November 2017.

[29]

S. Suzuki and J. P. O’Doherty. Breaking human social decision making into multiple components and then putting them together again. Cortex, 127:221–230, June 2020.

[30]

J. von Neumann, O. Morgenstern, and A. Rubinstein. Theory of Games and Economic Behavior (60th Anniversary Commemorative Edition). Princeton University Press. Original proposal records 1944; edition/year correspondence requires confirmation.

[31]

M. P. Woodward and R. J. Wood. Learning from Humans as an I-POMDP. Original reference records five pages; venue and year unspecified.

[32]

D. L. K. Yamins, H. Hong, C. F. Cadieu, E. A. Solomon, D. Seibert, and J. J. DiCarlo. Performance-optimized hierarchical models predict neural responses in higher visual cortex. Proceedings of the National Academy of Sciences, 111(23):8619–8624, June 2014.

[33]

L. Zhu, K. E. Mathewson, and M. Hsu. Dissociable neural representations of reinforcement and belief prediction errors underlie strategic learning. Proceedings of the National Academy of Sciences, January 2012.