Multiagent learning & coordination

Sandy Tanwisuth

Independent researcher in theoretical AI and pluralistic alignment

Sandy smiling on a beach at sunset, pointing toward a goat.

What do agents need to know about one another to coordinate, preserve their independence, and help each other flourish?

I study multiagent learning and strategic reasoning: which distinctions between agents matter for a decision, how those distinctions can be learned through interaction, and when uncertainty should lead an agent to seek more information or abstain. I am particularly interested in how agents whose goals do not fully align can coordinate in ways that expand each participant’s agency and capacity to flourish.

About

I’m an independent researcher supported by a BlueDot Impact Career Transition Grant. I also facilitate technical AI safety courses and mentor projects with BlueDot Impact. At SPAR, I mentor Representation Diagnostics for LLM Safety, investigating whether internal model signals distinguish refusal, clarification, and other safety behaviors across changing contexts. I write and teach Multi-Agent Learning at ILIAD.

I was a research scholar in the MATS program and a research intern at the Center for Human-Compatible AI at UC Berkeley. My earlier research examined human learning and decision-making at UC Berkeley and Caltech.

My curiosity centers on the boundary of trust. What do we entrust to one another, and how should that change as we learn? Interaction is a two-way process: people and AI systems shape one another’s understanding, choices, and possibilities. I want to understand how these relationships can expand what each participant is able to do, while preserving the freedom to question, disagree, and renegotiate.

Acting sensibly in isolation can still leave everyone with an outcome they would prefer to avoid. What changes when agents can make commitments, learn shared conventions, or revise how they respond to one another? I study when these possibilities produce mutual gains, whether those gains survive mistakes and changing incentives, and where trust becomes dependence or vulnerability. The aim is coordination that supports empowerment and flourishing: expanding what participants can achieve together while leaving each room to grow and shape what comes next.

Research questions

Which differences in another agent’s behavior change how we should respond?

I investigate representations that retain information relevant to incentives and decisions. How can we recognize when a useful approximation has discarded a distinction that matters?

When does an agent know enough to act, and when should it defer?

I study uncertainty and arbitration among conflicting perspectives. This work asks how an abstraction can preserve both a proposed action and the reasons for confidence in it.

How can agents coordinate when their partners and objectives keep changing?

I explore how interaction can make agents’ strategic effects easier to interpret, with open questions about evaluation, stability, and preserving individual autonomy.

Selected publications and reports

Google Scholar ↗

My recent work examines strategic abstraction, decisions under uncertainty, and the representations that guide AI behavior. Earlier collaborative work studied human learning, habits, and aesthetic value.

2026

Sisyphus in the loop: What Makes an LLM Persist?

Digital Minds Research Sprint, August 2026. Sprint research report.

Mohan W. Gupta, Xingyu Shirley Liu, and Sandy Tanwisuth.

What makes an LLM continue a task or stop? Using Qwen3.5-4B in a bandit task, we combine incentive changes, linear probes, and activation steering. Future return was decodable, but steering the identified return direction did not measurably change persistence; steering a persistence-aligned direction did. The study separates represented information from its tested influence on action.

2025

Uncertainty-Aware Policy-Preserving Abstractions with Abstention for One-Shot Decisions

Second Workshop on Aligning Reinforcement Learning Experimentalists and Theorists (ARLET), NeurIPS 2025. Workshop position paper.

Sandy Tanwisuth and Daniel K. Leja.

When should a simplified decision model retain enough uncertainty to justify abstention? We propose representing the margin between competing actions alongside the preferred action, and identify computational, statistical, and epistemic questions for deciding when to defer.

2023

Neural mechanisms underlying the hierarchical construction of perceived aesthetic value

Nature Communications, 14, article 127. Published 24 January 2023.

Kiyohito Iigaya, Sanghyun Yi, Iman A. Wahle, Sandy Tanwisuth, Logan Cross, and John P. O'Doherty.

How does the brain turn the features of an image into an aesthetic judgment? Computational modeling and functional neuroimaging provide evidence for a hierarchy that combines visual features into perceived value.

2022

Reinforcement learning with associative or discriminative generalization across states and actions: fMRI at 3 T and 7 T

Human Brain Mapping, 43(15), 4750–4790. First published 21 July 2022. Scott T. Grafton and John P. O'Doherty are co-senior authors.

Jaron T. Colas et al. (with Sandy Tanwisuth).

All authors and contributions

Jaron T. Colas, Neil M. Dundon, Raphael T. Gerraty, Natalie M. Saragosa-Harris, Karol P. Szymula, Sandy Tanwisuth, J. Michael Tyszka, Camilla van Geen, Harang Ju, Arthur W. Toga, Joshua I. Gold, Dani S. Bassett, Catherine A. Hartley, Daphna Shohamy, Scott T. Grafton, and John P. O'Doherty.

How do people use feedback about one event to learn about another? This study examines generalization across states and actions through computational modeling of behavior and fMRI during a structured reversal-learning task.

2022

Determining the effects of training duration on the behavioral expression of habitual control in humans: a multilaboratory investigation

Learning & Memory, 29(1), 16–28. The first eight authors contributed equally.

Eva R. Pool et al. (with Sandy Tanwisuth).

All authors and contributions

Eva R. Pool, Rani Gera, Aniek Fransen, Omar D. Perez, Anna Cremer, Mladena Aleksic, Sandy Tanwisuth, Stephanie Quail, Ahmet O. Ceceli, Dylan A. Manfredi, Gideon Nave, Elizabeth Tricomi, Bernard Balleine, Tom Schonberg, Lars Schwabe, and John P. O'Doherty.

Does longer training make human behavior more habitual? This collaboration across laboratories tested how training duration affects sensitivity to outcome devaluation. It found no statistical evidence for a main effect of training duration on habit expression in the studied task.

Papers and writing

Other work on Google Scholar

Working manuscripts

Current drafts on strategic representation, arbitration, and learning through interaction.

  • Strategic Representations for Multiagent CoordinationSandy Tanwisuth · Revised working manuscript · 2026

    Which differences between partners must a representation preserve to support a useful response? This manuscript develops full-policy results under explicit assumptions and connects response variation, successor features, and information gained through interaction.

  • Well-Founded Arbitration for Coalitional AgencySandy Tanwisuth · Revised working manuscript · 2026

    When do an agent’s internal perspectives agree enough to justify a decision? This manuscript studies response-preserving abstractions and action or abstention criteria, with explicit limits on sampling certificates and epistemic interpretations.

  • Strategic OpenendednessSandy Tanwisuth · Revised working manuscript · 2026

    Which interactions teach an agent something useful about a partner? An explicit Bayesian model defines information gain about response-relevant partner types; extending it to changing populations and testing the proposed learning method remain open.

Notes and earlier work

CV and contact

I welcome questions about my work and conversations about research collaboration.

My CV includes research, teaching, service, and earlier work in cognitive neuroscience.

Download CV (PDF)

The contact form is not available yet. Please check back soon.

Funders, Affiliations, and Communities

I am grateful to the organizations and communities that have supported my research, learning, and teaching, and to the people who have shared their time and ideas.