Revised working edition

Literature Review on Classifications of Multi-Agent Bandits and Multi-Agent Reinforcement Learning

Samyak Parajuli

Sandy Tanwisuth

Original review: May 2021
Editorial revision for author review: 9 October 2026

Contents

1 Introduction

What can multi-agent bandit theory and multi-agent reinforcement learning learn from one another? Both fields study agents that learn while interacting, but they often emphasize different objects: regret and information sharing in bandits, and sequential behavior and learned representations in reinforcement learning. This review places selected work from both fields within a common set of themes.

The review first describes the settings and objectives. It then compares approaches to emergent behavior, communication, cooperation and competition, and inference about other agents. The categories are not mutually exclusive. They organize a comparison of assumptions and mechanisms rather than a ranking of algorithms. The aim is to make connections between theoretical bandit models and practical deep reinforcement-learning methods visible enough to support further questions.

Scope and editorial status. The literature cutoff is the original May 2021 review. This revision preserves every substantive topic, algorithm, reported bound, and reference, while clarifying prose and separating model classes that the source occasionally conflates. Mathematical review notes flag incomplete assumptions, ambiguous formulas, and claims that require source-level checking. It is not an updated survey of all work through 2026 or a verification of every cited theorem.

2 Formalism

2.1 General settings and objectives

2.1.1 Single-agent bandits

A stochastic multi-armed bandit has an arm set 𝒜\mathcal A and a reward distribution F(a)F(a) for each arm [1]. In the bounded-reward setting considered here, rewards lie in [0,1][0,1]. Let μa\mu_a be the mean reward of arm aa and μ*=max⁡aμa\mu^*=\max_a\mu_a. The expected pseudo-regret of a policy choosing arm ata_t is 𝔼[∑t=1T(μ*−μat)].\mathbb{E}\left[\sum_{t=1}^T(\mu^*-\mu_{a_t})\right]. The expectation is needed when the selected arms are random. This makes explicit the expectation implicit in the source’s objective [2].

2.1.2 Multi-agent bandits

One cooperative model has a set DD of mm agents, joint arm space 𝒜=∏i∈D𝒜i\mathcal A=\prod_{i\in D}\mathcal A_i, and a stochastic global reward F(a)F(a) associated with the joint arm [3]. A communication graph G=(V,E)G=(V,E) describes which agents can exchange information: vertices represent agents and edges represent communication links.

A related distributed bandit model lets agents draw arms and receive samples from arm-specific reward distributions, then share information between rounds. The original review describes independent identically distributed arm rewards immediately after the joint-reward definition. These are different possible models. A shared joint reward does not imply independent rewards for individual agents; the cited algorithm’s observation and reward assumptions must determine which model is in use.

2.1.3 Single-agent reinforcement learning

A Markov decision process (MDP) represents sequential interaction through states 𝒮\mathcal S, actions 𝒜\mathcal A, a transition kernel T(s′∣s,a)T(s'\mid s,a), and a reward function R(s,a,s′)R(s,a,s'). The transition probabilities satisfy ∑s′T(s′∣s,a)=1.\sum_{s'}T(s'\mid s,a)=1. The agent seeks to maximize expected discounted return, 𝔼[∑t=0∞γtrt].\mathbb{E}\left[\sum_{t=0}^{\infty}\gamma^t r_t\right]. For bounded rewards and γ<1\gamma<1, this return is finite. The case γ=1\gamma=1 requires a finite horizon, termination, or another condition ensuring a well-defined objective.

The original tuple ⟨𝒮,O,𝒜,T,R⟩\langle\mathcal S,O,\mathcal A,T,R\rangle includes observations. In a fully observed MDP, the agent observes the state. If it receives only observations generated from states, an observation space and observation kernel are needed, producing a partially observable MDP (POMDP). This distinction matters because a policy may depend on the history of observations and actions, not just the current observation.

2.1.4 Multi-agent reinforcement learning

Markov games extend sequential decision-making to multiple agents [4]. A general partially observable, general-sum game has a state space, per-agent action and observation spaces, a transition kernel depending on the joint action, per-agent observation mechanisms, and reward functions RiR_i. Different rewards allow agents’ objectives to differ. A POMDP alone is a single-agent model; using it for a multi-agent problem requires specifying what is treated as part of the environment or what additional interactive structure is modeled [6].

Sequential social dilemmas (SSDs) are one family of multi-agent tasks [5]. They place social incentives in temporally extended interactions. Cooperation can then be a property of a policy or trajectory rather than a single action, and agents may have only partial information about one another [5], [6]. These tasks motivate partially observable Markov-game models, but do not exhaust MARL.

For NN agents, write the joint action as a=(a1,…,aN)∈∏i𝒜ia=(a_1,\ldots,a_N)\in\prod_i\mathcal A_i. A transition kernel has the form T:𝒮×𝒜×𝒮→[0,1],T:\mathcal S\times\mathcal A\times\mathcal S\longrightarrow[0,1], and agent ii receives Ri(s,a)R_i(s,a). This is a stochastic kernel; deterministic transitions are the special case concentrated on one successor state. The original text calls this kernel deterministic without imposing that restriction.

An observation function may be a deterministic map Oi:𝒮→ℝdiO_i:\mathcal S\to\mathbb{R}^{d_i} or a stochastic kernel. A reactive deterministic policy has form πi:Oi→𝒜i\pi_i:O_i\to\mathcal A_i, as in the source’s simplified description. More general policies can be stochastic and history-dependent. This flexibility is relevant when learners exhibit behavior that changes between strategies or combines them.

Independent learners often treat other agents’ changing behavior as nonstationarity in the environment [5], [7], [8]. This can avoid explicit recursive modeling, while obscuring information useful for predicting other agents [6]. It is a learning choice rather than a defining property of every SSD. Similarly, maximizing reward does not generally mean minimizing delay unless the task’s reward encodes delay.

3 Classifications of behavioral policies

Surveys of multi-agent learning offer several organizing frameworks [9], [10], [11], [12], [13], [14]. Shoham et al. [10] distinguish five agendas:

  1. Computational: use iterative procedures to compute properties of a game.

  2. Descriptive: model learning in a way that agrees with observed human behavior.

  3. Normative: study which repeated-game strategies are in equilibrium.

  4. Prescriptive cooperative: design decentralized control, joint policies, and resource allocation for a shared objective.

  5. Prescriptive noncooperative: determine how an individual agent should act to obtain high reward.

Stone [12] questions whether a game-theoretic framing covers the full range of multi-agent learning problems. Later surveys incorporate the challenges raised by deep learning, nonstationarity, and explicit models of other agents [13], [14], [15].

Following the framework used in the original review, the comparison below uses four overlapping themes: emergent behavior, communication, cooperation and competition, and inference about other agents. The review describes their implications and mechanisms without attempting a complete implementation-level comparison.

3.1 Emergent behavior

This category concerns behavior that develops through learning and interaction, often studied empirically with deep reinforcement learning. In the bandit papers selected for the original review, there is no direct counterpart to the rich sequential behavioral studies discussed here. That is a statement about this selection, not proof that bandit models cannot exhibit emergent collective behavior.

Deep Q-networks (DQN) [16] and proximal policy optimization (PPO) [17] serve as important underlying methods. Their single-agent forms help explain how multi-agent experiments are constructed.

3.1.1 Q-learning and deep Q-networks

An action-value function Q:𝒮×𝒜→ℝQ:\mathcal S\times\mathcal A\to\mathbb{R} estimates expected return, rather than simply the immediate scalar reward. Tabular Q-learning [18] updates the value of an experienced state-action pair using a sampled transition. Given learning rate α\alpha, discount γ\gamma, reward rr, and next state s′s', the standard update is Q(s,a)←(1−α)Q(s,a)+α[r+γmaxa′Q(s′,a′)].Q(s,a)\leftarrow(1-\alpha)Q(s,a)+\alpha\left[r+\gamma\max_{a'}Q(s',a')\right]. The algorithm initializes QQ, repeatedly chooses an action using a behavior policy, observes a reward and next state, applies this update, and advances to the next state until termination. A greedy action is arg⁡max⁡aQ(s,a)\arg\max_aQ(s,a), but learning guarantees generally require adequate exploration. The original pseudocode specifies only greediness and therefore does not supply such a guarantee.

Review note. The source updates Q(s′,a)Q(s',a) while the right-hand side is the temporal-difference update for Q(s,a)Q(s,a). The displayed correction repairs that index error explicitly. The original example learning rate α=0.1\alpha=0.1 is a tuning example, not a general convergence condition. DQN additionally uses neural function approximation and training mechanisms; tabular pseudocode alone does not define the full DQN algorithm.

3.1.2 Proximal policy optimization

PPO is an on-policy policy-gradient method [17]. It can be used with discrete or continuous action spaces, given an appropriate policy parameterization. The original review also cites the Gym environment framework [19] in this context.

Initialize policy parameters θ0\theta_0 and value parameters ϕ0\phi_0. At iteration kk, collect trajectories DkD_k under πθk\pi_{\theta_k}, compute rewards-to-go R̂t\widehat R_t, and estimate advantages Ât\widehat A_t using the current value function. With likelihood ratio rt(θ)=πθ(at∣st)πθk(at∣st),r_t(\theta)=\frac{\pi_\theta(a_t\mid s_t)}{\pi_{\theta_k}(a_t\mid s_t)}, the clipped objective averages min⁡{rt(θ)Ât,clip(rt(θ),1−ϵ,1+ϵ)Ât}\min\!\left\{r_t(\theta)\widehat A_t, \operatorname{clip}(r_t(\theta),1-\epsilon,1+\epsilon)\widehat A_t\right\} over the sampled transitions. The policy is updated approximately, typically by stochastic gradient ascent. The value function is fitted by minimizing an average squared error, 1|ℬk|∑(st,R̂t)∈ℬk(Vϕ(st)−R̂t)2,\frac{1}{|\mathcal B_k|}\sum_{(s_t,\widehat R_t)\in\mathcal B_k} \bigl(V_\phi(s_t)-\widehat R_t\bigr)^2, where ℬk\mathcal B_k denotes the collected transitions and targets.

The source expresses the second branch of the clipped objective as g(ϵ,A)g(\epsilon,A) without defining gg. The equivalent piecewise expression is g(ϵ,A)=(1+ϵ)Ag(\epsilon,A)=(1+\epsilon)A for A≥0A\geq0 and (1−ϵ)A(1-\epsilon)A for A<0A<0. Writing an empirical average also avoids the source’s inconsistent use of a T+1T+1-term sum with denominator TT. The original pseudocode’s final “return QQ” is replaced by the trained policy and value function; QQ is not an output of the displayed PPO routine.

3.1.3 Emergent behavior in MARL

The cited experiments build on DQN-based methods [5], [8], [20], [7] or PPO-based methods [21], [22]. In self-play, learning agents generate a changing training environment for one another [20], [21]. Cooperative and competitive patterns can emerge from those interactions. The appropriate interpretation depends on the task, incentives, and evaluation: a behavior that looks cooperative is not automatically evidence of a general cooperative objective.

3.2 Communication

How does sharing information change what agents can learn? Communication protocols may be local or global, discrete or continuous. Cooperative settings often use them to improve a shared objective, but the communication graph and message budget constrain the benefit.

3.2.1 Distributed bandits

Landgren et al. [23] study cooperative confidence-based methods whose uncertainty estimates account for information received from neighbors and its propagation through a network. A generic index has form Qik(t−1)=μ̂ik(t−1)+Cik(t−1).Q_i^k(t-1)=\widehat\mu_i^k(t-1)+C_i^k(t-1). The original review records the following confidence terms: coopUCB1:Cik(t−1)=σg2γG(η)n̂ik(t−1)+ϵckMn̂ik(t−1)log⁡(t−1)n̂ik(t−1),coopUCB2:Cik(t−1)=σg2γG(η)n̂ik(t−1)+f(t−1)Mn̂ik(t−1)log⁡(t−1)n̂ik(t−1).\begin{align} \text{coopUCB1:}\quad C_i^k(t-1) &=\sigma_g\sqrt{\frac{2\gamma}{G(\eta)} \frac{\widehat n_i^k(t-1)+\epsilon_c^k}{M\widehat n_i^k(t-1)} \frac{\log(t-1)}{\widehat n_i^k(t-1)}},\\ \text{coopUCB2:}\quad C_i^k(t-1) &=\sigma_g\sqrt{\frac{2\gamma}{G(\eta)} \frac{\widehat n_i^k(t-1)+f(t-1)}{M\widehat n_i^k(t-1)} \frac{\log(t-1)}{\widehat n_i^k(t-1)}}. \end{align} Here MM is the number of agents. The comparison in the source is that coopUCB2 replaces the graph-dependent ϵck\epsilon_c^k term with a function of time. These formulas reproduce the source’s intended grouping, with its missing closing parenthesis in G(η)G(\eta) repaired.

Review note. The source does not define G(η)G(\eta), σg\sigma_g, γ\gamma, n̂ik\widehat n_i^k, ϵck\epsilon_c^k, or ff sufficiently to implement or check the bounds. They must be matched to [23]; the displays are a record of the reviewed method, not self-contained confidence guarantees. Initialization is also needed where counts vanish or t=1t=1.

The Bayesian counterpart, coopUCL, uses an upper credible limit: Qik(t−1)=ν̂i(t−1)+σ̂i(t−1)Φ−1(1−α(t−1)).Q_i^k(t-1)=\widehat\nu_i(t-1)+\widehat\sigma_i(t-1) \Phi^{-1}\bigl(1-\alpha(t-1)\bigr). The source denotes posterior mean and standard deviation by ν̂i\widehat\nu_i and σ̂i\widehat\sigma_i. If Φ\Phi is the standard normal distribution function and the corresponding Gaussian posterior assumption holds, this is a 1−α(t−1)1-\alpha(t-1) posterior quantile. The original claim that it holds with probability α(t−1)\alpha(t-1) reverses the stated tail convention; a credible interval should also not be treated as a frequentist confidence guarantee without further argument.

3.2.2 Loose coupling and factored rewards

The joint action space grows exponentially with the number of agents when each has a fixed number of actions. Sparse interactions can reduce the burden when the global reward decomposes into local factors [24]: f(a)=∑e=1ρfe(ae),μ(a)=∑e=1ρμe(ae).f(a)=\sum_{e=1}^{\rho}f^e(a^e),\qquad \mu(a)=\sum_{e=1}^{\rho}\mu^e(a^e). Each factor depends on a subset of agents. A bipartite coordination graph connects agents to the reward factors they influence. The source assumes observable noisy local rewards and independence for its discussed factored model; this is an additional modeling condition, not a consequence of additive notation.

Overlapping factors can prefer conflicting actions for the same agent. Maximizing factors separately is therefore insufficient. Variable elimination can optimize the sum without explicitly enumerating every joint action, by eliminating variables while retaining their effects on neighboring factors. Its computational cost still depends on the induced factor structure.

3.2.2.1 Multi-Agent Upper Confidence Exploration (MAUCE).

MAUCE selects a joint action using estimated local rewards and exploration bonuses. The review relates it to combinatorial bandits [25], where a super-arm combines elementary arms with unknown reward distributions. It reports an improvement involving a harmonic mean of local upper-confidence quantities rather than their sum. The exact quantities and theorem conditions are not given in the source and remain to be checked before quoting a formal regret guarantee.

3.2.2.2 Multi-Agent Thompson Sampling (MATS).

At round tt, MATS samples each local mean μte(ae)\mu_t^e(a^e) from its posterior given the history of local actions and rewards. It then uses variable elimination to find the joint action maximizing the sampled total reward [24]. This repairs the source’s phrase “variable estimation.” The original review describes regret as sublinear in a factor involving ATAT, where AA denotes the number of local arms, but supplies no complete bound. That incomplete scaling description is preserved as a question for verification, not expanded into an invented theorem.

3.2.3 Communication in MARL

The cited MARL approaches often study cooperative partially observable games and optimize joint performance [26], [27], [22]. Parameter sharing trains agents using shared network parameters [27]; memory-based approaches provide a shared medium that agents can read and write [28]. Parameter sharing and an explicit execution-time communication channel are distinct mechanisms, even when an architecture uses both. The review’s classification highlights this design comparison rather than identifying the two.

3.3 Cooperation and competition

Explicit messages are useful in some cooperative problems but are not necessary for all cooperation. The reviewed settings range from fully cooperative to fully competitive and mixed-incentive interaction.

3.3.1 Fairness

If agents value arms differently, maximizing one scalar arm mean may not express the desired collective objective. Hossain et al. [29] study Nash social welfare. For NN agents, KK arms, mean utilities μi,j*\mu_{i,j}^*, and distribution pp, the objective is NSW⁡(p,μ*)=∏i=1N(∑j=1Kpjμi,j*).\operatorname{NSW}(p,\mu^*)= \prod_{i=1}^N\left(\sum_{j=1}^Kp_j\mu_{i,j}^*\right). An agent’s expected utility is the entire sum over arms, not a single term pjμi,jp_j\mu_{i,j}. Maximizing this product is one specified fairness criterion; it does not define fairness uniquely.

3.3.1.1 Explore-first.

The algorithm begins with exploration so that every arm is sampled LL times, then estimates a welfare-maximizing distribution p̂\widehat p and uses it during exploitation. The source compares the reported order N2/3K1/3T2/3(log⁡T)1/3N^{2/3}K^{1/3}T^{2/3}(\log T)^{1/3} with a single-agent explore-first scale lacking the N2/3N^{2/3} factor.

3.3.1.2 Epsilon-greedy.

At round tt, the algorithm explores with probability ϵt\epsilon_t, cycling through arms, and otherwise chooses the distribution maximizing estimated Nash social welfare. The source reports a horizon-independent schedule with scale ϵt=N2/3K1/3t−1/3(log⁡(NKt))1/3,\epsilon_t=N^{2/3}K^{1/3}t^{-1/3}\bigl(\log(NKt)\bigr)^{1/3}, and expected-regret scale N2/3K1/3T2/3(log⁡(NKt))1/3.N^{2/3}K^{1/3}T^{2/3}\bigl(\log(NKt)\bigr)^{1/3}.

Review note. A probability must be at most one, so the displayed schedule needs clipping or a restricted range. The regret display mixes the horizon TT with a free tt. Constants, admissible ranges, and the intended logarithm must be checked in [29].

3.3.1.3 UCB.

The welfare-based UCB approach optimizes an estimated Nash social welfare plus a confidence term involving a linear combination of agent-specific confidence intervals. The source states that choosing αt=N\alpha_t=N gives expected-regret scale NKTlog⁡(NKT).NK\sqrt T\log(NKT). This is a reported expression from the review; the confidence construction, assumptions, and exact logarithmic dependence are not reproduced there.

3.3.1.4 Heavy-tailed rewards.

Reward models are often bounded or sub-Gaussian, assumptions that support concentration. They may be unsuitable for some applications, including distributed load estimation and supply-chain models. The original discussion also mentions biased communication; bias and heavy tails are separate issues and should not be conflated.

Dubey and Pentland [30] study cooperative bandits with heavy-tailed rewards using robust mean estimators and a method called MP-UCB. The review describes controlling estimation error across a communication graph when ordinary consensus-based estimates are inadequate. It reports a group-regret lower-bound scale KΔ−1/ϵlog⁡T.K\Delta^{-1/\epsilon}\log T. The moment parameter ϵ\epsilon, the gap convention, and other assumptions are missing from the review and must be recovered from the cited work. The label “heavy-tailed” alone does not specify which moments exist or what rates can be obtained.

3.3.2 Cooperation and competition in MARL

SSDs provide a setting for studying cooperative and defecting policies over extended trajectories [5]. Experiments can show cooperation-like and competition-like patterns [5], [8], [7], [21], [22]. Some cited methods emphasize cooperation [8], [22], whereas others examine mixed interactions or heterogeneous preferences [5], [31], [7]. The task-level definition of cooperation should be kept separate from a claim that an agent possesses a general cooperative disposition.

3.4 Inference about other agents

Explicitly modeling others can support goal inference and account for changes in their learning behavior. It also raises a practical question: when should an agent trust another agent’s information?

3.4.1 Robustness to malicious agents

Vial et al. [32] consider a bandit network with nn honest and mm malicious agents. Motivating examples include faulty nodes in distributed systems and adversarial recommendations. Honest agents can adapt their communication based on the usefulness of past recommendations.

The review describes a blocking rule: when a recommended arm performs poorly at time tt, the recipient ignores that recommender until time t2t^2. Short blocks early in learning reduce the penalty for an honest error caused by noise; repeated poor recommendations later cause longer exclusions. The original review records an upper-bound scale (m+kn)log⁡(T/Δ),\left(m+\frac{k}{n}\right)\log(T/\Delta), where Δ\Delta is described as an arm gap. The exact gap dependence, communication schedule, and assumptions are not specified sufficiently to treat this display as a complete theorem. They remain a source-checking obligation.

3.4.2 Inference in MARL

The reviewed work includes social-influence objectives [33], inequity-aversion objectives [8], and methods involving joint-policy structure [31], [7], [22]. These overlap with the communication and cooperation themes. Some methods use interventions or counterfactual predictions to measure an agent’s influence; others shape preferences or policies without identifying another agent’s intention. The source’s broad claim that all such methods infer intentions through causality should therefore be narrowed to the actual mechanism of each method [8], [33], [31], [7], [21].

4 Discussion

The comparison organizes selected bandit and reinforcement-learning methods around four recurring questions: what behavior emerges, what information is communicated, which collective or individual objective is optimized, and how other agents are modeled. Many bandit methods address several of these questions at once. The more constrained settings often support explicit regret analysis, while the selected deep MARL studies explore richer sequential behavior whose theoretical guarantees are less fully characterized.

Those differences suggest directions for exchange rather than a simple division between theory and practice. Can a MARL problem be decomposed so that bandit-style guarantees become useful? Which communication and fairness assumptions remain appropriate when agents learn a policy over time? Which empirical coordination failures reveal distinctions that a simpler model should retain? The review leaves these as questions for future work.

Outstanding source checks

The review does not verify the reported constants or asymptotic expressions for coopUCB, MAUCE, MATS, welfare-based bandits, heavy-tailed bandits, or malicious-agent robustness. Definitions and probability assumptions must be aligned with each cited paper before those expressions are used as theorem statements. The earlier model distinctions, Q-learning index repair, PPO clipping notation, and posterior-tail convention are explicit editorial corrections. References with incomplete metadata are retained as supplied rather than filled with guesses.

References

[1]

W. R. Thompson. On the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples. Biometrika, 25(3–4):285–294, 1933.

[2]

S. Agrawal and N. Goyal. Analysis of Thompson Sampling for the Multi-Armed Bandit Problem. arXiv:1111.1797, 2011.

[3]

E. Bargiacchi, T. Verstraeten, D. Roijers, A. Nowe, and H. van Hasselt. Learning to Coordinate with Coordination Graphs in Repeated Single-Stage Multi-Agent Decision Problems. July 2018. Venue unspecified in the original reference.

[4]

M. L. Littman. Markov Games as a Framework for Multi-Agent Reinforcement Learning. Machine Learning Proceedings 1994, pp. 157–163, 1994.

[5]

J. Z. Leibo, V. Zambaldi, M. Lanctot, J. Marecki, and T. Graepel. Multi-agent Reinforcement Learning in Sequential Social Dilemmas. arXiv:1702.03037, 2017.

[6]

P. J. Gmytrasiewicz and P. Doshi. A Framework for Sequential Planning in Multi-Agent Settings. Journal of Artificial Intelligence Research, 24:49–79, 2005.

[7]

R. Koster, K. R. McKee, R. Everett, L. Weidinger, W. S. Isaac, E. Hughes, E. A. Dueñez-Guzman, T. Graepel, M. Botvinick, and J. Z. Leibo. Model-free Conventions in Multi-agent Reinforcement Learning with Heterogeneous Preferences. arXiv:2010.09054, version 1, 2020.

[8]

E. Hughes, J. Z. Leibo, M. G. Philips, K. Tuyls, E. A. Dueñez-Guzman, A. G. Castañeda, I. Dunning, T. Zhu, K. R. McKee, R. Koster, H. Roff, and T. Graepel. Inequity Aversion Resolves Intertemporal Social Dilemmas. arXiv:1803.08884, 2018. Title and author spellings follow the original review’s cited version.

[9]

L. Panait and S. Luke. Cooperative Multi-Agent Learning: The State of the Art. Autonomous Agents and Multi-Agent Systems, 11(3):387–434, 2005.

[10]

Y. Shoham, R. Powers, and T. Grenager. If Multi-agent Learning Is the Answer, What Is the Question? Artificial Intelligence, 171(7):365–377, 2007.

[11]

L. Busoniu, R. Babuska, and B. De Schutter. A Comprehensive Survey of Multiagent Reinforcement Learning. IEEE Transactions on Systems, Man, and Cybernetics, Part C, 38(2):156–172, 2008.

[12]

P. Stone. Multiagent Learning Is Not the Answer. It Is the Question. Artificial Intelligence, 171:402–405, 2007.

[13]

P. Hernandez-Leal, M. Kaisers, T. Baarslag, and E. M. de Cote. A Survey of Learning in Multiagent Environments: Dealing with Non-Stationarity. Original reference records 64 pages; publication metadata incomplete.

[14]

P. Hernandez-Leal, B. Kartal, and M. E. Taylor. A Survey and Critique of Multiagent Deep Reinforcement Learning. Autonomous Agents and Multi-Agent Systems, 33(6):750–797, 2019.

[15]

S. V. Albrecht and P. Stone. Autonomous Agents Modelling Other Agents: A Comprehensive Survey and Open Problems. arXiv:1709.08071, 2017.

[16]

V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al. Human-level Control through Deep Reinforcement Learning. Nature, 518(7540):529–533, 2015. The original reference abbreviates the remaining author list.

[17]

J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal Policy Optimization Algorithms. arXiv:1707.06347, 2017.

[18]

R. S. Sutton and A. G. Barto. Reinforcement Learning: An Introduction, second edition. MIT Press, 2018.

[19]

G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba. OpenAI Gym. 2016.

[20]

J. Z. Leibo, E. Hughes, M. Lanctot, and T. Graepel. Autocurricula and the Emergence of Innovation from Social Interaction: A Manifesto for Multi-Agent Intelligence Research. arXiv:1903.00742, 2019.

[21]

B. Baker, I. Kanitscheider, T. Markov, Y. Wu, G. Powell, B. McGrew, and I. Mordatch. Emergent Tool Use from Multi-Agent Autocurricula. arXiv:1909.07528, version cited February 2020.

[22]

B. Baker. Emergent Reciprocity and Team Formation from Randomized Uncertain Social Preferences. Original reference records 14 pages; publication metadata incomplete.

[23]

P. Landgren, V. Srivastava, and N. E. Leonard. On Distributed Cooperative Decision-making in Multiarmed Bandits. 2019.

[24]

T. Verstraeten, E. Bargiacchi, P. Libin, D. Roijers, J. Helsen, and A. Nowe. Multi-agent Thompson Sampling for Bandit Applications with Sparse Neighbourhood Structures. Scientific Reports, 10:6728, 2020.

[25]

W. Chen, Y. Wang, and Y. Yuan. Combinatorial Multi-armed Bandit: General Framework and Applications. Proceedings of ICML, PMLR 28(1):151–159, 2013.

[26]

S. Sukhbaatar, A. Szlam, and R. Fergus. Learning Multiagent Communication with Backpropagation. arXiv:1605.07736, 2016.

[27]

J. N. Foerster, Y. M. Assael, N. de Freitas, and S. Whiteson. Learning to Communicate with Deep Multi-agent Reinforcement Learning. Advances in Neural Information Processing Systems, pp. 2145–2153, 2016.

[28]

E. Pesce and G. Montana. Improving Coordination in Multi-agent Deep Reinforcement Learning through Memory-driven Communication. arXiv:1901.03887, 2019.

[29]

S. Hossain, E. Micha, and N. Shah. Fair Algorithms for Multi-agent Multi-armed Bandits. 2021.

[30]

A. Dubey and A. Pentland. Cooperative Multi-agent Bandits with Heavy Tails. 2020.

[31]

A. S. Vezhnevets, Y. Wu, R. Leblond, and J. Z. Leibo. Options as Responses: Grounding Behavioural Hierarchies in Multi-agent RL. arXiv:1906.01470, 2019.

[32]

D. Vial, S. Shakkottai, and R. Srikant. Robust Multi-agent Multi-armed Bandits. 2020.

[33]

N. Jaques, A. Lazaridou, E. Hughes, C. Gulcehre, P. A. Ortega, D. Strouse, J. Z. Leibo, and N. de Freitas. Intrinsic Social Motivation via Causal Influence in Multi-agent RL. arXiv:1810.08647, 2018.