Revised working edition
Well-Founded Arbitration for Coalitional Agency
Consolidated revision: October 9, 2026
Earlier drafts: July 23 and July 31, 2025
Abstract
When can an agent act on several internal judgments without concealing a disagreement that changes its response? We study a one-step arbitration model with fixed expert scores, a specified reference response, and uncertainty over admissible expert weights. The model separates two questions: whether an abstraction preserves the reference response, and whether alternative admissible weights lead to similar responses. We characterize the coarsest exact response-preserving partition and show that it is independent of positive temperature when the score map is fixed. A greedy construction preserves a checked stability constraint but need not find a largest stable set. For sampled checks, a coverage theorem gives an explicit conditional bound on the gap between a sampled maximum and its supremum. Corrected examples clarify the limits of scale normalization, rare-expert sampling, and hierarchical composition. The resulting gate is a criterion for response preservation and stability; connecting it to accuracy, common knowledge, or justified action requires additional models.
Contents
Working manuscript. Consolidates the two arbitration drafts with reviewed corrections.
1 When does internal agreement justify action?
An agent can receive several plausible recommendations without having a settled basis for choosing among them. A neural system may combine modules with conflicting outputs; a human may weigh competing judgments; a team may coordinate despite different objectives. These examples motivate a question about arbitration: when can an agent compress its internal predictions and choose an action without hiding a disagreement that matters to that choice?
The proposal studies this question separately from learning which predictions are correct. Classical decision and reinforcement-learning models supply useful tools for choosing actions under a specified objective [7], [8]. Here, the immediate problem is how to combine fixed inputs when their relative authority remains unsettled. Discussions of plural agency, hierarchical organization, and well-foundedness motivate this setting [6], [2], [9], [17]. Transformer components and systems trained from multiple feedback signals offer possible applications [3], [5]; applying the model to them requires a separate empirical account of what an expert score represents.
Strategic equivalence distinguishes co-player policies according to their best responses [10]. The present proposal transfers the organizing question to internal prediction profiles: which distinctions can be ignored without changing the selected response? It considers two checks. Sufficiency asks whether compression preserves that response. Stability asks whether admissible choices of expert weights produce similar responses. Abstention is the proposed output when either check fails. Selective prediction provides a related action/rejection perspective [14], [15], although no external error guarantee follows here without a model of outcomes.
1.1 The one-step pluralist bandit
Let be a finite, nonempty action set and let . An expert profile is Each is a fixed score. There is no feedback, reward observation, or iterative learning in the core setting. A map assigns profiles to abstraction classes. A reference score map induces Equation (1) The reference score map is fixed and specified as part of each instance. Equation equation (1) makes this assumption explicit without restricting the general results to one aggregation rule. The numerical example below uses . Results about the quotient hold for any fixed finite score map; what the abstraction preserves is the response induced by that map.
A Boltzmann response supplies a smooth relation between scores and action probabilities [21]. It does not, by itself, identify probabilities of expert correctness. Agreement about a response and accuracy about the world remain different properties [4], [13].
2 What should an abstraction preserve?
2.1 Canonical decision equivalence
Definition 1 (Canonical response equivalence). For fixed , write when . The canonical abstraction maps to its equivalence class .
Equality of distributions is reflexive, symmetric, and transitive, so this defines a partition. For an arbitrary abstraction , write and define the local discrepancy Use when a global guarantee is intended. A supremum avoids assuming that an arbitrary fiber is compact or that its maximum is attained.
Proposition 1 (Coarsest exact response-preserving partition). . An abstraction has if and only if every fiber of is contained in a fiber of .
Proof. Within a canonical fiber, the response distributions are equal, and their KL divergence is zero. Conversely, if two profiles share a -fiber and , their KL divergence is zero. Since finite-score softmax distributions have full support, equality follows. The profiles therefore share a canonical fiber. For sufficiency of the refinement condition, profiles in a common -fiber also share a canonical fiber, so every divergence in that fiber is zero and hence . ◻
This is uniqueness of a partition, up to renaming its labels. A safe abstraction may separate profiles that the canonical map merges; it may not merge profiles with different responses. The result resembles the construction of a minimal response-relevant statistic, but statistical sufficiency for an unknown data-generating parameter would require an additional model [23].
2.2 Temperature and computation
Proposition 2 (Score-shift characterization). For finite real score vectors and any , for some real . Consequently the exact response partition is independent of positive temperature when is fixed independently of temperature.
Proof. For necessity, equality of probability ratios for actions gives , hence . Fixing a reference action proves the constant-shift condition. For sufficiency, cancels from numerator and denominator of softmax. The resulting criterion does not depend on which finite positive temperature is used. ◻
Centered scores therefore provide an explicit invariant without enumerating equivalence classes. Temperature still changes approximate discrepancies. If the score map itself depends on temperature, the proposition does not imply that its equivalence classes remain fixed.
Tractable approximations remain relevant when scores, fibers, or worst-case discrepancies cannot be computed cheaply. It should be distinguished from an assertion that evaluation of the exact score invariant is inherently intractable.
2.3 Approximate compression
A proposed information-constrained objective is and an empirical penalty of the form The constrained specification and learned penalty serve different roles [22]. Their implementation must define a distribution on profiles, the abstraction family, estimators, and the relationship between sampled and worst-case discrepancies. Deterministic continuous outputs can have infinite mutual information. A finite penalty does not automatically enforce a hard constraint, and a validation-sample minimum cannot certify global feasibility. These are missing modeling and optimization obligations, not resolved guarantees.
3 How should uncertainty about expert weights be handled?
For , define Given an independently specified admissible set , the stability discrepancy is This checks variation across weights at a fixed profile, whereas sufficiency checks profiles within an abstraction class. Distributed aggregation and social choice motivate the distinction [1], [24], [25]; they do not establish agreement for arbitrary expert scores.
3.1 Three proposed weight families
Three candidate specifications encode different sources of uncertainty: Competence scores may encode historical validation accuracy, calibration, or domain evidence. The inverse-uncertainty family requires and a definition of what varies. Variance across an expert’s action scores is not the same as uncertainty about those scores. A floor or other convention for zero values changes the family and must be specified. Their union is one proposed way to represent several sources of trust uncertainty. These sets are generally continuous; finite implementations require grids or sampling with stated error.
3.2 Adaptive construction and its limitation
A greedy construction starts from the uniform weight and an ordered sequence of candidate batches . For the current profile , write Set . At stage , let
Proposition 3 (Conditional stability of greedy acceptance). For , every set accepted by this construction has discrepancy below , provided each supremum comparison is exact.
Proof. The singleton has discrepancy zero. At every later step, either the set is unchanged or its replacement is accepted only after the required bound has been checked. Induction proves the claim. ◻
The same conclusion holds with a sound upper-bound certificate in place of the exact supremum. Passing a sampled maximum alone is not such a certificate. A family selected from also has a different interpretation from a trust family fixed independently of .
The procedure need not find a largest stable set. Consider binary action probabilities and threshold on both directed divergences. The first is compatible with each other probability. The response conflicts with and , which are compatible with each other. Accepting first yields a set of two responses, while another stable set contains three. These probabilities are realizable using weighted scores of experts and at temperature one. Maximum-cardinality selection is a separate optimization problem, and the candidate ordering is part of the greedy method.
4 What do the two checks establish?
The proposed operational gate permits action when and abstains otherwise. Threshold convention matters at equality and should remain consistent throughout an implementation. The criterion expresses response preservation and robustness relative to the selected map, weights, temperature, and tolerance. It does not establish the correctness of an action, the acceptability of its consequences, or a preference among conflicting objectives.
4.1 A coordination interpretation
One coordination interpretation treats a canonical abstraction as a shared norm, its preservation as norm coherence, and stability across admissible weights as norm consensus. Convention emergence and learned norms provide relevant context [31], [26], [27], [28], [29], [30]. The interpretation uses three conditions:
a canonical abstraction is well-defined;
the abstraction preserves relevant distinctions;
admissible interpretations induce sufficiently similar action distributions.
If these terms are defined to be the two numerical checks and existence of their underlying map, equivalence to the operational gate is definitional. A theorem about common knowledge would instead require agents, possible worlds, accessibility relations, and knowledge operators [11], [12], [19], [20]. A comparison with epistemic responsibility or trust can be useful without identifying numerical agreement with such a theorem [16], [13]. The contribution here is an operational definition with a coordination interpretation. A proposed framework equivalence leaves a question for future formalization: under which epistemic model, if any, do the two numerical checks correspond to a common-knowledge property? No such equivalence is proved by the present definitions.
The same distinction limits proposed applications to neural modules, human–AI teams, and coalitions. Memory abstractions describe what information is represented [18]; the gate asks whether a specified response survives selected internal distinctions. Neither alone supplies a full theory of justified action.
5 A diagnostic example, with corrected arithmetic
Take three actions and three experts. In profile , all experts report . In , only the second expert changes, reporting . The means are and , respectively. At , For the second profile, trusting only expert one yields , while trusting only expert two gives The directed divergences are and . Neither exceeds .
The two profiles do not share the same exact-mean abstraction, so their comparison does not demonstrate a sufficiency failure. Retaining the mean is exactly sufficient when the reference response is softmax of that mean. This is a sensitivity calculation; these comparisons alone give no reason to abstain at threshold . They also do not certify every possible profile or weight pair.
6 Can the gate be evaluated efficiently?
An empirical maximum over sampled profile or weight pairs is a lower bound on the corresponding exact supremum. A sampled violation can reject the strict gate, but a passing sample cannot certify the supremum without further assumptions. The following result states one sufficient coverage condition.
Proposition 4 (Finite-cover sampling certificate). Let be a nonempty metric space and let be -Lipschitz. Suppose balls , , with centers in cover . Let a proposal law satisfy for every . For independent samples , define and . Then, with probability at least , In particular, on this event, certifies .
Proof. The probability of missing any particular covering ball is at most . A union bound shows that all balls contain a sample except with probability at most . On that event, for any , choose a covering center within of and a sample within of that center. The sample is at distance at most from , so Lipschitz continuity gives . Taking a supremum gives the upper bound; the lower bound follows because every sample belongs to . ◻
For failure probability at most , it is sufficient that . This is an unconditional high-probability certificate over the declared sampling procedure, not a posterior probability conditional on accepting. Data-dependent choice of a fiber, reuse of samples to learn the abstraction, or repeated testing needs an appropriate joint guarantee or independent evaluation.
To apply Proposition 4, the domain is a specified set of profile pairs for sufficiency, or weight pairs for stability; is the directed response KL on that domain. The cover, metric, Lipschitz constant, and proposal mass must be justified for that particular domain. A Lipschitz bound on softmax alone does not supply these ingredients. Fixed finite scores and compact weight sets can make smoothness arguments available, but arbitrary abstraction fibers or an arbitrary reference map need not meet a proposed uniform sampling bound.
6.1 A direct score-range certificate
Proposition 5 (A finite bound for softmax disagreement). For and with finite scores, put . Then Consequently weighted averages of scores in have pairwise KL at most . If for every , every pair of such weighted responses has KL at most .
Proof. The softmax normalizers give . The first term is at most and the second is at least . For bounded weighted scores, each coordinate difference lies in . Under the stated near-consensus condition, any two weighted averages at a fixed action differ by at most , giving . ◻
This bound is conservative but explicit. For example, certifies the stability check under near-consensus. It does not certify sufficiency across different profiles, and it does not automatically bound an unrestricted reference score map.
6.2 Cost and implementation
Computing one fresh dense weighted score costs ; softmax and KL cost after scores are available. Evaluating new dense aggregations therefore generally costs . Lower cost requires a specified cache, sparse weights, or a different sample representation. Adaptive sampling, cached evaluations, parallel comparisons, and early rejection are implementation options with their own error and cost accounting.
A singleton abstraction fiber has zero sufficiency discrepancy. Lower temperature does not universally imply more rejection: different fixed score vectors sharing a unique maximizer can both concentrate on that same action as temperature tends to zero. Finite-temperature tolerances still require a quantitative calculation.
7 Extensions and unresolved questions
A possible feedback-based trust update is The signal is deliberately unspecified in this proposal. If denotes regret, the displayed update increases weights on experts assigning larger scores to the regretted action. A feedback model must determine whether this sign is appropriate before the update is used as a learning rule. Sequential settings also need transition, feedback, and abstention-cost models. Strategic experts introduce a further question: what prevents advantageous misreporting, and can truthful elicitation coexist with the gate?
The immediate research question is narrower than universal agency: under which specified prediction, weighting, and sampling models can response preservation and stability be certified at a useful cost? Establishing that result would clarify which broader coordination interpretations the model can support.
Acknowledgments and provenance
I thank Richard Ngo for advice and support, and Manfred Diaz, Su Hyeong Lee, and Jeffrey Heninger for conversations. The earlier drafts described this work as independently authored and unaffiliated. The original dedication to Meridian, Cassiopeia, and Chris is retained: thank you for showing that the deepest care often needs no words. This manuscript consolidates the July 23 and July 31, 2025 drafts. Their originals and separate reconstructions are retained as historical versions. Arithmetic corrections, counterexamples, and the distinction between established results and proposed extensions are recorded in the accompanying mathematical review. No new experiment is reported here.
8 Scale, sampling, and hierarchical composition
Can the same arbitration interface be reused across atomic components, subsystems, and larger coalitions? The July 31 extension considered action-set size, expert count, hierarchical depth, and representation dimension. Those questions remain useful, but each change requires a defined transformation and a property to preserve.
8.1 Parameter normalization
Candidate normalizations are with seed constants , and a trust radius . These are candidate parameterizations, not a demonstrated invariance theorem. They require or a separately defined singleton-action convention.
A proposed justification for these normalizations assumed . This is false: for scores and at temperature , the divergence is about , greater than . Changing temperature also does not multiply every KL divergence by the same factor. Exact canonical partitions remain unchanged at positive temperature for the score-shift reason proved earlier; approximate gate decisions require another argument. Renaming actions, duplicating experts, adding new actions, and nesting gates are different operations. Proposition 5 gives a valid score/temperature bound, but does not establish invariance under changing the action set. The normalizations remain quantities to evaluate.
8.2 Paired expert sampling
The proposed algorithm samples distinct expert indices, compares their response distributions, rejects when a discrepancy exceeds the scaled threshold, and otherwise accepts after comparisons. Its runtime is given random access to expert predictions. This accounting does not include constructing those predictions or certifying arbitrary weighted aggregations.
A sampled maximum is not an unbiased estimator of the maximum over all pairs. With one exceptional expert among , a uniform pair includes it with probability . Independent sampling misses it with probability . For and , this probability is about . A uniform worst-case detection guarantee cannot therefore use that fixed sample budget independently of . If a declared pair-sampling law gives violating pairs mass at least , independent comparisons miss them with probability at most . This distributional guarantee depends on its stated lower-mass condition. It does not certify every arbitrary weighted aggregation.
8.3 Recursive interfaces and composition
A gate at level accepts expert scores, local parameters, and an abstraction; it returns a policy or abstains. One proposal feeds a child policy into the next level as an expert score and sets . The interface is a possible design, but the proposed scalar reduction does not follow from softmax: at unit temperatures, A hierarchy therefore needs an explicit composition rule, a meaning for child abstention, and an error-propagation analysis. Summing local costs gives a computational total only after the number of gates, input construction, and local sample budgets are specified.
8.4 What can be checked without an invariance claim?
Consistent permutation of actions preserves softmax probabilities up to the same permutation and leaves KL unchanged. A gate built entirely from such equivariant operations can preserve its accept/abstain decision under relabeling. This observation does not imply invariance when the number of actions changes.
A verification plan includes expert-count comparisons, action-set changes, hierarchical wrapping, and shared seed parameters. These remain useful implementation experiments if outputs, costs, and failure frequencies are recorded. Equality of configuration fields or interface types is an interface test, not evidence that all epistemic guarantees survive a transformation. Which measurable invariance should the hierarchical framework target first?
References
- [1]
-
Kenneth J. Arrow. Social Choice and Individual Values. Wiley, 1951.
- [2]
-
Richard Ngo. Towards a scale-free theory of agency. AI Alignment Forum, 2025. https://www.alignmentforum.org/posts/gHefoxiznGfsbiAu9/towards-a-scale-free-theory-of-agency.
- [3]
-
Nelson Elhage et al. A mathematical framework for transformer circuits. Transformer Circuits Thread, 2021.
- [4]
-
Stuart Russell. Human Compatible. Viking, 2019.
- [5]
-
Yuntao Bai et al. Constitutional AI: Harmlessness from AI feedback. arXiv:2212.08073, 2022.
- [6]
-
Christian List and Philip Pettit. Group Agency. Oxford University Press, 2011.
- [7]
-
Leonard J. Savage. The Foundations of Statistics. Wiley, 1954.
- [8]
-
Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction, second edition. MIT Press, 2018.
- [9]
-
Richard Ngo. Well-foundedness as an organizing principle of healthy minds and societies. Mind the Future, 2025.
- [10]
-
Niklas Lauffer, Ameesh Shah, Micah Carroll, Michael Dennis, and Stuart Russell. Who Needs to Know? Minimal Knowledge for Optimal Coordination. 2023. https://arxiv.org/abs/2306.09309.
- [11]
-
Ronald Fagin, Joseph Y. Halpern, Yoram Moses, and Moshe Y. Vardi. Reasoning About Knowledge. MIT Press, 1995.
- [12]
-
Brian F. Chellas. Modal Logic: An Introduction. Cambridge University Press, 1980.
- [13]
-
Alvin I. Goldman. Knowledge in a Social World. Oxford University Press, 1999.
- [14]
-
C. K. Chow. On optimum recognition error and reject tradeoff. IEEE Transactions on Information Theory, 16(1):41–46, 1970.
- [15]
-
Yonatan Geifman and Ran El-Yaniv. Selective classification for deep neural networks. NeurIPS, 2017.
- [16]
-
Richard Ngo. Trust develops gradually via making bids and setting boundaries. LessWrong, 2023.
- [17]
-
Jan Kulveit. Hierarchical agency: A missing piece in AI alignment. AI Alignment Forum, 2024.
- [18]
-
Aaron Kirtland, Alexander Ivanov, Cameron Allen, Michael L. Littman, and George Konidaris. Memory as state abstraction over trajectories. 2025.
- [19]
-
Robert J. Aumann. Agreeing to disagree. The Annals of Statistics, 4(6):1236–1239, 1976.
- [20]
-
Yoram Moses and Moshe Tennenholtz. Distributed epistemic algorithms: Knowledge, coordination and common knowledge. TARK, 1995.
- [21]
-
Edwin T. Jaynes. Probability Theory: The Logic of Science. Cambridge University Press, 2003.
- [22]
-
Thomas M. Cover and Joy A. Thomas. Elements of Information Theory. Wiley, 2006.
- [23]
-
Erich L. Lehmann and George Casella. Theory of Point Estimation. Springer, 2005.
- [24]
-
Morris H. DeGroot. Reaching a consensus. Journal of the American Statistical Association, 69(345):118–121, 1974.
- [25]
-
Michel Grabisch, Jean-Luc Marichal, Radko Mesiar, and Endre Pap. Aggregation Functions. Cambridge University Press, 2009.
- [26]
-
Abram Demski. Learning normativity: A research agenda. AI Alignment Forum, 2020.
- [27]
-
Abram Demski. Four motivations for learning normativity. AI Alignment Forum, 2021.
- [28]
-
Ninell Oldenburg and Tan Zhi-Xuan. Learning and sustaining shared normative systems via Bayesian rule induction in Markov games. AAMAS, 2024.
- [29]
-
Zhi-Xuan Tan and Desmond C. Ong. Bayesian inference of social norms as shared constraints on behavior. arXiv:1905.11110, 2019.
- [30]
-
Zhi-Xuan Tan, Jake Brawer, and Brian Scassellati. That’s mine! Learning ownership relations and norms for robots. arXiv:1812.02576, 2019.
- [31]
-
David K. Lewis. Convention: A Philosophical Study. Harvard University Press, 1969.