Revised working edition

Well-Founded Arbitration for Coalitional Agency

Sandy Tanwisuth

Consolidated revision: October 9, 2026
Earlier drafts: July 23 and July 31, 2025

Abstract

When can an agent act on several internal judgments without concealing a disagreement that changes its response? We study a one-step arbitration model with fixed expert scores, a specified reference response, and uncertainty over admissible expert weights. The model separates two questions: whether an abstraction preserves the reference response, and whether alternative admissible weights lead to similar responses. We characterize the coarsest exact response-preserving partition and show that it is independent of positive temperature when the score map is fixed. A greedy construction preserves a checked stability constraint but need not find a largest stable set. For sampled checks, a coverage theorem gives an explicit conditional bound on the gap between a sampled maximum and its supremum. Corrected examples clarify the limits of scale normalization, rare-expert sampling, and hierarchical composition. The resulting gate is a criterion for response preservation and stability; connecting it to accuracy, common knowledge, or justified action requires additional models.

Contents

Working manuscript. Consolidates the two arbitration drafts with reviewed corrections.

1 When does internal agreement justify action?

An agent can receive several plausible recommendations without having a settled basis for choosing among them. A neural system may combine modules with conflicting outputs; a human may weigh competing judgments; a team may coordinate despite different objectives. These examples motivate a question about arbitration: when can an agent compress its internal predictions and choose an action without hiding a disagreement that matters to that choice?

The proposal studies this question separately from learning which predictions are correct. Classical decision and reinforcement-learning models supply useful tools for choosing actions under a specified objective [7], [8]. Here, the immediate problem is how to combine fixed inputs when their relative authority remains unsettled. Discussions of plural agency, hierarchical organization, and well-foundedness motivate this setting [6], [2], [9], [17]. Transformer components and systems trained from multiple feedback signals offer possible applications [3], [5]; applying the model to them requires a separate empirical account of what an expert score represents.

Strategic equivalence distinguishes co-player policies according to their best responses [10]. The present proposal transfers the organizing question to internal prediction profiles: which distinctions can be ignored without changing the selected response? It considers two checks. Sufficiency asks whether compression preserves that response. Stability asks whether admissible choices of expert weights produce similar responses. Abstention is the proposed output when either check fails. Selective prediction provides a related action/rejection perspective [14], [15], although no external error guarantee follows here without a model of outcomes.

1.1 The one-step pluralist bandit

Let AA be a finite, nonempty action set and let K≥1K\geq1. An expert profile is Q=(Q1,…,QK)∈[0,1]K×|A|.Q=(Q_1,\ldots,Q_K)\in[0,1]^{K\times |A|}. Each Qk(a)Q_k(a) is a fixed score. There is no feedback, reward observation, or iterative learning in the core setting. A map ϕ\phi assigns profiles to abstraction classes. A reference score map x(Q)∈ℝ|A|x(Q)\in\mathbb{R}^{|A|} induces Equation (1)πraw,τ(a∣Q)=exp⁡(xa(Q)/τ)∑b∈Aexp⁡(xb(Q)/τ),τ>0.\begin{equation} \label{eq:raw} \pi_{\mathrm{raw},\tau}(a\mid Q)= \frac{\exp(x_a(Q)/\tau)}{\sum_{b\in A}\exp(x_b(Q)/\tau)},\qquad \tau>0. \end{equation} The reference score map is fixed and specified as part of each instance. Equation equation (1) makes this assumption explicit without restricting the general results to one aggregation rule. The numerical example below uses xa(Q)=K−1∑kQk(a)x_a(Q)=K^{-1}\sum_kQ_k(a). Results about the quotient hold for any fixed finite score map; what the abstraction preserves is the response induced by that map.

A Boltzmann response supplies a smooth relation between scores and action probabilities [21]. It does not, by itself, identify probabilities of expert correctness. Agreement about a response and accuracy about the world remain different properties [4], [13].

2 What should an abstraction preserve?

2.1 Canonical decision equivalence

Definition 1 (Canonical response equivalence). For fixed τ>0\tau>0, write Q∼τQ′Q\sim_\tau Q' when πraw,τ(⋅∣Q)=πraw,τ(⋅∣Q′)\pi_{\mathrm{raw},\tau}(\cdot\mid Q)=\pi_{\mathrm{raw},\tau}(\cdot\mid Q'). The canonical abstraction maps QQ to its equivalence class [Q]∼τ[Q]_{\sim_\tau}.

Equality of distributions is reflexive, symmetric, and transitive, so this defines a partition. For an arbitrary abstraction ϕ\phi, write Fϕ(q)={Q:ϕ(Q)=ϕ(q)}F_\phi(q)=\{Q:\phi(Q)=\phi(q)\} and define the local discrepancy dϕ(q)=supQ,Q′∈Fϕ(q)DKL(πraw,τ(⋅∣Q)‖πraw,τ(⋅∣Q′)).d_\phi(q)=\sup_{Q,Q'\in F_\phi(q)} D_{\mathrm{KL}}\bigl(\pi_{\mathrm{raw},\tau}(\cdot\mid Q)\Vert \pi_{\mathrm{raw},\tau}(\cdot\mid Q')\bigr). Use Dϕ=sup⁡qdϕ(q)D_\phi=\sup_q d_\phi(q) when a global guarantee is intended. A supremum avoids assuming that an arbitrary fiber is compact or that its maximum is attained.

Proposition 1 (Coarsest exact response-preserving partition). Dϕcan=0D_{\phi_{\mathrm{can}}}=0. An abstraction ϕ\phi has Dϕ=0D_\phi=0 if and only if every fiber of ϕ\phi is contained in a fiber of ϕcan\phi_{\mathrm{can}}.

Proof. Within a canonical fiber, the response distributions are equal, and their KL divergence is zero. Conversely, if two profiles share a ϕ\phi-fiber and Dϕ=0D_\phi=0, their KL divergence is zero. Since finite-score softmax distributions have full support, equality follows. The profiles therefore share a canonical fiber. For sufficiency of the refinement condition, profiles in a common ϕ\phi-fiber also share a canonical fiber, so every divergence in that fiber is zero and hence Dϕ=0D_\phi=0. ◻

This is uniqueness of a partition, up to renaming its labels. A safe abstraction may separate profiles that the canonical map merges; it may not merge profiles with different responses. The result resembles the construction of a minimal response-relevant statistic, but statistical sufficiency for an unknown data-generating parameter would require an additional model [23].

2.2 Temperature and computation

Proposition 2 (Score-shift characterization). For finite real score vectors x,yx,y and any τ>0\tau>0, softmax⁡(x/τ)=softmax⁡(y/τ)⇔x−y=c𝟏\operatorname{softmax}(x/\tau)=\operatorname{softmax}(y/\tau) \quad\Longleftrightarrow\quad x-y=c\mathbf1 for some real cc. Consequently the exact response partition is independent of positive temperature when x(Q)x(Q) is fixed independently of temperature.

Proof. For necessity, equality of probability ratios for actions a,ba,b gives e(xa−xb)/τ=e(ya−yb)/τe^{(x_a-x_b)/\tau}=e^{(y_a-y_b)/\tau}, hence xa−ya=xb−ybx_a-y_a=x_b-y_b. Fixing a reference action proves the constant-shift condition. For sufficiency, ec/τe^{c/\tau} cancels from numerator and denominator of softmax. The resulting criterion does not depend on which finite positive temperature is used. ◻

Centered scores x(Q)−xa0(Q)𝟏x(Q)-x_{a_0}(Q)\mathbf1 therefore provide an explicit invariant without enumerating equivalence classes. Temperature still changes approximate discrepancies. If the score map itself depends on temperature, the proposition does not imply that its equivalence classes remain fixed.

Tractable approximations remain relevant when scores, fibers, or worst-case discrepancies cannot be computed cheaply. It should be distinguished from an assertion that evaluation of the exact score invariant is inherently intractable.

2.3 Approximate compression

A proposed information-constrained objective is minϕ∈ℱI(Q;ϕ(Q))subject toDϕ≤ε,\min_{\phi\in\mathcal F} I(Q;\phi(Q)) \quad\text{subject to}\quad D_\phi\leq\varepsilon, and an empirical penalty of the form Ĥ[ϕ(Q)]−βÎ[πϕ;Q]+λD̂ϕ.\widehat H[\phi(Q)]-\beta\widehat I[\pi_\phi;Q] +\lambda\widehat D_\phi. The constrained specification and learned penalty serve different roles [22]. Their implementation must define a distribution on profiles, the abstraction family, estimators, and the relationship between sampled and worst-case discrepancies. Deterministic continuous outputs can have infinite mutual information. A finite penalty does not automatically enforce a hard constraint, and a validation-sample minimum cannot certify global feasibility. These are missing modeling and optimization obligations, not resolved guarantees.

3 How should uncertainty about expert weights be handled?

For W∈ΔKW\in\Delta_K, define πW,τ(a∣Q)=exp⁡(∑kWkQk(a)/τ)∑b∈Aexp⁡(∑kWkQk(b)/τ).\pi_{W,\tau}(a\mid Q)= \frac{\exp(\sum_kW_kQ_k(a)/\tau)} {\sum_{b\in A}\exp(\sum_kW_kQ_k(b)/\tau)}. Given an independently specified admissible set 𝒲\mathcal W, the stability discrepancy is dplural(Q)=supW,V∈𝒲DKL(πW,τ(⋅∣Q)‖πV,τ(⋅∣Q)).d_{\mathrm{plural}}(Q)=\sup_{W,V\in\mathcal W} D_{\mathrm{KL}}(\pi_{W,\tau}(\cdot\mid Q)\Vert\pi_{V,\tau}(\cdot\mid Q)). This checks variation across weights at a fixed profile, whereas sufficiency checks profiles within an abstraction class. Distributed aggregation and social choice motivate the distinction [1], [24], [25]; they do not establish agreement for arbitrary expert scores.

3.1 Three proposed weight families

Three candidate specifications encode different sources of uncertainty: 𝒲comp={W:Wk=eαck∑jeαcj,α∈[0,αmax]},𝒲unc={W:Wk=uk−β∑juj−β,β∈[0,βmax]},𝒲robust={W∈ΔK:∥W−Wuniform∥1≤ρ}.\begin{align*} \mathcal W_{\mathrm{comp}}&=\left\{W: W_k=\frac{e^{\alpha c_k}}{\sum_je^{\alpha c_j}},\quad \alpha\in[0,\alpha_{\max}]\right\},\\ \mathcal W_{\mathrm{unc}}&=\left\{W: W_k=\frac{u_k^{-\beta}}{\sum_ju_j^{-\beta}},\quad \beta\in[0,\beta_{\max}]\right\},\\ \mathcal W_{\mathrm{robust}}&=\{W\in\Delta_K: \|W-W_{\mathrm{uniform}}\|_1\leq\rho\}. \end{align*} Competence scores ckc_k may encode historical validation accuracy, calibration, or domain evidence. The inverse-uncertainty family requires uk>0u_k>0 and a definition of what varies. Variance across an expert’s action scores is not the same as uncertainty about those scores. A floor or other convention for zero values changes the family and must be specified. Their union is one proposed way to represent several sources of trust uncertainty. These sets are generally continuous; finite implementations require grids or sampling with stated error.

3.2 Adaptive construction and its limitation

A greedy construction starts from the uniform weight W(0)W^{(0)} and an ordered sequence of candidate batches B1,…,BJ⊆ΔKB_1,\ldots,B_J\subseteq\Delta_K. For the current profile QQ, write dQ(𝒱)=supW,V∈𝒱DKL(πW,τ(⋅∣Q)‖πV,τ(⋅∣Q)).d_Q(\mathcal V)=\sup_{W,V\in\mathcal V} D_{\mathrm{KL}}(\pi_{W,\tau}(\cdot\mid Q)\Vert\pi_{V,\tau}(\cdot\mid Q)). Set 𝒱0={W(0)}\mathcal V_0=\{W^{(0)}\}. At stage jj, let 𝒱j={𝒱j−1∪Bj,dQ(𝒱j−1∪Bj)<ε,𝒱j−1,otherwise.\mathcal V_j=\begin{cases} \mathcal V_{j-1}\cup B_j,&d_Q(\mathcal V_{j-1}\cup B_j)<\varepsilon,\\ \mathcal V_{j-1},&\text{otherwise}. \end{cases}

Proposition 3 (Conditional stability of greedy acceptance). For ε>0\varepsilon>0, every set accepted by this construction has discrepancy below ε\varepsilon, provided each supremum comparison is exact.

Proof. The singleton has discrepancy zero. At every later step, either the set is unchanged or its replacement is accepted only after the required bound has been checked. Induction proves the claim. ◻

The same conclusion holds with a sound upper-bound certificate in place of the exact supremum. Passing a sampled maximum alone is not such a certificate. A family selected from QQ also has a different interpretation from a trust family fixed independently of QQ.

The procedure need not find a largest stable set. Consider binary action probabilities 0.50,0.65,0.35,0.400.50,0.65,0.35,0.40 and threshold 0.060.06 on both directed divergences. The first is compatible with each other probability. The 0.650.65 response conflicts with 0.350.35 and 0.400.40, which are compatible with each other. Accepting 0.650.65 first yields a set of two responses, while another stable set contains three. These probabilities are realizable using weighted scores of experts (1,0)(1,0) and (0,1)(0,1) at temperature one. Maximum-cardinality selection is a separate optimization problem, and the candidate ordering is part of the greedy method.

4 What do the two checks establish?

The proposed operational gate permits action when dϕ(Q)<εanddplural(Q)<ε,d_\phi(Q)<\varepsilon\quad\text{and}\quad d_{\mathrm{plural}}(Q)<\varepsilon, and abstains otherwise. Threshold convention matters at equality and should remain consistent throughout an implementation. The criterion expresses response preservation and robustness relative to the selected map, weights, temperature, and tolerance. It does not establish the correctness of an action, the acceptability of its consequences, or a preference among conflicting objectives.

4.1 A coordination interpretation

One coordination interpretation treats a canonical abstraction as a shared norm, its preservation as norm coherence, and stability across admissible weights as norm consensus. Convention emergence and learned norms provide relevant context [31], [26], [27], [28], [29], [30]. The interpretation uses three conditions:

  1. a canonical abstraction is well-defined;

  2. the abstraction preserves relevant distinctions;

  3. admissible interpretations induce sufficiently similar action distributions.

If these terms are defined to be the two numerical checks and existence of their underlying map, equivalence to the operational gate is definitional. A theorem about common knowledge would instead require agents, possible worlds, accessibility relations, and knowledge operators [11], [12], [19], [20]. A comparison with epistemic responsibility or trust can be useful without identifying numerical agreement with such a theorem [16], [13]. The contribution here is an operational definition with a coordination interpretation. A proposed framework equivalence leaves a question for future formalization: under which epistemic model, if any, do the two numerical checks correspond to a common-knowledge property? No such equivalence is proved by the present definitions.

The same distinction limits proposed applications to neural modules, human–AI teams, and coalitions. Memory abstractions describe what information is represented [18]; the gate asks whether a specified response survives selected internal distinctions. Neither alone supplies a full theory of justified action.

5 A diagnostic example, with corrected arithmetic

Take three actions and three experts. In profile Q(1)Q^{(1)}, all experts report (1,0,0)(1,0,0). In Q(2)Q^{(2)}, only the second expert changes, reporting (1,0.1,0)(1,0.1,0). The means are (1,0,0)(1,0,0) and (1,1/30,0)(1,1/30,0), respectively. At τ=0.2\tau=0.2, p=softmax⁡(5,0,0)=(0.986703,0.006648,0.006648),q=softmax⁡(5,1/6,0)=(0.985515,0.007845,0.006640),DKL(p‖q)≃0.000096963.\begin{align*} p&=\operatorname{softmax}(5,0,0)=(0.986703,0.006648,0.006648),\\ q&=\operatorname{softmax}(5,1/6,0)=(0.985515,0.007845,0.006640),\\ D_{\mathrm{KL}}(p\Vert q)&\simeq0.000096963. \end{align*} For the second profile, trusting only expert one yields pp, while trusting only expert two gives r=softmax⁡(5,0.5,0)=(0.982466,0.010914,0.006620).r=\operatorname{softmax}(5,0.5,0)=(0.982466,0.010914,0.006620). The directed divergences are DKL(p‖r)≃0.000979478D_{\mathrm{KL}}(p\Vert r)\simeq0.000979478 and DKL(r‖p)≃0.001153451D_{\mathrm{KL}}(r\Vert p)\simeq0.001153451. Neither exceeds 0.010.01.

The two profiles do not share the same exact-mean abstraction, so their comparison does not demonstrate a sufficiency failure. Retaining the mean is exactly sufficient when the reference response is softmax of that mean. This is a sensitivity calculation; these comparisons alone give no reason to abstain at threshold 0.010.01. They also do not certify every possible profile or weight pair.

6 Can the gate be evaluated efficiently?

An empirical maximum over sampled profile or weight pairs is a lower bound on the corresponding exact supremum. A sampled violation can reject the strict gate, but a passing sample cannot certify the supremum without further assumptions. The following result states one sufficient coverage condition.

Proposition 4 (Finite-cover sampling certificate). Let (D,d)(D,d) be a nonempty metric space and let f:D→ℝf:D\to\mathbb{R} be LL-Lipschitz. Suppose balls B(cj,r)B(c_j,r), j=1,…,Nj=1,\ldots,N, with centers in DD cover DD. Let a proposal law PP satisfy P(B(cj,r))≥m>0P(B(c_j,r))\geq m>0 for every jj. For SS independent samples Z1,…,ZS∼PZ_1,\ldots,Z_S\sim P, define M̂S=max⁡if(Zi)\widehat M_S=\max_i f(Z_i) and M=sup⁡z∈Df(z)M=\sup_{z\in D}f(z). Then, with probability at least 1−Ne−Sm1-Ne^{-Sm}, 0≤M−M̂S≤2Lr.0\leq M-\widehat M_S\leq2Lr. In particular, on this event, M̂S+2Lr<ε\widehat M_S+2Lr<\varepsilon certifies M<εM<\varepsilon.

Proof. The probability of missing any particular covering ball is at most (1−m)S≤e−Sm(1-m)^S\leq e^{-Sm}. A union bound shows that all NN balls contain a sample except with probability at most Ne−SmNe^{-Sm}. On that event, for any z∈Dz\in D, choose a covering center within rr of zz and a sample within rr of that center. The sample is at distance at most 2r2r from zz, so Lipschitz continuity gives f(z)≤M̂S+2Lrf(z)\leq\widehat M_S+2Lr. Taking a supremum gives the upper bound; the lower bound follows because every sample belongs to DD. ◻

For failure probability at most δ∈(0,1)\delta\in(0,1), it is sufficient that S≥m−1log⁡(N/δ)S\geq m^{-1}\log(N/\delta). This is an unconditional high-probability certificate over the declared sampling procedure, not a posterior probability conditional on accepting. Data-dependent choice of a fiber, reuse of samples to learn the abstraction, or repeated testing needs an appropriate joint guarantee or independent evaluation.

To apply Proposition 4, the domain is a specified set of profile pairs for sufficiency, or weight pairs for stability; ff is the directed response KL on that domain. The cover, metric, Lipschitz constant, and proposal mass must be justified for that particular domain. A Lipschitz bound on softmax alone does not supply these ingredients. Fixed finite scores and compact weight sets can make smoothness arguments available, but arbitrary abstraction fibers or an arbitrary reference map need not meet a proposed uniform sampling bound.

6.1 A direct score-range certificate

Proposition 5 (A finite bound for softmax disagreement). For p=softmax⁡(x/τ)p=\operatorname{softmax}(x/\tau) and q=softmax⁡(y/τ)q=\operatorname{softmax}(y/\tau) with finite scores, put va=(xa−ya)/τv_a=(x_a-y_a)/\tau. Then DKL(p‖q)≤maxava−minava.D_{\mathrm{KL}}(p\Vert q)\leq \max_a v_a-\min_a v_a. Consequently weighted averages of scores in [0,1][0,1] have pairwise KL at most 2/τ2/\tau. If |Qk(a)−Ql(a)|≤η|Q_k(a)-Q_l(a)|\leq\eta for every k,l,ak,l,a, every pair of such weighted responses has KL at most 2η/τ2\eta/\tau.

Proof. The softmax normalizers give DKL(p‖q)=𝔼p[v]−log⁡𝔼q[ev]D_{\mathrm{KL}}(p\Vert q)=\mathbb{E}_p[v]-\log\mathbb{E}_q[e^v]. The first term is at most max⁡v\max v and the second is at least min⁡v\min v. For bounded weighted scores, each coordinate difference lies in [−1,1][-1,1]. Under the stated near-consensus condition, any two weighted averages at a fixed action differ by at most η\eta, giving va∈[−η/τ,η/τ]v_a\in[-\eta/\tau,\eta/\tau]. ◻

This bound is conservative but explicit. For example, 2η/τ<ε2\eta/\tau<\varepsilon certifies the stability check under near-consensus. It does not certify sufficiency across different profiles, and it does not automatically bound an unrestricted reference score map.

6.2 Cost and implementation

Computing one fresh dense weighted score costs O(K|A|)O(K|A|); softmax and KL cost O(|A|)O(|A|) after scores are available. Evaluating SS new dense aggregations therefore generally costs O(SK|A|)O(SK|A|). Lower cost requires a specified cache, sparse weights, or a different sample representation. Adaptive sampling, cached evaluations, parallel comparisons, and early rejection are implementation options with their own error and cost accounting.

A singleton abstraction fiber has zero sufficiency discrepancy. Lower temperature does not universally imply more rejection: different fixed score vectors sharing a unique maximizer can both concentrate on that same action as temperature tends to zero. Finite-temperature tolerances still require a quantitative calculation.

7 Extensions and unresolved questions

A possible feedback-based trust update is wk′=wkexp⁡{η𝟏{s=1}Qk(at)}∑jwjexp⁡{η𝟏{s=1}Qj(at)}.w_k'= \frac{w_k\exp\{\eta\,\mathbf1\{s=1\}Q_k(a_t)\}} {\sum_j w_j\exp\{\eta\,\mathbf1\{s=1\}Q_j(a_t)\}}. The signal ss is deliberately unspecified in this proposal. If s=1s=1 denotes regret, the displayed update increases weights on experts assigning larger scores to the regretted action. A feedback model must determine whether this sign is appropriate before the update is used as a learning rule. Sequential settings also need transition, feedback, and abstention-cost models. Strategic experts introduce a further question: what prevents advantageous misreporting, and can truthful elicitation coexist with the gate?

The immediate research question is narrower than universal agency: under which specified prediction, weighting, and sampling models can response preservation and stability be certified at a useful cost? Establishing that result would clarify which broader coordination interpretations the model can support.

Acknowledgments and provenance

I thank Richard Ngo for advice and support, and Manfred Diaz, Su Hyeong Lee, and Jeffrey Heninger for conversations. The earlier drafts described this work as independently authored and unaffiliated. The original dedication to Meridian, Cassiopeia, and Chris is retained: thank you for showing that the deepest care often needs no words. This manuscript consolidates the July 23 and July 31, 2025 drafts. Their originals and separate reconstructions are retained as historical versions. Arithmetic corrections, counterexamples, and the distinction between established results and proposed extensions are recorded in the accompanying mathematical review. No new experiment is reported here.

8 Scale, sampling, and hierarchical composition

Can the same arbitration interface be reused across atomic components, subsystems, and larger coalitions? The July 31 extension considered action-set size, expert count, hierarchical depth, and representation dimension. Those questions remain useful, but each change requires a defined transformation and a property to preserve.

8.1 Parameter normalization

Candidate normalizations are τscaled=τ0log⁡|A|,εscaled=ε0log⁡|A|,\tau_{\mathrm{scaled}}=\frac{\tau_0}{\log|A|},\qquad \varepsilon_{\mathrm{scaled}}=\frac{\varepsilon_0}{\log|A|}, with seed constants τ0,ε0>0\tau_0,\varepsilon_0>0, and a trust radius ρ0/K\rho_0/\sqrt K. These are candidate parameterizations, not a demonstrated invariance theorem. They require |A|>1|A|>1 or a separately defined singleton-action convention.

A proposed justification for these normalizations assumed DKL(p‖q)≤log⁡|A|D_{\mathrm{KL}}(p\Vert q)\leq\log|A|. This is false: for scores (1,0)(1,0) and (0,1)(0,1) at temperature 0.20.2, the divergence is about 4.933074.93307, greater than log⁡2\log2. Changing temperature also does not multiply every KL divergence by the same factor. Exact canonical partitions remain unchanged at positive temperature for the score-shift reason proved earlier; approximate gate decisions require another argument. Renaming actions, duplicating experts, adding new actions, and nesting gates are different operations. Proposition 5 gives a valid score/temperature bound, but does not establish invariance under changing the action set. The normalizations remain quantities to evaluate.

8.2 Paired expert sampling

The proposed algorithm samples distinct expert indices, compares their response distributions, rejects when a discrepancy exceeds the scaled threshold, and otherwise accepts after SS comparisons. Its runtime is O(S|A|)O(S|A|) given random access to expert predictions. This accounting does not include constructing those predictions or certifying arbitrary weighted aggregations.

A sampled maximum is not an unbiased estimator of the maximum over all pairs. With one exceptional expert among KK, a uniform pair includes it with probability 2/K2/K. Independent sampling misses it with probability (1−2/K)S(1-2/K)^S. For K=1000K=1000 and S=100S=100, this probability is about 0.8190.819. A uniform worst-case detection guarantee cannot therefore use that fixed sample budget independently of KK. If a declared pair-sampling law gives violating pairs mass at least v>0v>0, SS independent comparisons miss them with probability at most (1−v)S≤e−Sv(1-v)^S\leq e^{-Sv}. This distributional guarantee depends on its stated lower-mass condition. It does not certify every arbitrary weighted aggregation.

8.3 Recursive interfaces and composition

A gate at level ℓ\ell accepts expert scores, local parameters, and an abstraction; it returns a policy or abstains. One proposal feeds a child policy into the next level as an expert score and sets τtotal=∑ℓτℓ\tau_{\mathrm{total}}=\sum_\ell\tau_\ell. The interface is a possible design, but the proposed scalar reduction does not follow from softmax: at unit temperatures, softmax⁡(softmax⁡(1,0))≃(0.613516,0.386484),softmax⁡((1,0)/2)≃(0.622459,0.377541).\operatorname{softmax}(\operatorname{softmax}(1,0))\simeq(0.613516,0.386484),\qquad \operatorname{softmax}((1,0)/2)\simeq(0.622459,0.377541). A hierarchy therefore needs an explicit composition rule, a meaning for child abstention, and an error-propagation analysis. Summing local costs gives a computational total only after the number of gates, input construction, and local sample budgets are specified.

8.4 What can be checked without an invariance claim?

Consistent permutation of actions preserves softmax probabilities up to the same permutation and leaves KL unchanged. A gate built entirely from such equivariant operations can preserve its accept/abstain decision under relabeling. This observation does not imply invariance when the number of actions changes.

A verification plan includes expert-count comparisons, action-set changes, hierarchical wrapping, and shared seed parameters. These remain useful implementation experiments if outputs, costs, and failure frequencies are recorded. Equality of configuration fields or interface types is an interface test, not evidence that all epistemic guarantees survive a transformation. Which measurable invariance should the hierarchical framework target first?

References

[1]

Kenneth J. Arrow. Social Choice and Individual Values. Wiley, 1951.

[2]

Richard Ngo. Towards a scale-free theory of agency. AI Alignment Forum, 2025. https://www.alignmentforum.org/posts/gHefoxiznGfsbiAu9/towards-a-scale-free-theory-of-agency.

[3]

Nelson Elhage et al. A mathematical framework for transformer circuits. Transformer Circuits Thread, 2021.

[4]

Stuart Russell. Human Compatible. Viking, 2019.

[5]

Yuntao Bai et al. Constitutional AI: Harmlessness from AI feedback. arXiv:2212.08073, 2022.

[6]

Christian List and Philip Pettit. Group Agency. Oxford University Press, 2011.

[7]

Leonard J. Savage. The Foundations of Statistics. Wiley, 1954.

[8]

Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction, second edition. MIT Press, 2018.

[9]

Richard Ngo. Well-foundedness as an organizing principle of healthy minds and societies. Mind the Future, 2025.

[10]

Niklas Lauffer, Ameesh Shah, Micah Carroll, Michael Dennis, and Stuart Russell. Who Needs to Know? Minimal Knowledge for Optimal Coordination. 2023. https://arxiv.org/abs/2306.09309.

[11]

Ronald Fagin, Joseph Y. Halpern, Yoram Moses, and Moshe Y. Vardi. Reasoning About Knowledge. MIT Press, 1995.

[12]

Brian F. Chellas. Modal Logic: An Introduction. Cambridge University Press, 1980.

[13]

Alvin I. Goldman. Knowledge in a Social World. Oxford University Press, 1999.

[14]

C. K. Chow. On optimum recognition error and reject tradeoff. IEEE Transactions on Information Theory, 16(1):41–46, 1970.

[15]

Yonatan Geifman and Ran El-Yaniv. Selective classification for deep neural networks. NeurIPS, 2017.

[16]

Richard Ngo. Trust develops gradually via making bids and setting boundaries. LessWrong, 2023.

[17]

Jan Kulveit. Hierarchical agency: A missing piece in AI alignment. AI Alignment Forum, 2024.

[18]

Aaron Kirtland, Alexander Ivanov, Cameron Allen, Michael L. Littman, and George Konidaris. Memory as state abstraction over trajectories. 2025.

[19]

Robert J. Aumann. Agreeing to disagree. The Annals of Statistics, 4(6):1236–1239, 1976.

[20]

Yoram Moses and Moshe Tennenholtz. Distributed epistemic algorithms: Knowledge, coordination and common knowledge. TARK, 1995.

[21]

Edwin T. Jaynes. Probability Theory: The Logic of Science. Cambridge University Press, 2003.

[22]

Thomas M. Cover and Joy A. Thomas. Elements of Information Theory. Wiley, 2006.

[23]

Erich L. Lehmann and George Casella. Theory of Point Estimation. Springer, 2005.

[24]

Morris H. DeGroot. Reaching a consensus. Journal of the American Statistical Association, 69(345):118–121, 1974.

[25]

Michel Grabisch, Jean-Luc Marichal, Radko Mesiar, and Endre Pap. Aggregation Functions. Cambridge University Press, 2009.

[26]

Abram Demski. Learning normativity: A research agenda. AI Alignment Forum, 2020.

[27]

Abram Demski. Four motivations for learning normativity. AI Alignment Forum, 2021.

[28]

Ninell Oldenburg and Tan Zhi-Xuan. Learning and sustaining shared normative systems via Bayesian rule induction in Markov games. AAMAS, 2024.

[29]

Zhi-Xuan Tan and Desmond C. Ong. Bayesian inference of social norms as shared constraints on behavior. arXiv:1905.11110, 2019.

[30]

Zhi-Xuan Tan, Jake Brawer, and Brian Scassellati. That’s mine! Learning ownership relations and norms for robots. arXiv:1812.02576, 2019.

[31]

David K. Lewis. Convention: A Philosophical Study. Harvard University Press, 1969.