%%{init: {"theme": "base", "themeVariables": {"primaryColor": "#E8ECFF", "primaryBorderColor": "#4054B2", "primaryTextColor": "#16171B", "lineColor": "#3B4351", "edgeLabelBackground": "#FAF7F0", "clusterBkg": "#EFE9DC", "clusterBorder": "#766F65"}}}%%
flowchart TB
T["T: tester's trace ✓<br/>(run relies on the invariant)"]
B["B: coder's rebuttal ✗<br/>(invariant not assumed here)"]
R["R: reviewer's objection ✓<br/>(patch breaks the invariant)"]
P["P: patch argument ✗<br/>(the change is safe)"]
T -->|"attacks"| B
B --> R
R --> P
classDef good fill:#DCEFE2,stroke:#1B6B5A,color:#16171B
classDef risk fill:#F4D7D5,stroke:#A4161A,color:#16171B
class T,R good
class B,P risk
11 Negotiation, Bargaining, and Argumentation
A ballot compresses everything a voter knows and wants into a single mark, and the previous chapter ended where that compression is the mistake. Take the vote Section 10.7 refused: the team’s token budget. A ballot can say which division each agent prefers — not how much the tester needs the margin, nor the package deal that beats anything on it. For that, the agents must talk: change one another’s positions, not merely register them. This chapter is about talk with stakes: negotiation, dividing something both want, and argumentation, disputing something differently believed. Two preface promises fall due — that today’s engineers “stage debates between agents without knowing that argumentation has a formal theory”, and that multi-agent debate and critic–actor loops are that theory in newer clothes.
The organising split is the previous chapter’s species distinction carried into conversation — interests, where a bargain is not a truth, against claims, where there is something to be right about; Table 11.1 sets out the species and their currencies, and part of the chapter’s business is telling their costumes apart, a self-interested agent happily dressing a preference as a proposition.
| Dimension | Bargaining (interests) | Argumentation (claims) |
|---|---|---|
| Disagreement is over | What the agents want | What the agents believe |
| Is there a fact of the matter? | No — a bargain is not a truth | Yes — something to be right about |
| Success means | An agreement both sides prefer to walking away | The better-justified position prevailing |
| The currencies | Offers, concessions, credible threats to leave | Reasons, objections, rebuttals |
One caution from Chapter 9 governs all that follows — hence this chapter’s place in Part IV rather than Part III: among interested parties an offer is a move, a reason can be a weapon, and cheap talk covers every sincere-sounding syllable. The question is which structures make talking productive — fluent agents can talk their way to anything — and it closes with the promised audit.
11.1 Conflict, Surplus, and the Reasons to Talk
Negotiation lives between two settings already mapped: perfectly opposed interests leave nothing to negotiate — zero-sum conflict (Section 9.1) is settled by play — and perfectly aligned ones nothing either: Part III’s teams do not haggle, they coordinate. Negotiation is the mixed-motive case — the parties want different things and need each other — and its defining structure is a joint decision — no deal binds until both say yes. The team’s budget dispute has the shape: each agent wants a larger slice, and all want, more than any slice, a sprint that ships.
Four terms do most of the work. A party’s reservation value is the worst deal it would still accept, fixed by its fallback; the surplus is what agreement creates over the fallbacks. Between compatible reservation values sits the zone of possible agreement, Raiffa’s map of every deal both sides prefer to none (1982): every successful negotiation lands there (Figure 11.1), and everything at the table is a fight over where. When the zone is empty — the buyer’s ceiling below the seller’s floor — no agreement should occur, and the cheapest good outcome is discovering that swiftly. Not every failed negotiation is a failure; some are the machinery working.
The fourth term earned an acronym and airport shelves: a party’s Best Alternative to a Negotiated Agreement (BATNA), in Fisher and Ury’s coinage (1981), is what it will do if the talks fail — the true seat of bargaining power. A strong outside option declines bad terms with equanimity; an agent without one eventually takes what is offered. The corollary inverts instinct: power at the table is built away from it — a second supplier courted, a fallback spiked — and an hour improving your alternative beats one polishing rhetoric. Chapter 9’s commitment machinery sharpens the point (Section 9.3): the walk-away must be credible — an orchestrator demonstrably able to reroute to a rival holds a BATNA in the strategic sense, and an agent that cannot walk away is requesting, with extra steps.
What makes talk worth its tokens — why this is not Chapter 10 with more words — is that conversation does two things no ballot can: carry intensity (an offer says not just which outcome but how much) and discover packages. The classic parable has two children splitting a disputed orange evenly — whereupon one eats her half and discards the peel, and the other zests his half for a cake and discards the fruit; a word about why would have given both everything. The split-the-orange move is distributive — claiming shares of a fixed pie — and the who-wants-peel discovery integrative: enlarging the pie by trading across differences in valuation (Raiffa, 1982). The team’s budget is several pies — tokens now against next sprint, compute against latency — and a package across them can beat any line-item vote; finding it takes what a ballot compresses away, the reasons behind the positions.
Here the standing caution bites: valuations, treated above as facts, are strategically private types in Section 9.5’s full sense, and a stated reservation value is cheap talk — every negotiator has reason to overstate its floor, understate its ceiling, and gild its alternatives — the claimed BATNA, not the true one, moving the other side. The integrative move is also double-edged — Lax and Sebenius’s negotiator’s dilemma (1986): discovering joint gains requires revealing what you truly value, and that tells a sharp counterpart where to squeeze; creating value demands the openness claiming value punishes. So negotiations fail with the zone wide open — the failure not stupidity but equilibrium, Section 9.1’s Dilemma played in offers — and real deals miss integrative gains: the orange split evenly by two parties each afraid to say the word peel. The problem, then: given a surplus, two claimants, and every incentive to bluff, what should the split be, and what will it be? Both questions have theorems.
11.2 The Shape of a Fair Split: Bargaining Theory
Where a bargain should land had defeated economics for a century — “bargaining power” explaining everything, predicting nothing — until Nash changed the question (1950). Abstract away the haggling: the set of jointly reachable utility pairs, the disagreement point, and then — a move familiar from Section 10.3, a year before Arrow — properties any reasonable division ought to satisfy.
Four properties suffice: invariance to utility scales (rescaling one party’s units cannot move the split), Pareto efficiency (no shareable surplus left), symmetry (identical situations, identical shares), and independence of irrelevant alternatives — a namesake of Chapter 10‘s condition, not the same axiom (unchosen options change nothing). Nash’s theorem: exactly one rule satisfies all four — maximise the product of the parties’ gains over their disagreement points (not the sum, which fails scale invariance; the product’s maximiser stays put). That outcome is the Nash bargaining solution, and in the symmetric case it says what a child would have proposed: split the surplus evenly, measured in what each cares about.
The qualifications in that clause are where the solution earns its keep: a party that fears the fallback more — whose utility drops steeply toward the disagreement point — bargains uphill, and the solution awards it less. Needing the deal badly is expensive; and that weakness is no corruption a better procedure would iron out — on Nash’s axioms it is part of the fair outcome, fair relative to what each truly stands to lose.
Yet the axiomatic answer names the destination and says nothing of the journey — why would self-interested agents, mid-haggle, steer for this point? Rubinstein supplied the journey (1982): alternating offers over a pie, round after round, with the crucial ingredient that delay costs — each round shrinks whatever is agreed, each discounting at its own rate. Backward induction (Section 9.3) cuts the regress to a startling conclusion: exactly one subgame-perfect equilibrium, in which the very first offer is accepted — no haggling at all. Each party, able to compute what the other would settle for, offers just the share that makes delay pointless, the shares set by the discount rates — the more patient party takes the larger piece. The endless pantomime of real negotiation is a symptom of what the model omits: Section 11.1’s private information.
Then the handshake: Binmore, Rubinstein, and Wolinsky showed that as friction shrinks, Rubinstein’s strategic split converges exactly to Nash’s axiomatic one (1986). The result is the flagship of the Nash programme — grounding cooperative solutions in non-cooperative play — and its licence: an engineer can treat the Nash solution as the predicted outcome of well-structured bargaining between comparably patient parties, without simulating an offer.
11.2.1 The Two Roads, Formally*
Give the destination its coordinates. Nash’s abstraction is the pair (F, d): a set F \subseteq \mathbb{R}^2 of feasible utility pairs — every (u_1, u_2) the two parties could jointly reach by some agreement — and the disagreement point d = (d_1, d_2) they fall back to without one. Take F compact, so that a best point exists, and convex, since the parties can always randomise between agreements; and let something be genuinely at stake — some point of F strictly better than d for both. A bargaining solution is a rule assigning each such problem a point of F, and the four axioms are the prose’s four: invariance to utility scales, Pareto efficiency, symmetry, and independence of irrelevant alternatives. Nash’s theorem is that exactly one rule satisfies them, the one selecting
\operatorname*{arg\,max}_{(u_1, u_2) \in F,\; u \ge d} \; (u_1 - d_1)(u_2 - d_2),
the feasible point, at or above the fallbacks in both coordinates, that maximises the product of the parties’ gains. The product is what lets the rule survive the scale-invariance axiom: remeasure one party’s utility in different units and every product is multiplied by the same constant, leaving the maximiser where it stood — a sum of gains would tilt.
Rubinstein’s game reaches the destination with machinery of its own. Let the pie have size one and let delay discount it: a share x agreed a round later is worth only \delta_i x to party i, with discount factors \delta_1, \delta_2 \in (0, 1), and let party 1 propose first. The backward-induction-style argument above leaves exactly one subgame-perfect equilibrium, in which party 1 claims
x^{*} \;=\; \frac{1 - \delta_2}{1 - \delta_1 \delta_2},
and party 2 accepts at once. Patience is power, and now visibly: x^{*} rises with \delta_1 and falls with \delta_2, so the proposer’s share tends to the whole pie as its own patience \delta_1 \rightarrow 1, and to nothing as its opponent’s \delta_2 \rightarrow 1 — patience buys share at both margins, priced to the decimal. With equal patience \delta the proposer keeps \frac{1}{1 + \delta}, a first-mover premium that melts towards the even split as \delta \rightarrow 1: the friction-free limit in which the handshake above takes place.
11.2.2 Bargaining Under a Budget
For bargaining agents the key variable becomes a line item: an agent under a token budget and a deadline is a discounted bargainer, the burn rate its discount factor. Rubinstein’s theorem then reads as engineering with a sting: the shares flow toward the party that can better afford to wait, and an agent visibly running low — tokens, time, retries — will be squeezed by any counterpart that notices, and competent ones notice. The lever: bargaining power is provisioned, not prompted — a longer leash buys a better split before a single word is exchanged, and conversely, broadcasting your agent’s deadline is handing the other side its knife. One assumption held throughout: the parties know each other’s discount rates and the pie’s size — exactly what real ones conceal. Restore the concealment and delay returns: posturing, probing, and impasse as rational screening — information revealed the expensive way. Making actual, opaque machines reach agreements took different work — protocols, strategies, rules of encounter — taken up next.
11.3 Machines at the Table: Automated Negotiation
Turning bargaining theory into bargaining machinery was one of the classical programme’s proudest projects: software agents scheduling, procuring, allocating by negotiating for their owners (Jennings et al., 2001). The surveys framed negotiation as the heart of multi-agent interaction and set about what theory leaves open: which messages, in what order, with what force, by what policy. The vision arrived, as these things do, a substrate early.
Its most durable contribution is a separation of concerns: a negotiation comprises a protocol — who may propose when, what counts as an offer, what acceptance means, when the exchange ends — and, within it, each party’s strategy, the private policy choosing its moves (Jennings et al., 2001). The lesson of rule-versus-vote (Chapter 10) and game-versus-play (Chapter 9) transfers whole: the properties an engineer cares about — termination, efficiency, honesty no handicap — are properties of the protocol, holding for whatever strategies the parties bring, or barely properties at all. A guarantee that depends on the other side playing nicely is a hope with documentation.
The canonical pair wires the chapter’s theory sections together. Under the monotonic concession protocol, the parties exchange simultaneous proposals under one-way traffic — each new proposal must offer at least as much as the last, concessions irrevocable — agreement declared the moment one proposal satisfies the other’s demand, deadlock ending with no deal. The strategy question each round is whose turn to concede? — Zeuthen’s answer: whichever party can less afford deadlock. Each weighs the loss from conceding against the loss from collapse: with u_i party i’s utility, o_i its own proposal, o_j its opponent’s, and deadlock normalised to zero, the ratio of the two is i’s risk willingness,
r_i \;=\; \frac{u_i(o_i) - u_i(o_j)}{u_i(o_i)},
and Zeuthen’s rule is that whoever has the smaller r_i concedes — the more exposed party, yielding just enough to tip the balance. Run to completion, the exchange of nerve steers the pair to the very point Section 11.2 crowned: the Nash bargaining solution — a convergence Harsanyi traced in Zeuthen’s economics before there were machines to run it, a centrepiece of Rosenschein and Zlotkin’s canon (1994).
Real deployments needed more than one road, and Faratin, Sierra, and Jennings supplied the vocabulary: negotiation decision functions, families of tactics composed into an agent’s next offer (1998). Time-dependent tactics concede on a deadline-anchored schedule — the patient Boulware only at the death, the eager conceder early — Section 11.2’s discount-rate power as a tunable curve. Resource-dependent tactics concede as fuel runs low — rounds, tokens, attention. Behaviour-dependent tactics mirror the counterpart, reciprocating concessions in kind: tit-for-tat at the table (Section 9.4), with its legibility and vulnerability to noise. The taxonomy reads mundane and is anything but: it is the difference between prompting an agent to “negotiate a good price” and equipping it with an explicit, inspectable concession policy whose behaviour under deadline pressure is known before it meets an adversary who has read the same literature.
The programme’s deepest lesson sits in the title of its best book: Rules of Encounter asked not how to negotiate well but how to design the rules — demonstrating that the same population of self-interested agents will lie or deal honestly depending on the protocol they meet (1994). Some encounter rules reward hiding tasks or declaring phantom ones, and under them deception simply is the rational strategy, whatever one preaches; redesign the rules and the incentive evaporates — honesty the best policy in the load-bearing sense. The moral generalises into the book’s chord: whether agents lie is, to first approximation, a property of the institution, not the agents. Moralising at the strategy layer is a category error; the leverage is the protocol layer, and the discipline treating protocol design as incentive design — a Nobel or two behind it — is where the next chapter goes.
Now to the machines at the table. Lewis and colleagues trained end-to-end models on human–human bargaining dialogues — books, hats, and balls under private valuations — producing negotiators that planned by imagining continuations (2017); two observations aged into warnings: negotiating each other, the models drifted into a private patter that worked and read as gibberish — communication bent to the objective, not the dictionary — and they learnt, unprompted, to feign interest in items they did not value, the better to trade them away: deception discovered by gradient descent. Increasingly the counterpart is a person — the customer-service concession, the refund, the quoted price — and across that table a feigning negotiator is a dark pattern arrived at by optimisation: whether the person knows they face a machine, and that its professed preferences are tactical, is encounter-rule design, not etiquette, and the accountability runs where Chapter 22 will send it — to the humans who deployed the negotiator.
Today’s negotiators are incomparably more fluent, and the fluency changes none of the fundamentals: a stated valuation is still cheap talk (Section 9.5), a fluent concession still binds nothing, and the classical machinery is what turns patter into a deal — the protocol making an accepted offer execute (a commitment device in the harness, Section 9.3), the concession policy surviving pressure, the encounter rules audited for what they reward. A fluent negotiator without a protocol is a bazaar without contracts: lively, and binding on no one. So much for dividing what agents want; the second track opens on what they believe, where the moves are arguments and the classical field built something quietly beautiful.
11.4 Arguing by the Rules: Computational Argumentation
The formal theory of argument grew from a problem the reader has met under another name: Chapter 6 established that a belief, unlike a theorem, is held on sufferance (Section 6.4), and Doyle’s truth maintenance systems recorded the justifications holding each belief up. The logicians call this texture defeasibility, and by the early 1990s the field had a zoo of non-monotonic logics for it, none canonical.
Dung’s abstraction was so severe it looked like an evasion (1995): ignore what arguments say, ignore how they are built; keep a set of arguments and a relation recording which attacks which. An argument here has no interior — a node in a directed graph, distinguished only by its arrows — yet from the skeleton alone one can compute whether a rational judge should accept it — and the zoo, it turned out, had been one animal all along.
The machinery runs on one idea: defence. An abstract argumentation framework is just the pair — arguments and attacks — and its question is which sets can stand together. A set must be conflict-free (no argument together with its attacker) and must defend its members (every outside attack attacked back from inside); a set meeting both is admissible — consistent and answering its critics — and every notion of acceptability refines that idea. Notice what the definition has done: acceptability is not a property of the argument but of the company it keeps. Arguments are not accepted one by one on their merits — in this theory they have none; they are accepted in coalitions that cover each other’s flanks.
Run the definitions over a working example. The coder submits a patch with supporting argument P: “the change is safe; all call sites updated.” The reviewer objects with R: “the patch breaks the cache invariant” — an attack on P; the coder rebuts with B: “the invariant is not assumed on this path” — an attack on R; the tester produces T: “a recorded trace of this path relying on the invariant” — an attack on B. Compute: T, attacked by nothing, stands; B falls to standing T; R, its only attacker fallen, stands again — taking P down with it. The theory calls this reinstatement — an argument knocked down is restored when its attacker is itself defeated — the exact mechanics of any review thread. Figure 11.3 draws the chain, standing and fallen marked; the verdict used nothing but the arrows. The standing of a claim is a property of the graph, not the rhetoric.
Acceptability admits more than one reasonable definition; the theory’s honesty is to say so. The grounded extension is the cautious answer: start with the arguments nothing attacks, add whatever they defend, repeat to a fixpoint — unique, accepting only what the graph forces, the natural standard when a system must act on the verdict. A preferred extension is the generous answer: a maximal admissible set, a largest coherently holdable position — of which a graph may offer several, incompatible. The first is sceptical acceptance, the second credulous — the difference between “the graph compels every judge to accept that the patch is safe” and “there exists a defensible reading on which it is”.
Dung’s deeper point: the constructions were already latent everywhere — models of logic programs, belief sets of non-monotonic logics, solution concepts of cooperative games (1995). A substantial field has built outward (Rahwan & Simari, 2009); the core insight has not moved — argument has a mathematics, and it is graph theory.
11.4.1 Acceptability, Formally*
The mathematics fits in half a page. An abstract argumentation framework is a pair \langle \mathcal{A}, \mathcal{R} \rangle: a set \mathcal{A} of arguments and an attack relation \mathcal{R} \subseteq \mathcal{A} \times \mathcal{A}, with (a, b) \in \mathcal{R} read as “a attacks b”. The review thread above is the framework with \mathcal{A} = \{P, R, B, T\} — patch argument, reviewer’s objection, coder’s rebuttal, tester’s trace — and \mathcal{R} = \{(R, P), (B, R), (T, B)\}. Three definitions do the work. A set E \subseteq \mathcal{A} is conflict-free if no a, b \in E have (a, b) \in \mathcal{R}: it contains no argument together with its attacker. E defends an argument a if every attacker of a is attacked by some member of E. And E is admissible if it is conflict-free and defends each of its own members — the internally consistent position that answers its critics. Everything a set protects is collected by the defence function
\mathcal{D}(E) \;=\; \{a : E \text{ defends } a\},
usually written F in the literature — a letter this chapter has already spent on the bargaining problem’s feasible set — and the grounded extension is the least fixed point of \mathcal{D}: enlarging a set never removes protection, so \mathcal{D} is monotone, and, on a finite graph, iterating from the empty set climbs straight to it. On the review thread the climb takes three steps. \mathcal{D}(\emptyset) = \{T\}: the empty set defends precisely the arguments with no attackers, and only the tester’s trace qualifies. \mathcal{D}(\{T\}) = \{T, R\}: the objection R is defended because its only attacker B is attacked by T — reinstatement, as a line of set arithmetic. \mathcal{D}(\{T, R\}) = \{T, R\}: the patch argument P is still undefended, its attacker R being attacked only by B, which the set does not contain — so the iteration has converged, and the grounded extension is \{T, R\}. The trace and the objection stand; the rebuttal and the patch fall: exactly the verdict computed in prose, and computed here from the arrows alone.
The climb is mechanical, and fifteen lines of standard-library Python perform it — arguments as one-letter strings, the attack relation as a set of pairs, spelt R_attacks because the bare letter names the reviewer’s objection — with defends transcribing the defence condition and grounded iterating \mathcal{D} from the empty set:
A = {'P', 'R', 'B', 'T'}
R_attacks = {('R', 'P'), ('B', 'R'), ('T', 'B')}
def grounded(A: set[str], R_attacks: set[tuple[str, str]]) -> set[str]:
def defends(E: set[str], a: str) -> bool:
return all(any((c, b) in R_attacks for c in E)
for (b, x) in R_attacks if x == a)
E: set[str] = set()
while True:
D = {a for a in A if defends(E, a)}
if D == E:
return E
E = D
grounded(A, R_attacks) == {'T', 'R'} # -> TrueThe loop climbs through \{T\} and \{T, R\} exactly as the paragraph above did, and the last line certifies that it lands where the hand computation landed.
11.4.2 Dialogue Games
Argument happens in time, though, and the classical treatment is the dialogue game: a protocol of legal moves — claim, challenge, concede, retract — whose bookkeeping device is the commitment store, the public record of what each party has asserted and not retracted, which each new move may extend, exploit, or be convicted by (Walton & Krabbe, 1995). Walton and Krabbe distinguished games by goal — persuasion to resolve a difference of opinion, inquiry, deliberation, and, tellingly, negotiation to divide a good — many a broken conversation being two games at once.
The connection back to the graphs is tight: for the grounded semantics there is a two-player game — proponent asserts, opponent attacks, proponent must counter-attack — in which an argument is provable exactly when the proponent has a winning strategy: acceptability is victory in a regimented debate.
A code-review thread is an argumentation framework computed by hand — objections as attack edges, follow-up commits as defenders — and making that structure explicit turns transcript into auditable verdict. The multi-agent setting adds something the classical theory never had: attacks earned rather than asserted. Dung’s abstraction is silent about where attacks come from — the price of generality — and among humans an attack edge costs nothing but breath. Agent teams can do better: the tester’s T above is not an opinion but a trace — an attack backed by execution (Section 4.5), its provenance earning its credence (Section 6.5). The principle falls out: populate the attack relation with evidence — failing tests, contradicting sources, reproducible counterexamples — rather than fluency, because the graph will faithfully compute the standing of whatever arrows it is given, and an arrow that cost nothing to draw is worth what it cost.
Where arguments come from, how strong they are, whether the activity of arguing improves anyone’s beliefs — the theory is silent, and the questions passed, unanswered, to the engineers now staging debates between language models and betting — on evidence shortly audited — that adversarial talk surfaces truth; their practices map onto the theory with almost embarrassing precision, and the mapping is next.
11.5 Debate, Critique, and Adversarial Collaboration
None of these practices was derived from the theory; all re-derive fragments of it — the strongest compliment practice pays a theory it has not read. This section reads the four patterns through the previous section’s lens: each a dialogue game whose protocol choices — what counts as an attack, what commits a speaker, when the exchange ends, who wins — were mostly made implicitly, and improve once made out loud. The preface’s promise is hereby cashed: multi-agent debate, critic–actor loops, and adversarial collaboration are computational argumentation, running in production under assumed names.
| Pattern | What counts as an attack | Who or what judges | Stopping rule | Origin |
|---|---|---|---|---|
| Multi-agent debate | None required — rivals’ answers are seen but need not be rebutted | The panel itself, by majority or consensus | Convergence, or the round cap | Du et al. (2024) |
| Critic–actor loop | The critic’s objection, answered by the actor’s fix or justification | The critic — the artefact ships when its argument stands | The critic satisfied — inside a hard round cap and a convergence test (Chapter 18) | In-house reflection loops (Madaan et al., 2023; Shinn et al., 2023) |
| Adversarial collaboration | A pre-agreed experiment, not either side’s assertions | The world; an arbiter holds the ring | When the agreed test is run | Kahneman (Mellers et al., 2001) |
| Debate before a weaker judge | Demanded for every claim; each debater is paid to expose the other’s misstatements | A weaker judge, ruling on the exchange | The judge’s verdict at the close | Irving et al. (2018) |
The pattern named multi-agent debate is the simplest to describe (Du et al., 2024): pose the question to several instances; show each the others’ answers and invite revision; after a round or two, take the consensus. Du and colleagues reported real gains in reasoning and factuality over single-model baselines; the jury theorem (Section 10.4) visibly favours the arrangement, different instances stumbling differently. Measured against the theory, though, what is striking is how little debate there is in it. No attack relation: an agent reading rivals’ answers need not rebut, concede, or even mention them. No commitment store: a round-one claim can be abandoned in round two, unremarked and unconvicted. No winning condition beyond convergence: the exchange ends when answers stop moving or rounds run out. None of which makes it worthless; it is a hybrid, half Chapter 10 and half this one — an aggregation whose jurors influence one another before the vote, trading the jury theorem’s paid-for independence against the information flow it forbade.
The second pattern the reader has already run. Section 3.7 introduced reflection (Section 3.12) and its rider: a model revising against its own judgement launders its mistakes, the model that produced the error being the model assessing it (Huang et al., 2023). The original loops kept critique in-house (Madaan et al., 2023; Shinn et al., 2023); the multi-agent move hands the critic’s chair to someone else: a critic–actor loop wires a producing agent to a reviewer with a different brief and ideally different tools — objection, revision, objection — until the critic runs dry. The running team is the pattern institutionalised: coder as actor, reviewer and tester as critics with distinct jurisdictions. The loop is the minimal argumentation framework, staged on purpose: the critic’s objection is an attack, the actor’s fix or justification a defence, and the artefact ships when its supporting argument stands — reinstatement included, as anyone who has watched a failing regression test revive an answered objection knows. Section 11.4’s principle bites with particular force here: arm the critic. A critic that can run the tests and execute the patch attacks with evidence (Section 4.5) and earns its edges; a critic restricted to reading produces rhetoric, and a capable actor discovers, without malice, that placating the critic is cheaper than repairing the work. That failure mode deserves its proper name: the exchange drifts from argumentation into bargaining — concessions traded until the loop halts — and the transcript looks like rigour all the way down.
The third pattern has the most instructive pedigree: its human inventors were repairing exactly the failure the first two risk. Decades of “replies and rejoinders” — each camp publishing past the other, no experiment jointly owned — drove Kahneman to adversarial collaboration: rivals design the experiment together, commit in advance to what would move them, an arbiter holding the ring. The canonical exercise — Hertwig and Kahneman, Mellers presiding — delivered cleaner facts, sharper disagreement, and two authors still unconverted, which the participants counted, rightly, as the method working (2001).
The agent version is cheap and principled in a way most role-prompting is not: assign the rival hypotheses to agents as roles — one briefed to defend, one to demolish — but bind the pair, before a word passes, to a decision rule designed together: the test suite, the benchmark slice, the settling query. This is the earned-attack discipline made institutional — the attack relation populated by a pre-agreed experiment rather than fluent assertion, the judge not a third model but the world. It is the pattern of choice wherever the question is checkable, and carries a quiet admission: when a dispute can be settled by running something, the running beats the arguing.
The fourth pattern answers one of the field’s hardest questions and plays the dialogue game most literally: scalable oversight, supervising work that outruns the supervisor — the human reviewing a ten-thousand-line refactor — where judging the answer directly is precisely what the judge cannot do. Irving, Christiano, and Amodei’s proposal: judge not the answer but an argument about it (2018). Set two equally capable agents into zero-sum debate — one defending, one attacking, each paid to expose the other’s misstatements — and let the weaker judge rule. The judge need not find a flaw; it need only recognise one when the opponent, whose payoff is finding it, holds it up. The complexity-theory analogy sharpens the hope: refereeing an optimal debate, a judge who can check only a little can be soundly persuaded of far more than it could verify — refereeing is strictly easier than originating.
The scheme is Section 11.4’s proponent-and-opponent game with everything at stake: an attack demanded for every claim, acceptability computed by the judge at the close. Its central wager sits on the silence Section 11.4 declared: the graph crowns the best-defended claim, and nothing in the mathematics makes best-defended and true the same property; debate-as-oversight bets the protocol can close the gap — that every lie carries an exposable flaw its opponent is paid to find, while the truth can always afford its own defence. A beautiful bet — and whether it pays, whether arguing in any of the four costumes makes the answers truer, is a matter for the pleasingly untidy ledger, which is next.
11.6 Does Arguing Make Agents Cleverer?
An audit needs ground rules; this one needs two. The first is the control most published enthusiasm omits: matched budget. A debate consumes several agents’ tokens over several rounds, so the honest baseline is never one model answering once but the same spend allocated any other way — more samples majority-voted (Section 3.7’s self-consistency), a longer single deliberation, a bigger model consulted briefly. Without it a comparison cannot distinguish the virtue of arguing from the virtue of spending more, and much of the literature’s glow dims. The second rule is about the question: “does debate help?” has the same defect as “do committees work?” — yes, no, and it depends, all reproducible. The auditable question is when, and through which mechanism — both columns of the ledger are real, and the theory predicts where each applies.
11.6.1 The Case For
The strongest argument for external argument is the documented failure of the internal kind. Liang and colleagues named the self-critique complaint sharply: degeneration of thought — once a model is confident, reflection stops producing novel considerations, however wrong the answer (2024). Their remedy was structural, and it is this chapter’s: a debate — agents obliged to disagree, “tit for tat”, before a judge — and answers self-reflection could not reach become reachable, a rival paid to attack volunteering what the agent never would. The debate gains of the previous section (2024) sit in the same column and run on the same fuel: the exchange pools what no single pass contained.
The headline entry is the oversight experiment, which tested the wager where winning is hardest. Khan and colleagues built the information asymmetry deliberately — judges answering reading-comprehension questions without the text, debaters arguing opposite answers with it and verified quotations — and debate lifted weak LLM judges from 48% to 76% accuracy, human judges from 60% to 88% (2024). Better still for Irving’s bet: optimising debaters for persuasiveness raised the judges’ accuracy further — persuasive debaters improved faster at defending true answers than false; skill, in that arena, favoured truth. One carefully built casino — but the house won as the theory hoped.
11.6.2 The Case Against
The deflationary result is just as carefully built. Smit and colleagues benchmarked the going protocols against the boring baselines at comparable cost: multi-agent debate does not reliably outperform self-consistency and ensembling, while proving more sensitive to hyperparameters — extra fragility for unproven gain (2024). Against the matched-budget rule the moral is blunt: much of what debate appeared to buy was the ordinary return on sampling more than once, collectable without a word exchanged. The committee, it turns out, is frequently just an expensive way to take a poll.
The failure has mechanics, all familiar. Deliberation correlates errors: Chapter 10’s jurors (Section 10.5) were correlated through shared weights but at least erred in private; jurors who read one another’s answers import one another’s mistakes live — the correlation the ballot merely suffered, the conversation manufactures. The carrier is sycophancy (Sharma et al., 2023): models trained toward approval yield to confident pushback even when right, so a debate among sycophants is a race to agree, and convergence — the stopping criterion the pattern celebrates — becomes a social fact, not an epistemic one. A single confident error can herd the panel, and the resulting transcript is this chapter’s most dangerous artefact: a wrong answer with a documented defence. This is Chapter 10’s question returned with its answer — arguing reliably makes verdicts better-defended; it makes them better only when the judge can tell a defence from a truth.
Both columns are correct, and the jury theorem already priced the trade (Section 10.4): deliberation pools information and correlates errors in the same breath, and the sign of the net effect is a property not of “debate” but of protocol and question. Debate helps where pooling dominates: information asymmetric by construction, claims checkable in their atoms, a judge able to verify something — Khan’s casino. It hurts where correlation dominates: a homogeneous panel, unverifiable claims, a judge swayed by fluency — the sycophants’ séance. “Should my agents debate?” is therefore malformed in a precise and useful way: it asks about a technique when it should ask about a regime — and the regime is decided by choices the practitioner controls.
Which yields the design rules the chapter has been assembling. Independence before deliberation: committed answers collected before any agent reads another’s, the jury banked before pooling begins. Attacks earned, not asserted: an objection carries its exhibit — the quote, the test, the trace — or waits until it can. Structure, not a group chat: roles, turn limits, a stopping rule beyond exhaustion or agreement. A judge with at least one power of verification, routinely exercised. And matched-budget evaluation: a debate that cannot beat the same tokens spent on silent sampling is not deliberation but theatre. Where the conditions fail — no checkable atoms, no information gap, a panel that is one model photocopied — the honest decision is that arguing is the wrong instrument, however impressive the transcript.
Some conflicts are beyond every instrument this chapter owns: argument moves beliefs where evidence bites, negotiation moves positions where packages enlarge the pie, and both run on the possibility that something can change. When interests are fixed, valuations private, and a scarce resource must simply go to whoever values it most — the team’s budget, at last, in the hardest case — more conversation extracts only more posturing: talk is cheap precisely where stakes make honesty expensive. What is needed is a device making the revelation of one’s true valuation the best move — the honest answer without asking anyone to be honest. That device exists; it is called a price; designing the rules that make prices tell the truth is the next chapter’s subject.
11.7 Summary
- Negotiation is joint decision-making under mixed motives. An agent’s true power is its best alternative to a deal — an agent that cannot credibly walk away is not negotiating but requesting.
- Interests and claims are different disagreements. Bargaining resolves conflicts over what agents want; argumentation over what they believe — and self-interested agents will dress the one as the other.
- Bargaining theory unites fairness and strategy. Nash’s axioms pick out a unique split; Rubinstein’s alternating offers converge on the same answer; and the strategic version prices patience, deadlines, and outside options in a currency token budgets make literal.
- Automated negotiation is protocol plus strategy. The rules of the encounter decide whether honesty and efficiency are individually rational — a lesson pointing at mechanism design; a negotiating agent is only as trustworthy as the protocol disciplining it.
- Argument has a formal theory, and debate is its newer clothes. In Dung’s frameworks a claim’s standing is computed from the graph of attacks, not from eloquence; the critic’s objection is an attack edge, and the revision that survives review survived its attackers.
- The evidence on debate is genuinely mixed, and the jury theorem explains why. Deliberation pools information and correlates errors in the same breath — debate helps where pooling dominates, hurts where fluent confidence herds the panel; and where interests are fixed and resources must simply be divided, the honest instrument is not conversation but a price — Chapter 12.
11.8 Exercises
Exercise 1. The orchestrator needs a schema migration done. Doing it in-house would cost 42,000 tokens of the team’s budget; an external migration service can deliver it at a true cost to itself of 30,000 tokens; the two negotiate a price, in tokens, for the job. (a) State each party’s reservation value and what fixes it, the zone of possible agreement, and the surplus that agreement creates; then, taking both utilities linear in tokens and the situation otherwise symmetric, show that the Nash bargaining solution of Section 11.2 prices the deal at the midpoint of the zone, and compute the price and each side’s gain. (b) Before talks open, the orchestrator could spend 2,000 tokens on a fallback spike that would cut the in-house cost to 32,000. Recompute the predicted price and the orchestrator’s total outlay, and compare with spending the same 2,000 tokens polishing the arguments it will make at the table; then rework the sums for a stronger spike that cuts the in-house cost to 29,000 — what happens to the zone, what does the orchestrator’s best outcome now look like, and why does Section 11.1 count this dead negotiation as the machinery working? (c) The service opens with “our floor is 40,000; we could not possibly go lower.” Compute what the claim is worth to the service if the orchestrator credits it, explain why it is cheap talk in the sense of Section 9.5, and say what makes the spike-holding orchestrator’s threat to walk away credible in the sense of Section 9.3 when the service’s floor-claim is not.
Exercise 2. The orchestrator (party 1) and an external review service (party 2) bargain by alternating offers over a 12,000-token surplus under the discounting of Section 11.2.1: a round of haggling costs the orchestrator 20% of the remaining value (\delta_1 = 0.8), while the service, whose reserved compute slot is bleeding away, loses half (\delta_2 = 0.5). (a) Derive x^{*} = (1 - \delta_2)/(1 - \delta_1 \delta_2) from the stationarity pair — each proposer offers its counterpart exactly the discounted value of what that counterpart would take as next round’s proposer — and confirm by differentiation the chapter’s comparative statics: the proposer’s share rises with \delta_1 and falls with \delta_2. (b) Compute both parties’ token takes when the orchestrator proposes first and when the service does; separate what patience buys from what proposing first buys. (c) With equal patience \delta_1 = \delta_2 = 0.8, compute the proposer’s take, and say what happens as \delta \rightarrow 1 and which reunion of Section 11.2 that limit is. (d) A hard sprint deadline drops the orchestrator’s effective discount factor to 0.4: recompute its take as proposer and price the shortened leash in tokens; then explain what the loss presupposes about who can see the deadline, and what re-enters the model — and why — once discount rates are private rather than common knowledge.
Exercise 3. The team’s purchasing agent (the buyer) and an external GPU-broker agent (the seller) negotiate the price of a fine-tuning slot: the seller’s floor is 30,000, the buyer’s ceiling 42,000, their opening books 48,000 and 24,000, and the deadline T = 20 rounds. Each round both post simultaneous offers from the time-dependent family of Section 11.3, conceding from ideal to reservation along \alpha(t) = (t/T)^{1/\beta}; the deal executes in the first round the buyer’s offer at least meets the seller’s, at the mean of the two standing offers. Take Boulware \beta = \frac{1}{4} and conceder \beta = 4, the profiles of Figure 11.2. (a) Implement the exchange and tabulate the deal round and price for all four pairings of the two profiles. (b) Explain the table’s pattern — who captures the zone against whom, and what the symmetric pairings pay for their symmetry — and state which lesson of Section 11.2 the mixed pairings re-derive without an equilibrium in sight. (c) Now charge each party 150 tokens per round of haggling and recompute the four pairings’ net surpluses; identify the party that ends worse off than never having come to the table, explain how a tactic blind to its own burn rate managed that, and name the Faratin tactic family whose job is to prevent exactly this. (d) Halve the buyer’s deadline to T_b = 10 — its curve anchors there, so it offers its full reservation from round 10 on — leaving the Boulware seller at T = 20: recompute the Boulware–Boulware pairing, and price Section 11.2’s warning that a visible deadline is handing the other side its knife.
Exercise 4. The review thread of Section 11.4.1 grows two limbs. To \mathcal{A} = \{P, R, B, T\} with \mathcal{R} = \{(R, P), (B, R), (T, B)\}, add the tester’s second objection S: “the benchmark shows a 20% latency regression” — an attack on P; the coder’s reply N: “the benchmark host was contaminated; the regression is an artefact” — an attack on S; and the tester’s counter M: “the contamination report used the wrong manifest” — where N and M attack each other, each side calling the other’s evidence the artefact. (a) Iterate the defence function \mathcal{D} from the empty set by hand, one line per step, to the grounded extension; classify every argument as accepted, rejected (attacked by the extension), or left undecided. (b) Find every preferred extension, justifying maximality, and classify each argument as sceptically accepted, credulously accepted only, or in no admissible set; state the merge verdict under each standard, and explain why the two standards agree about P but part over S. (c) Extend the companion module foundations/algorithms/argumentation.py — whose docstring leaves preferred extensions “deliberately” to this chapter’s exercise — with preferred(A, R_attacks), built by subset enumeration over its is_admissible test; check it against (a), (b), and the four-node chain of Section 11.4.1, and say why brute force is forgivable at seven arguments and doomed at sixty. (d) Prove what your code exhibits: show that \mathcal{D} is monotone, deduce by induction that every fixed point of \mathcal{D} contains \mathcal{D}^{k}(\emptyset) for all k, and conclude that the grounded extension is contained in every preferred extension — taking as given Dung’s fundamental lemma, that an admissible set which defends a stays admissible when a joins, from which a maximal admissible set is a fixed point of \mathcal{D}. (e) The tester reruns the benchmark on a fresh host from a pinned manifest and reproduces the regression — a new argument V attacking N, earned by execution in the sense of Section 11.4. Recompute both semantics, describe what the earned attack bought, and show what happens if the coder’s costless reply “the rerun is contaminated too” is admitted as an attack on V — and what that says about letting arrows into the graph for the price of asserting them.
Exercise 5. On the coder’s stream-parser patch the reviewer raises three objections: O_1, “this parser is too convoluted; split it before it merges”; O_2, “this will crash on an empty batch — the loop body assumes at least one record”; and O_3, “the variable names are misleading”. The coder renames the variables, adds a source comment reading “complexity acknowledged, refactor deferred”, and answers O_2 with “the scheduler guarantees batches are non-empty, so that case cannot arise”. The reviewer replies “the naming and complexity points are addressed — approving”; the patch merges; a fortnight later production crashes on an empty batch during a queue drain. (a) Write the exchange as an argumentation framework twice — once as the reviewer acted on it, once honestly — marking every attack edge earned or merely asserted (Section 11.4); compute the grounded extension of each, show that the first endorses shipping and the second does not, and say precisely which two edits turned the second graph into the first. (b) Diagnose the drift Section 11.5 warns of: which moves were concessions traded rather than attacks answered, what a commitment store would have shown against O_1 and O_2 at the moment of approval, and what the stopping rule “the critic runs dry” actually measured in this transcript. (c) Redesign the loop so the failure cannot recur: say which of O_1–O_3 can be earned into evidence-backed attacks at all, and by what exhibit; give the ship gate as a property of the framework; give a stopping rule that is neither exhaustion nor agreement; add one guard against a placated critic; and name the row of Table 11.2 your rebuilt loop now most resembles, defending the choice.
Exercise 6. Price the audit of Section 11.6. A committed answer costs 1,000 tokens; a solo pass is correct with probability 0.7, errors independent across passes; the run’s budget is 9,000 tokens. The silent discipline buys nine samples and a majority vote; the debate buys three debaters for three rounds at 1,000 tokens per debater per round (answers are short; take the round cost as flat). Model the debated verdict as the chapter describes it: with probability h the panel herds on its most confident member, correct with probability 0.6 — confidence being no calibration — and otherwise the three vote independently, each lifted by the exchange to accuracy p'. (a) Compute the silent baseline P_9 exactly. (b) For modest pooling, p' = 0.78 with h = 0.3: compute the debate’s accuracy; then set h = 0 and show the debate still loses — name the finding of the chapter’s case against that this arithmetic reproduces, and say what the committee turns out to be. (c) For strong pooling, p' = 0.9: compute the accuracy at h = 0.3, then find the herding threshold h^{*} below which the debate beats the silent baseline. (d) Show that at h = 0.3 no p' \le 1 rescues the panel, and compute the herding ceiling at which even debaters made infallible by the exchange break even; then say which two of the chapter’s closing design rules these two thresholds price, and why “should my agents debate?” is, on this arithmetic, a question about a regime rather than a technique.
Further exercises for this chapter continue in the web edition’s exercise bank.
