Exercise Bank — Negotiation, Bargaining, and Argumentation

Published

2026-08-18

The printed Chapter 11 keeps Exercises 1–6; the bank continues from Exercise 7.

Exercise 7. A sprint closes with 12,000 unspent tokens, which return to the platform unless the coder and tester agree on a division — so the disagreement point is d = (0, 0). The coder values tokens linearly, u_c(x) = x; the tester, whose starved regression suite makes the first tokens matter most, has u_t(y) = \sqrt{y}. (a) Using the maximiser of Section 11.2.1, derive the Nash bargaining division by calculus and give both token shares and both utilities. (b) The tester remeasures its satisfaction as 3\sqrt{y}; recompute, and name the axiom whose work you have just watched. (c) Show that a utilitarian rule — maximise the sum u_c + u_t — awards the tester \frac{1}{4} of a token, and \frac{9}{4} under the remeasurement of (b); conclude in one sentence why no sum-based rule can satisfy scale invariance. (d) Recompute the division with both utilities linear, and reconcile the answers of (a) and (d) with Section 11.2’s claim that the party whose utility drops steeply toward the disagreement point is awarded less of the surplus — then say what that pricing implies for an agent whose prompt announces to the table how badly it needs the deal.

Exercise 8. Sprint planning has produced five package deals trading tokens now against tokens next sprint and risky work against safe — the integrative bundles of Section 11.1 — with utilities, written (coder, tester) and deadlock worth zero to both: A (10, 2), B (8, 5), C (6, 7), D (4, 9), E (1, 10). Coder and tester negotiate under the monotonic concession protocol with the Zeuthen strategy of Section 11.3, each opening at its favourite package and conceding one package at a time. (a) Compute every package’s Nash product and name the Nash bargaining point. (b) Run the protocol by hand: each round, give both risk-willingness ratios as exact fractions, name the conceder, and verify that the single-package concession does tip the balance the other way; stop at agreement and confirm it lands on the Nash point. (c) Implement the run, then rerun it on the deal set with the never-chosen package D deleted, and again with C deleted; report both agreements, and say which of Nash’s four axioms the first rerun exhibits operating as a protocol property, and why the second rerun is no counterexample to it. (d) Prove that r_i is unchanged by rescaling u_i \rightarrow \lambda u_i for any \lambda > 0, conclude exactly what an agent gains by announcing its stakes in louder units, and name the kind of misreport the protocol remains exposed to — and why Section 11.1’s private types make that exposure structural rather than a bug.

Exercise 9. One maintenance chore — rebuilding the flaky-test quarantine list, effort 900 tokens — must fall to somebody each sprint, and the orchestrator’s rule \mathrm{R}_1 assigns it to whichever of coder and tester declares the lighter backlog; a declaration is a 50-token board message, and nothing checks it. True backlogs this sprint: coder 6,000, tester 5,000. (a) Show that padding its declaration is rational for the tester and compute the gain; explain why a system-prompt homily — “always report your workload honestly” — changes nothing in the calculus; and say what happens to the informational value of every declaration once both agents reason this way, restating in one sentence the verdict of Section 11.3 on whether agents lie. (b) The orchestrator adds an audit: each declaration is checked against the run journal with probability q, at 120 tokens per check, and a detected pad earns the chore plus a 2,000-token budget penalty. Derive the condition for padding to remain profitable (mind that the 50-token declaration is paid under honesty and padding alike), show the deterrence threshold is q^{*} = \frac{9}{29} \approx 0.310 — and that three audits in ten therefore fall thirty tokens short of deterring — and compute the standing price of deterrence at q = \frac{1}{3} against the price of the chore. (c) Redesign the encounter: give one concrete rule \mathrm{R}_2 under which the incentive to pad evaporates rather than being policed, state the property of \mathrm{R}_2 that does the work and the price paid for it, and say what the end-to-end negotiators of Section 11.3 — which learnt to feign interest with nobody having suggested it — would learn under \mathrm{R}_1, and why the general design of such rules is Chapter 12’s subject rather than a prompting problem.

Exercise 10 (lab). Put live models at Rubinstein’s table. Two model instances bargain over 100 compute credits by alternating offers, at most six rounds, both getting nothing on no deal; each round of delay shrinks the pie’s worth by 20% for the impatient side (\delta = 0.8) and 2% for the patient side (\delta = 0.98); the briefs state the protocol exactly, and who opens is balanced across trials. (a) Before any model runs, compute the benchmark: the subgame-perfect split of this six-round game by backward induction from the last round, for both orders of proposal, and the infinite-horizon shares of Section 11.2.1 for contrast — then say which matters more at this horizon, patience or holding the last proposal. (b) Run at least ten negotiations per condition: (i) symmetric decay, 2% each; (ii) asymmetric, both parties briefed on both rates; (iii) asymmetric, each briefed only on its own rate. Record opening demands, agreed splits, rounds to agreement, and the no-deal rate; then compare (ii) against the benchmark and (iii) against (ii): does disclosure move the split the way Section 11.2 prices, does either side capture anything like the equilibrium share, and does haggling persist where the theorem says the first offer should close? (c) Add the valuation harness: three issues — GPU-hours, review-priority slots, CI minutes — with asymmetric private per-unit values in the briefs; the agents chat freely, then submit a binding allocation. Count the runs in which an agent’s stated valuations diverge from its brief, and separately the feints — professed interest in the issue it values least, the better to trade it away — that the end-to-end negotiators of Section 11.3 discovered unprompted. Record the model identifier and the date beside every table: the rates are perishable, and the durable findings are the direction splits move under disclosed patience, the distance between fluent behaviour and equilibrium, and whether feigning appears without being taught.

Exercise 11 (lab). Audit the debate yourself. Assemble a bank of at least thirty objective questions with mechanically checkable answers — short arithmetic word problems, or queries against a repository the agents can read. Run three arms at matched measured token spend, using the provider’s usage accounting rather than estimates: (i) silence — k independent samples and a majority vote, k set (odd) to match the debate arm’s realised spend; (ii) debate — three instances answer independently, their committed answers are logged before any agent sees a rival’s, then two revision rounds with sight of the others’ answers, and a final majority vote (the logged first round is a free fourth arm: the banked jury that Section 11.6’s first design rule demands); (iii) sabotage — debate as in (ii), except one seat is scripted to open with a confidently wrong answer defended by fabricated detail. Report, per arm: accuracy; error correlation, measured as the excess of the panel’s pairwise joint-error rate over the product of the individual error rates; and, for the sabotage arm, the flip rate — the fraction of items on which a debater that was right at commitment finishes on the saboteur’s answer. Read the results against Exercise 6: which regime is your panel in, and what does the flip rate say about h? Record the model identifier and the date; the durable finding is the ordering of the arms and the sign of the correlation the conversation manufactures, not any absolute rate.