23 Collective Intelligence and Open Problems
The book has been building, by degrees, a small society: agents that reason and act, speak and listen, plan together, coordinate without colliding, vote, bargain, trade, organise, learn, survive the night, withstand attack, and answer for themselves afterwards. Twenty-six chapters of machinery, and one question deferred at every step. It is later. The question is whether the assembled whole is ever more intelligent than its parts — not more productive, which is bookkeeping, nor more impressive, which is marketing, but better at the thing intelligence is for: reaching truer conclusions, better decisions, and sounder designs than any member could have reached alone. The preface called a committee of language models still a committee, and demanded minutes; the final chapter must now say when a committee is worth convening at all.
Two answers are on offer, and both are wrong in instructive ways. The enthusiast’s — ten agents are ten times the mind, a swarm a superintelligence with better branding — fails against the honest ledger: multiplied tokens for the same result, teams that fail at the joins, and the standing discovery that many multi-agent systems would be faster, cheaper, and sturdier as one agent with a good tool. The sceptic’s — the whole is always less than its parts — fails against crowds that outguess their experts, juries provably wiser than their jurors, and parallel teams that beat strong single agents on tasks that genuinely divide. The truth has conditions attached; the conditions are checkable; and most of this book has been, without always saying so, a manual for engineering them.
Behind the engineering question stands an older one. The field’s founding intuition is that intelligence was never a solo phenomenon to begin with: a mind is itself a society of lesser agents, cognition in the wild is distributed across people and instruments, and the human intelligence now building these systems is the product of a collective process no individual brain could have run. If that is right, the era’s question — one ever-larger mind, or a better-organised many? — is a question about composition, and composition is precisely what this book has a theory of. What the book attempted and did not achieve is then handed on the way a well-run team closes a sprint: written down honestly and delegated, with a scoped brief and the accounting kept, to the only agent the book has left — the reader.
23.1 Is the Whole Smarter? The Conditions of Collective Intelligence
A term this attractive needs a definition before the marketing department gets to it. Say that a group exhibits collective intelligence when its performance on cognitive tasks reliably exceeds what its best member could have achieved alone — reliably, because any committee can beat its best member on a lucky Tuesday, and best member, because beating the average is too cheap a bar: a group that merely dilutes its strongest mind with company has arranged an expensive way to be worse. The definition is this book’s operational choice, and it bites only under a fair comparison: the best member must stand alone with the same model, tools, and token budget the team enjoyed, or the team’s margin measures its allowance rather than its interaction (Chapter 21’s like-for-like discipline). Chapter 10 proved the narrow case: a majority of competent jurors whose errors are independent beats any one juror, the whole result balancing on the independence clause. But a working team is more than a jury convened at the end — it decomposes, allocates, works in parallel, checks, argues, and only then aggregates — and collective intelligence, if the term is to earn its keep, is a property of that entire loop: the jury theorem’s fine print, scaled up from the verdict to the whole enterprise.
Something in the vicinity can be measured. Woolley and colleagues put small human groups through varied task batteries and found a single statistical factor — a group’s general ability — that predicts performance across tasks (2010): only weakly predicted by the members’ average or smartest intelligence, the two quantities a recruiter would buy, and predicted instead by the interaction — social sensitivity, and how evenly the conversational turns were distributed. Replications dispute the c factor’s size and predictors, but the structural point — the interaction carries signal the members’ test scores do not — has weathered better than the coefficients, and every quantity in it is a design surface: who speaks, when, with what right of reply is Chapter 18’s control lens, and it appears to be where much of a group’s intelligence lives. The most consequential instance is the smallest mixed collective, one person and one agent, and the record there is usefully deflating: complementary performance, the mixed team beating both members alone, is exactly this section’s bar, and the early human–AI studies cleared it far less often than their transcripts suggested (Bansal et al., 2021). Chapter 22’s automation bias as a coefficient, and Section 21.4’s methods as the entry fee: the mixed team is measured as a unit or it is not measured at all.
Diversity’s claim comes with a famous theorem, and the theorem with a health warning. Hong and Page showed conditions under which a diverse collection of problem solvers outperforms a collection of the individually best (2004): the best solvers, selected on a common yardstick, tend to be similar, so they stall at the same local optima — a team of them is one search, repeated — while a diverse group stalls in different places, each member’s dead end somewhere another can restart from. The slogan it spawned, that diversity trumps ability, promptly outran the mathematics. The durable residue: diversity is a second, independent axis of group design, not a substitute for competence — the ensemble literature’s identical verdict, accurate and diverse, wrong in different places — and its value is precisely as large as the difference between the members’ blind spots.
Assembled, the conditions for collective intelligence read like a pre-flight check, each item with a chapter behind it. Competence: every contributor better than chance at its station, a measurable, per-domain quantity — Chapter 21 measures it, Chapter 22 calibrates it. Diversity: members wrong in different places, so that one agent’s blind spot falls within another’s field of view. Independence: the errors kept uncorrelated at the moment of judgement, which is a property of process — of who saw whose answer before committing to their own. Aggregation: a combining rule matched to the species of question, majority and median for facts, fairness-aware rules for preferences, with Chapter 10’s impossibility taxes paid knowingly. And two conditions the jury frame is too small to show. Decomposability: the task must actually divide, or the perspectives complement — Chapter 1’s first question, still the first question. Verification: cheap, trustworthy checking at the joins — Chapter 21’s harness — because diversity without verification merely produces disagreement, and disagreement is raw material, not intelligence. Six conditions; not one of them is headcount.
For teams of language-model agents, the checklist has a distinctive failure mode, named in Section 20.9 with a security bill attached: samples from one model share its training wholesale, and a thousand such jurors are one juror, photocopied. Turning up the temperature buys surface variety, not independence — the paraphrases differ, the blind spots are identical, since the questions a model gets confidently wrong are not the ones its sampler dithers over. Prompting one model to role-play five experts convenes five costumes around one mind. Real decorrelation has to be bought where the correlation lives — Section 20.9’s procurement list: different base models; different evidence — a different window on the problem, a different tool, a different slice of the corpus; different methods, enforced by role and harness rather than requested by adjective. And independence, once bought, must be protected by process: separate contexts, verdicts committed blind, aggregation before deliberation — letting the jury talk pools its information and correlates its errors in the same breath. None of this is free, which is rather the point: in this substrate, diversity is a procurement decision, and a team’s error correlation is a measurable quantity sitting in the journal — Chapter 21 can estimate it from disagreement rates — rather than a hope. Figure 23.1 plots what the correlation costs: the whole wisdom-of-crowds gain, draining to a single mind as errors align.
What does the ledger show when the conditions are respected? The most fully priced result on the public record has the right shape: a parallel, orchestrator-led team beat a strong single agent decisively on research tasks that divide into independent facets — and cost many times the tokens, and showed no advantage on tasks that fit in one context (Anthropic, 2025): Chapter 21’s canonical triple, and a clean instance of every condition above. Against it stands the failure census: systems that collapse not because members lacked competence but at the joins — which is to say, at the conditions (Cemri et al., 2025). Read together: collective intelligence is not what you get by convening agents; it is what you get by convening them correctly, and “correctly” has technical content — six conditions, each checkable, each buildable, none automatic. The whole can be smarter than its parts. Whether it goes on getting smarter as the parts multiply is the next question, and the answer will want a graph rather than a slogan.
23.2 More Agents, More Wire: How Capability and Risk Scale
The intuition this section exists to retire is additive: if one agent is useful, ten are ten times as useful. The refutation was published in 1975. Brooks’s The Mythical Man-Month observed that adding people to a late software project makes it later, and located the mechanism in the wiring: n contributors hold n(n-1)/2 potential channels, so the coordination burden grows quadratically while the hands grow linearly (1975). Men and months are interchangeable only when a task partitions perfectly among workers who need never speak — almost never — so the “man-month” is mythical exactly to the extent that the task is coupled: Section 23.1’s decomposability condition arriving from a different century. Substituting agents for programmers changes none of the structure and sharpens the accounting, because agents communicate in tokens and tokens are the budget — Chapter 8’s warning about the team that spends its whole allowance staying in step is Brooks’s law with a meter attached. The man-month was mythical; the agent-token is billable.
So capability does not scale with population; it scales with the task’s parallelism, and population merely sets a ceiling on how much of it can be harvested. Where the work divides into independent facets, gains are real and can approach linear up to the number of true subtasks — the research-synthesis win sits here (Anthropic, 2025). Where the work couples, the curve flattens fast: each additional agent brings less uncovered work and more wire, and the orchestrator’s context becomes the serial bottleneck. Past the flat maximum the returns turn negative, because the errors that kill multi-agent systems live at the interfaces — specification, hand-off, verification — and interfaces are exactly what a marginal agent adds (Cemri et al., 2025). An agent added to a coupled task contributes its work minus its wiring, and the subtraction changes sign quietly, with no alarm beyond a bill growing faster than a benchmark score (Figure 23.3). Diminishing returns is the good outcome; the bad one is a team that gets worse while getting bigger and dearer, which the joins data says is not a hypothetical.
Risk, meanwhile, runs up a different axis: capability scaling needs the task’s cooperation; risk scaling needs only the wiring. Add a channel between two agents and you acquire its dangers unconditionally. The channel is attack surface: Section 20.9 found that plurality multiplies every ingress point. It is a contagion path: the poisoned paragraph and the fashionable error propagate along the very connections installed for coordination, and Chapter 15’s cascade arithmetic showed how quickly a densely connected population converts two early mistakes into a thousand-agent consensus. And it is a correlation engine: agents that share a model, a context, or a retrieval corpus fail together, so the more thoroughly a population is knitted, the less its redundancy is worth — replicas on one substrate are not replicas. Connectivity taxes the very independence Section 23.1 identified as collective intelligence’s scarcest input. The wire carries the signal and the disease at the same rate, and only the signal needs an invitation.
The lever that sets both variables at once has been in the book’s hands since Chapter 8: topology. A topology, that chapter said, is a placement of coordination cost; the risk ledger adds that it is equally a placement of exposure. The star bounds contagion — spokes that cannot address one another cannot infect one another — and concentrates both bottleneck and blast radius at the hub. The peer mesh distributes the load and maximises the channels, buying resilience to any one failure at the price of a superb propagation fabric for the correlated kind. The hierarchy bounds the wiring per node and inserts firebreaks, at the usual cost in delay and rigidity. None of this changes Chapter 8’s conclusion; it doubles it: the cheapest dependency is still the one designed out. The practical rule is an audit discipline — every arrow is a token bill (Chapter 8), an attack path (Section 20.9), and a contagion channel (Chapter 15), so every arrow should correspond to a dependency the task actually has. Wiring installed because connection feels like collaboration is paying three taxes to deliver a feeling.
A book that has spent a chapter insisting on evidence over vibes (Chapter 21) owes its own central claims the same treatment, so the companion repository ends on an experiment. The capstone lab runs one batch of repository issues of graded difficulty across a grid — team sizes one to eight; single agent with good tools, star, chain, blackboard, and peer mesh — under the same token accounting, against the strong single-agent baseline. Four curves come out — quality, tokens, latency, and the failure-mode census — and the chapter’s hypotheses are exposed to refutation: a flat maximum at small team sizes for coupled tasks and later for parallel ones; cost rising faster than quality past it; failures migrating from competence errors to join errors as the population grows. The curves may land elsewhere — tasks differ, models improve, the flat maximum moves — and that would not embarrass the lab but vindicate it, because the portable finding is the method: the scaling question is empirical, per task, and cheap to answer with a harness you already own. Scaling by assumption is the one configuration the lab forbids.
What the two ledgers say together is easily compressed: more is not smarter; structured is smarter. The intelligence of a collective lives in its organisation, its independence, and its joins, all of which degrade with careless growth and none of which headcount supplies; small, well-wired, well-verified teams keep winning the honest evaluations because Section 23.1’s conditions are easier to hold at four agents than at forty. Which leaves the question waiting behind the engineering all along: if piling on agents is not the road to more intelligence, what is a collective’s contribution to intelligence in the first place — and might the answer explain where intelligence, including the reader’s, came from?
23.3 Societies of Mind: What Collectives Contribute to Intelligence
Chapter 1 opened this book with a wager borrowed from Minsky: that intelligence is not a possession but an organisation — the society of mind, in which a single mind is itself a society of mindless components and the cleverness lives in the arrangement (1986). The preceding two sections are the assessment: what separated a collective that thinks from an expensive committee was never the parts — it was the conditions, the topology, the joins, the aggregation rule; organisation, at every turn. And if intelligence is organisation, there is no privileged scale at which it must live. The evidence, assembled at three scales in Table 23.1, is that it never did live at the scale we flatter ourselves it does.
| Scale | Framework | Where the intelligence lives | Lesson for agent teams |
|---|---|---|---|
| One mind | Society of mind (Minsky, 1986) | In the arrangement of mindless components | Agent teams make the picture a literal engineering reality |
| One crew (a warship’s bridge) | Distributed cognition (Hutchins, 1995) | In the system — people, instruments, and procedure together | The journal, schema, protocol, and verifier are constituents of the system’s cognition, not accessories to it |
| One species | Cumulative culture / the collective brain (Henrich, 2016) | In innovations retained, transmitted, and recombined across generations | Agents still have no cross-run cumulative culture; building its full loop — the test and the ledger included — is open |
The first piece of evidence is a warship entering San Diego harbour. Hutchins, an anthropologist embedded on the navigation bridge, asked who is doing the navigating, and found no answer at the level of an individual survives the data (1995): the fix that lands on the chart is computed by a system — bearing-takers, plotter, instruments, chart, and a choreography of who reports what to whom — whose cognitive properties differ from any member’s. He called the framework distributed cognition, and its unit-of-analysis lesson transfers without modification: asking which agent in a well-built team is the intelligent one is the same category error as asking which sailor is the navigation. The reader has already built Hutchins’s bridge — the journal that remembers, the schema that disciplines, the protocol that choreographs, the verifier that checks are not accessories to the agents’ intelligence but constituents of the system’s.
The second piece of evidence is the reader. Henrich assembled the case that human intelligence is itself the product of a collective process (2016): stripped of their culture, individual humans are unimpressive apes — exploration history is punctuated by well-equipped expeditions starving in landscapes where local populations had flourished for millennia, because the local food required preparation knowledge no individual could re-derive on the spot. What fills the gap is cumulative culture: innovations retained, transmitted, recombined, and improved across generations, so that each mind starts where the last generation finished rather than where the species did. Henrich’s name for the engine is the collective brain, and its power grows with the size and interconnectedness of the group. Our intelligence, on this account, is not what built the collective; it is what the collective built. The existence proof for composition creating intelligence is the species reading this sentence.
But Henrich’s scaling claim and Section 23.2’s appear to point in opposite directions — connectivity as the collective brain’s fuel there, a triple-taxed liability here — and resolving the tension yields this section’s most practical insight. Cultural transmission is not a mesh forwarding everything; it is ferociously selective: learners copy the successful, teachers filter, failed practices die, and institutions retain what survived. Culture, in this book’s vocabulary, is connectivity plus verification plus memory — the wire, the test, and the ledger — and stripped of either of the latter two it degenerates into exactly the cascade machinery Chapter 15 warned of: fashion is what cumulative culture looks like without the test. This names what today’s agent systems are still missing. Within a run, the book has built the organs: the journal remembers, the verifier selects. Across runs, teams, and organisations there is as yet no cumulative culture of agents worth the name — each deployment re-derives its working practices, and what one team’s expensive failure taught is retained, if at all, in its humans. Chapter 16’s asset ledger and Chapter 20’s incident suite are first artefacts of such a culture; the open question — it reappears in the next section — is whether agent collectives can be given the full loop, selective transmission included, without also being given the fashions.
All of which allows the era’s loudest question — does the road to more general intelligence run through scaling one mind or composing many? — to be put with discipline, and three things said at textbook temperature. First, the alternatives are complements, not rivals: every composed system improves when its members do, and every large mind is deployed, in practice, into a system of tools, checks, and colleagues. Second, the one instance we possess of general intelligence actually arising — ours — arose by the collective route, with individual brains as the product of the process rather than its designer; an existence proof is not a roadmap, but it is a standing rebuke to the assumption that composition is the derivative programme. Third, composition supplies what scale alone does not, and what the later chapters argued is non-negotiable in deployment: division of cognitive labour, independent verification, auditability of the joins, institutions that keep capability answerable. What multi-agent systems contribute to the pursuit of general intelligence is not a shortcut and not a rival bet; it is the science of organisation that any outcome will need on the day it is deployed — because the destination will not be a mind in a vat. It will be a participant in a collective that includes us, and the collective is where intelligence has always lived. What we do not yet know about building such collectives well is, fittingly, the longest list in the book — and it is next.
23.4 The Open Problems
A place on this list has to be earned twice over: each problem below is open — unsolved by the classical tradition, the modern practice, and every chapter of this book — and each is sharpened, given a form precise enough to be attacked rather than merely admired. The provenance is the text itself: these are the places where the preceding chapters had to write “at the time of writing” or “nobody yet knows”, and the reader who felt the hedges accumulating may treat this section as their consolidated statement of account. Two warnings. Lists of open problems date in both directions — solved embarrassingly soon, or dissolved into questions never well posed — and both outcomes retire an entry honourably; what should outlast the list is the habit of stating problems sharply enough to suffer either fate. And the six headings, inherited from the preface’s promise, are a filing system rather than a taxonomy: the best problems live at the seams between them.
Theory (Table 23.2). The book’s formal machinery runs on idealisations the substrate violates daily, and the violations are now the interesting objects: a strategic theory of the player we actually deploy rather than the player game theory models; a compositional theory of team capability — the field’s own Amdahl’s law, so that Chapter 1’s opening question eventually gets an answer that is a calculation rather than a discovery made at four times the token bill; and the mathematics of juries as they actually convene.
Learning (Table 23.3). The substrate is periodically replaced by its provider, so every coordination habit a team has learned is an asset of uncertain survivorship; credit assignment grows a new face when the actions are paragraphs — needed forwards for training, backwards for Chapter 22’s accountability; and the efficiency–auditability frontier is real and unmapped. Section 23.3’s cheque falls due at L4: the cumulative culture problem — the full loop, test and ledger included, without the fashions — may be the most consequential entry on the page.
Engineering (Table 23.4). Three problems, each the residue of a chapter that did everything it could with what exists: end-to-end attenuation, which Chapter 19 named the thinnest treaty in the consolidation; a science of verification for open-ended outputs — for prose, designs, and judgement calls the ground-truth ladder of Chapter 21 runs out at a judge model holding an office Chapter 10 proved troublesome; and the regression science of substitution — when a member’s brain is swapped overnight, the prompts, the learned coordination, and Chapter 22’s trust curves are invalidated together, and re-certification is, at the time of writing, folklore plus vigilance.
Economics (Table 23.5). Mechanism design earned its guarantees against solitary deviation, and Chapter 12 was candid about the cartel-shaped hole; it sharpens viciously when the bidders share a language to scheme in and possibly a substrate to think with. Behind collusion stand two questions the book could only gesture at: the pricing of thought, and the missing commercial layer between strangers’ agents — Chapter 13’s institutions rebuilt for an open web; until it exists, every cross-boundary delegation is a handshake without a court.
Safety (Table 23.6). The composition problem is Section 20.9’s residue stated as a research programme: every agent passes its evaluations and the ensemble still colludes, cascades, or drifts, because alignment is certified per agent and needed per system. Scalable oversight — machine-speed work reviewed by human-speed judgement — keeps its standing as the field’s hardest supervision problem. And substrate monoculture wants its systemic theory: an economy of agents on a handful of model families is a correlated-failure machine of exactly the kind financial regulators learned to fear, and its contagion mathematics has yet to be written.
Governance (Table 23.7). Chapter 22 built answerability inside one organisation, and the open problems begin at its property line: the chain of authority that crosses companies — your agent, calling my tool, running on their model — currently has no owner, in law or in protocol; institutions at machine speed are a genuinely new design problem, not a faster copy of an old one; and the oversight regress asks where supervision grounds when the auditors are acquiring agents too — Chapter 22’s chain-terminates-at-a-named-human collides with the volume that made delegation attractive, and reconciling the two is nobody’s solved problem.
| ID | Open problem | What remains open | Rooted in |
|---|---|---|---|
| T1 | A solution concept for approximately rational agents | A strategic theory of approximately rational, language-conditioned agents, with equilibria robust to paraphrase | Chapter 9 |
| T2 | A compositional theory of team capability | Deriving optimal team size and topology from a task’s dependency structure — the field’s own Amdahl’s law | Chapter 21, Section 23.2, Chapter 1 |
| T3 | Juries as they actually convene | The mathematics of jurors partially correlated, miscalibrated, and trained to persuade | Chapter 10 |
| ID | Open problem | What remains open | Rooted in |
|---|---|---|---|
| L1 | Transfer of collective competence across substrate change | Preserving learned coordination when the provider replaces the underlying model | Chapter 14 |
| L2 | Credit assignment at transcript scale | Which utterance won or lost the task — forwards for training, backwards for accountability | Chapter 14, Chapter 22 |
| L3 | The efficiency–auditability frontier | Mapping the trade between agents’ compressed private codes and human legibility | Chapter 14 |
| L4 | The cumulative-culture problem | Retention, selective transmission, and recombination of what works — across runs, teams, and organisations, without the fashions | Section 23.3, Chapter 15 |
| ID | Open problem | What remains open | Rooted in |
|---|---|---|---|
| E1 | End-to-end attenuation | Authority that provably narrows across autonomous hops spanning different owners | Chapter 19 |
| E2 | Verification of open-ended outputs | A science of verification, with known error bars, for prose, designs, plans, and judgement calls | Chapter 21, Chapter 18, Chapter 10 |
| E3 | The regression science of substitution | A discipline for re-certifying a team after its substrate is swapped overnight | Chapter 22 |
| ID | Open problem | What remains open | Rooted in |
|---|---|---|---|
| Ec1 | Collusion-resistant mechanism design for natural-language agents | Detection and prevention when bidders share a language, and possibly a substrate, to scheme with | Chapter 12 |
| Ec2 | The pricing of thought | Who captures the surplus from delegated work, and what an economy of agent labour equilibrates to | none — only gestured at |
| Ec3 | A commercial layer between strangers’ agents | Binding commitment, escrow, liability, and reputation that travels across platforms | Chapter 19, Chapter 13 |
| ID | Open problem | What remains open | Rooted in |
|---|---|---|---|
| S1 | The composition problem | When safety composes, what each aggregation level must re-prove, and the safety case for a collective | Section 20.9 |
| S2 | Scalable oversight | Reviewing machine-speed, machine-volume work with human-speed judgement | Chapter 11, Chapter 22 |
| S3 | Substrate monoculture | The systemic contagion mathematics of a few model families — stress tests, exposure limits, diversity requirements | none — only gestured at |
| ID | Open problem | What remains open | Rooted in |
|---|---|---|---|
| G1 | The cross-company chain of authority | Provenance federation, liability allocation, and the three questions asked across organisational boundaries | Chapter 22 |
| G2 | Institutions at machine speed | Norms, sanctions, adjudication, and appeal at the violator’s clock, without becoming ungoverned automation | Chapter 22 |
| G3 | The oversight regress | Where supervision grounds when the auditors and judge models are themselves acquiring agents | Chapter 22 |
Nineteen problems, six headings, and a fair question from the reader: which one? The engineer’s criterion is the one this book has trusted throughout — pick the problem whose absence you have personally paid for, because the pain is a proof of consequence and the incident report is a head start on the related work. The researcher’s criterion is visible in the list’s structure: the strongest entries live where two chapters collide, and the parting methodological advice is to hunt at exactly those seams, since a problem with two parents inherits two literatures’ tools. The final exercise accordingly asks not for an implementation but for a research proposal. None of this list will survive the decade intact, and that is the correct fate for such a list; what should survive is what it demonstrates — the field is alive, its difficulties are specifiable, and the distance between a working system and an open problem is, in multi-agent systems, roughly one deployment. What remains is to say what the book believes it has shown.
23.5 Coda: No Longer Premature
The preface stated a thesis plainly: the classical theory of multi-agent systems was not wrong, nor obsolete — it was premature. A verdict is now owed, and the evidence has been the book’s own method: every time the modern practice ran into a wall, a result was found waiting on the other side, usually with the dust of decades on the cover — the supervisor pattern stood revealed as the Contract Net; the tool protocols re-derived the speech acts; the judge model walked into an office Arrow had described to the syllable; the ensemble convened Condorcet’s jury and inherited its fine print; the injected agent re-enacted the confused deputy; and the deployed fleet demanded, on schedule, the norms, institutions, and answerability the theory of agent societies had drafted for occupants it never met. Not every wall — the open problems are the honest remainder — but often enough, and centrally enough, for a verdict with an exact scope: where the book could put a correspondence to work — run the lab, price the assumptions, watch the theorem’s failure mode arrive on cue — the wager has paid out as a finding; where it could only exhibit the correspondence, it stands as a programme, now posed sharply enough to fail. The theory was a map drawn before the territory could be reached. The territory has arrived, the map has been substantially right wherever the book could survey it, and premature is a word the field may now retire.1
The preface’s other image can be retired with it. Two tribes, it said, had sailed past each other in the fog — the theorists with the vocabulary and no substrate, the engineers with the substrate and no vocabulary — and the book was offered as a bridge. But the bridge metaphor undersells what twenty-seven chapters of crossings keep suggesting: there was only ever one landmass — one field with a thirty-year gap between its specification and its hardware. The classical results discipline the modern practice, as advertised; what the theory tribe receives in return is the thing it lacked from the beginning: experiments. The canonical texts were written for agents their authors could only specify (Shoham & Leyton-Brown, 2009; Wooldridge, 2009); their successors can now run the specifications — game theory with subjects that actually play, social choice with reproducible juries, institutions whose sanctions execute — and every deployed fleet is, seen from the theory side, the largest multi-agent laboratory ever accidentally assembled. A field reunited with its own experiments does not merely apply its old results; it finds out which of them were wrong, the part of science the classical era had to postpone. That, and not nostalgia, is why the old literature deserves the reader’s attention: it is finally falsifiable.
Which leaves only the ending, and the book proposes to end in the one style it trusts: as a delegation, executed by its own rules. Chapter 7 set the requirements, and they are met. The brief is specified clearly enough to commit to — nineteen problems, six domains, one field; or, for the reader whose ambitions are a system rather than a thesis, simply this: build the smallest team that works, measure it against the strongest loner, keep the journal, and gate what cannot be undone. The commitment must be taken on rather than merely received, and that part was never the book’s to perform. And there is a channel for reporting back: this book is written in public, its margins are open, and beyond them lie the workshops and journals of a field that has spent thirty years being politely early and would be glad of the company. The accounting follows Chapter 22 to the letter: authority is hereby delegated, and answerability stays where it belongs — whatever you build is yours to answer for, and whatever this book got wrong remains, entirely and cheerfully, ours.
The preface asked for a cup of tea and a terminal. If you have followed the whole way, the tea is long cold and the terminal has history worth keeping — so put the kettle on again, and go and build some societies. Mind the conditions. Measure the whole. And do keep the minutes: as the book has been arguing since page one, they were never really minutes at all — they are where the intelligence lives.
23.6 Summary
- Collective intelligence is conditional, not automatic. A group outperforms its best member when its members are diverse, their errors independent, and the aggregation sound — the jury theorem’s fine print. Headcount is not a condition; an ensemble of clones sharing one context is a crowd of one, at chorus prices.
- The conditions are engineering targets. Diversity can be built — different models, prompts, tools, evidence; independence protected — separate contexts, blind judgements, aggregation before discussion; competence and error correlation measured from the journal.
- The whole is smarter only where the task divides and the errors decorrelate. The honest evidence shows real wins on genuinely parallelisable work, at a multiple of the cost, and no gain where one context suffices; teams still fail overwhelmingly at the joins.
- Capability scales sublinearly, risk can scale faster, and the scaling must be measured, never assumed. Coordination overhead grows with the wiring rather than the work — Brooks’s law survives the substitution of agents for programmers — while connectivity breeds correlated failure and cascades; the capstone experiment puts the book’s own central claim to its own standard of evidence.
- Intelligence was collective all along. A mind as a society of agents, cognition distributed across crews and instruments, human capability the dividend of cumulative culture: the existence proof that composition creates intelligence is us — and organising many is what this book has a theory of.
- The remainder is delegated. Nineteen open problems across six domains — each sharpened by a chapter of this book and solved by none — are handed to the reader with a scoped brief and the accounting kept. The classical theory was not wrong, and it is no longer premature; the field it was waiting for is the one you are now equipped to build.
23.7 Exercises
Exercise 1. The running team’s review gate currently draws five verdicts by having the one reviewer model post them in sequence in a single transcript, each verdict able to see those before it, and takes the majority. Six configurations are on the operator’s desk: (i) the current gate as described; (ii) one call in which the reviewer is prompted to role-play five named specialists and return five verdicts; (iii) five separate calls to the same model on the same evidence, each verdict committed before any is shared; (iv) five reviewers on five different base models trained by different hands, same evidence, verdicts blind, majority; (v) five calls to the same model, each given a different evidence window — the diff alone, the diff with the tests, with the specification, with the call sites, with the file’s recent history — blind, majority; (vi) the five reviewers of (iv), but sharing one discussion thread and converging on a consensus verdict, which is what gets reported. (a) For each configuration, say which of Section 23.1‘s six conditions it purchases, which it merely imitates, and which it damages, each verdict grounded in the section’s mechanics in a sentence or two. (b) Order the six by the error correlation \hat{\rho} that a disagreement-rate estimate from the journal would be expected to report, ties permitted, justifying every inequality — and say which configuration makes \hat{\rho} unmeasurable and why that is itself the diagnosis. (c) The journal shows the current gate’s majority verdicts scored 0.78 over the last forty decisions, its strongest member alone 0.81 on the same forty, and the members’ average 0.66: deliver the section’s verdict — does the gate exhibit collective intelligence, which comparison is the binding one and why is the other too cheap a bar, and what does a standard error of roughly \sqrt{0.8 \times 0.2 / 40} \approx 0.063 say about the word reliably, and about what Chapter 21’s paired design would have to establish before the panel’s five-fold token bill is justified?
Exercise 2. The diversity ablation in the companion repository’s frontier/scaling_lab/jury.py decorrelates a seven-juror panel of competence p = 0.7 rung by rung — from one shared context to fully split evidence — and at each rung measures the panel’s pairwise disagreement and estimates the error correlation back from it: exactly the audit Section 23.1 says the journal makes possible. Where Chapter 10‘s Exercise 6 built the correlated jury — a copy mechanism and the design-effect n_{\mathrm{eff}} — and took the correlation as given, this exercise recovers it from what a panel visibly does. (a) Derive the instrument: under the module’s photocopying mixture — with probability \rho a single shared draw is copied to every juror, otherwise all seven vote independently, each correct with probability p — show that the correlation between two jurors’ correctness indicators is exactly \rho, that the expected pairwise disagreement rate is d = 2p(1-p)(1-\rho), and hence that \hat{\rho} = 1 - d/\bigl(2p(1-p)\bigr) recovers the correlation from an observed disagreement rate; then show that forcing n_{\mathrm{eff}} = n/\bigl(1 + (n-1)\rho\bigr) to equal a target m requires \rho = (n-m)/\bigl(m(n-1)\bigr), and evaluate the design rungs for n = 7 at m = 1, 3, 5, 7. (b) Run the ablation (python -m frontier.scaling_lab.run prints it) and check its columns: reproduce the accuracy column analytically from the jury sum P_m(0.7) at m = 1, 3, 5, 7, and set the \hat{\rho} column against the design values from (a), explaining why the estimates land near but not on them — including why the split-evidence rung reports \hat{\rho} \approx 0.003 rather than zero. (c) Price the correlation: compute the wisdom-of-crowds gain at full independence, the fraction of it the panel keeps at the m = 3 and m = 5 rungs as the instrument reads it, and compare each fraction with 1 - \rho at that rung — which way does the inequality run? Then derive, from (a)’s mixture itself, the panel’s exact accuracy \rho p + (1 - \rho) P_n(p) and hence the exactly retained fraction 1 - \rho, and say what that makes of the instrument’s sub-proportional readings — and of any single formula trusted for a correlated panel’s accuracy. (d) Map the ladder onto the section’s procurement menu — temperature, role-play, separate contexts, split evidence, different base models — saying which purchases genuinely move a panel down a rung and which only redecorate the top one, and why the estimator’s input being a disagreement rate makes a team’s diversity an auditable number rather than a hope.
Exercise 3. In the scripted sweep, topology is a uniform factor that — as frontier/scaling_lab/sweep.py’s own comments admit — cannot move a column’s argmax, so every wiring peaks at the same team size, and the claim Section 23.2 most wants tested, that a topology is a placement of coordination cost, is designed out. Put it back: keep p = 0.9, set \kappa = 0.03, price the wiring per size as c_{\mathrm{star}}(n) = 2\kappa(n-1) (a report up and a digest down per worker), c_{\mathrm{chain}}(n) = \kappa(n-1) (one hand-off per link), c_{\mathrm{blackboard}}(n) = \kappa n (one post per member), and c_{\mathrm{peer}}(n) = \kappa n(n-1)/2 (every pair talks), with c(1) = 0 in every wiring, and evaluate S(n) = 1/\bigl((1-p) + p/n + c(n)\bigr). The single-wiring maximisers are already in hand: Chapter 1’s Exercise 2 found the fully connected mesh’s cube-root optimum n^{*} \approx (p/\kappa)^{1/3} and its Exercise 3 the star’s square-root one, so reuse those closed forms and spend this exercise on the placement. (a) Compute the four speedup columns on the lab’s grid, each wiring’s grid flat maximum, and its true maximiser over n = 1, \dots, 100, reporting the exact ties your search must break (the chain and the blackboard each tie two adjacent sizes) and what the lab’s _flat_argmax convention — the smallest size achieving the maximum — says about them. (b) Exhibit the two facts the uniform stand-in could not show: the wirings now peak at different sizes — compare the peer mesh’s true peak with the chain’s — and the mesh alone crosses below the single-agent baseline on the grid; find the crossing size and verify it against the inequality \kappa n(n-1)/2 > p(1 - 1/n), which is Figure 23.2’s caption made computable. (c) Show that the star and the mesh levy identical taxes exactly at n = 4 — solve 2\kappa(n-1) = \kappa n(n-1)/2 — so the running team of four pays the same token bill under either wiring; Exercise 4 reads the rest of the two invoices. (d) In a sentence or two, defend the repository’s choice to ship the argmax-invariant stand-in in a hermetic lab, and state what your per-size taxes now expose to refutation by a live sweep that the stand-in never risked.
Exercise 4. Every arrow in a topology diagram is a token bill, an attack path, and a contagion channel (Section 23.2); price the last two for the running team of four — orchestrator, coder, reviewer, tester — under the star (every channel meets the orchestrator), the chain (orchestrator–coder–reviewer–tester), and the peer mesh. A poisoned artefact enters one member, chosen uniformly at random, and spreads along shortest paths, surviving each hop independently with probability \beta — every relay is one more chance for the checks to catch it — so a colleague at distance d is corrupted with probability \beta^{d}. (a) Count each wiring’s channels; then, in exact fractions at \beta = 1/2, compute the expected number of corrupted colleagues for every origin and the uniform average per wiring — and set the result beside Exercise 3(c): at n = 4 the star and the mesh charge the same token tax, while the mesh carries twice the ingress channels and four-thirds the average contagion. (b) Recompute the three averages at \beta = 9/10 and rank the wirings; the chain now beats the star on the average while the star’s hub remains the worst single origin on the board — reconcile both facts with the section’s claim that the star bounds contagion, saying precisely what is bounded and what is merely relocated. (c) Grant the team one perfect verifier — a member that can be neither corrupted nor used as a relay — and place it optimally under each wiring: show that the expected corruption over the remaining origins falls to 0 in the star (verifier at the hub), 1/3 in the chain (verifier at an interior post), and 1 in the mesh, and state what the star’s concentration buys that the mesh’s egalitarian wiring cannot at any placement. (d) One sentence each: why the shortest-path model flatters the mesh, and which correlated-failure channel appears in no wiring diagram and survives every rewiring — the one whose coefficient Exercise 2 taught you to estimate.
Exercise 5. The final exercise is the delegation Section 23.4 promised: write a research proposal, four pages at most, for one numbered entry of the ledger of Table 23.2 through Table 23.7. It must contain: (i) the problem restated as a falsifiable claim or a computable question, sharpened at least one turn beyond the table’s wording; (ii) the two parent literatures — the section’s advice is to hunt at the seams, and a problem with two parents inherits two literatures’ tools — each with one named result the proposal will stand on and a sentence on why its tools transfer; (iii) the first experiment or the first theorem: an experiment designed to Chapter 21’s discipline, with the strong baseline named, the comparison paired, and the claim it would license stated as the canonical triple — better at this, at this price, not elsewhere — or a theorem with its statement, the sharpest counterexample not yet ruled out, and the toy case to be proved first; (iv) the kill condition — the result that would end the direction rather than adjust it — together with a perishability audit separating what in the proposal is durable from what is tied to today’s substrate; and (v) a closing paragraph answering the section’s two selection criteria: the absence you have personally paid for, or the seam whose two parent literatures you can actually read.
Further exercises for this chapter continue in the web edition’s exercise bank.

The field’s own bench concurs: in August 2026, at IJCAI–ECAI in Bremen, the discipline’s career honour — the Award for Research Excellence, first given to John McCarthy — went to Jennings, co-author of Chapter 1’s canonical definition, for a body of multi-agent work this book has drawn on throughout; his award lecture presented agentic AI, multi-agent systems, and agent-based computing as one subject, and located the next shift not in smarter agents but in the societies they form (Jennings, 2026).↩︎