Chapter 4. The Factory That Assembled This Book
Everything you have read up to this page and will read after it was produced by the this chapter shows from the inside. Not "roughly like it," not "similar in spirit" — this very one. One calendar day, July 25, 2026, five tool entities, eleven runs. The day has a ship's log with run identifiers, and we are publishing it almost in full: with the , the minutes, the incidents, and one vendor that had to be struck from the total.
Let's start with the night shift.
00:22. The Swarm Reads the Archive
At 00:22 the first launches: 67 get the task of reading the 191 files of the pilot project's chaotic folder — transcripts, registries, chapter drafts, , contract documents. Forty-eight readers walk the corpus in parallel, seven agents go out to the web for external practices, one strong agent assembles a of the official methodology, six "gap hunters" look for what the methodology lacks, and five agents synthesize the findings. After 86.6 minutes the swarm stops. The platform counter shows 8,624,353 , 727 , zero errors across all 67 launches. The output: a consolidated methodology core, 51 anonymized lessons, and a map of the gaps.
Next comes a mode we call the "clean laboratory." The pilot project is packed with sensitive material: names, company figures, respondents' confessions. Before a single line goes into the public contour, two specialized swarms — four sanitization agents and three verification agents — scrub the corpus, and the repository's git history is recreated from scratch so that physically no trace of the source material remains in it. The verification swarm finds and closes five residual leak risks. On top of the machine check come three full grep sweeps over a ban list of roughly forty-five patterns, including Cyrillic homoglyphs: the letter "х" typed in Cyrillic does not equal the Latin "x" to a search engine, and that once already broke an entire layer of automated control.
This is also where the day's first incident happened. An argument-substitution bug in the prompts broke the paths: four agents received garbage instead of addresses as input, adapted correctly, and finished the job — but four files landed in improvised directories, and it was me who moved them, by hand, watching a clock that read past two in the morning. An important detail for the record: the agents did not err once — the did, which is to say, I did. We logged it as an incident, not a footnote, because events happen to someone, and in our factory that someone is still a human.
One Agent Against the Swarm
By four in the morning the night shift's executive summary was ready, and by regulation it cannot be handed in without an outside look. A single, solitary was set loose on the summary — a "cold" reviewer with a clean context who had seen no drafts and taken part in no . It spent 99,665 and six minutes — the cheapest worker of the day. And it was the one who caught what 67 readers and I myself had walked past: the summary carried a salesy "30×" multiplier where the data defended only "20–30×." Verdict PASS_WITH_NOTES, five edits, multiplier lowered.
Six minutes of a loner against an hour and a half of a swarm — and a catch of equal value. This is not a paradox but a property of the : generation and verification are different professions, and the second is radically cheaper when it has a clean context and a mandate to spoil the mood. Why our entire trust contour is built this way is the subject of Chapter 3; what matters here is that the rule worked on the edition itself.
Five Tools, Five Characters
During the day the remaining entities joined the factory. By evening each had developed a character — and a different degree of provability of what it had done.
The two Claude — the production session and a separate discovery — supplied the bulk of the measured numbers: every run has an identifier in the , every a counter. The discovery swarm closed the dossier layer in three waves: corpus, web, digest with a gap audit — 12 agents, 1,370,473 tokens, 321 calls, 36.8 minutes, zero errors and zero re-runs across all twelve. The "headline versus body" self-check — a pilot lesson described in Chapter 2 — converged twice.
Codex played the honesty engineer: it built the data schemas, the counter-aggregator script, and the edition's . Its session is a specimen of a different virtue: tests 24/24 PASS in 0.492 seconds in a clean environment, while the diagnostic gate at that point honestly did not open — FAIL, 808 errors in 12 categories on a corpus not yet migrated. The gate's failure was correct behavior: the corpus genuinely wasn't ready, and a gate that opens in that situation is worth nothing. At the same time, the Codex platform does not export tokens, cost, or the exact model identifier — so its passport carries null with the status not_instrumented, not a little zero and not a guess. Zero and "not measured" are different statements; a factory that confuses them lies.
Grok worked as the scout of social signals and the market — and worked deliberately alone: one agent, zero , six waves. That is a conscious trade-off: parallelism would have sped up the search but multiplied the risk of source-hygiene drift; for a layer where a single voice and a single link registry matter, depth beat speed. The Grok platform does not expose a token counter, so its entire numeric side is marked estimated — an order of magnitude reconstructed from the volume of artifacts, and nothing more precise. It has one distinction of its own: it also shot down one of our future cover numbers, but that story belongs entirely to Chapter 5.
Antigravity was billed as the orchestrator of cross-tool waves — and became the day's most instructive character. Its report is beautiful: thirteen subagents, three waves, "Tier-1-grade artifacts." The problem is that its own log for the entire day holds exactly one entry — initialization, two minutes in the morning — and there are no run identifiers, no per-agent numbers, no timecodes at all. The thirteen subagents exist only on the tool's own word.
I struck Antigravity from the measured total. Not because I don't believe it did anything — a synthesis cluster overlapping in time is visible in the artifacts of the other entities. But because "I believe" is not a category of . A beautiful self-report without a single run_id is worth exactly as much as the "sensors deployed" report from Chapter 1. A pity? Yes, a pity: the figure "182 + 13" would have looked more impressive. But let one row in on its word of honor, and the whole table of measured numbers turns into an opinion.
The contrast — a tool with telemetry versus a tool without it — is itself the main takeaway of this section. We weren't choosing whom to praise; we were choosing whose numbers can be defended before a hostile auditor.
The Day's Total, Counted to the Last Unit
Now the sum. Over the production day and the separate consolidation session — the one where a of 16 readers re-read the entire corpus, a swarm of 14 writers assembled the chapters, the verification and fix swarms ran the text through the distrust contour (their catch is dissected in Chapter 5), and a clean-context controller confirmed the defects closed — the Claude orchestrators dispatched 182 subagents: 115 during the day and 67 in consolidation. An attentive reader will notice that the night swarm also consisted of 67 agents — and suspect double counting. We suspected the same and re-checked by run_id: the coincidence is accidental. The night's 67 are part of the production day's 115; the consolidation's 67 are a different session with different identifiers: 16 readers, 14 writers, 22 reviewers, 9 fixers, and six service agents for revalidation and control.
The consolidation session deserves one detail, because it shows the fault-tolerance mechanism at work. Of the sixteen readers, fifteen delivered valid reports on the first try; one produced a structural format failure. Nobody patched its output by hand or pretended it had slipped through: the validator rejected the report, a separate agent re-ran the reading and restored all 87 claims of the failed digest down to the last one — the verdict REPAIRED is written into the log next to the original failure. Pipeline resilience is not the absence of failures but the guarantee that a failure neither passes unnoticed nor loses data. The total measured counter is 21,609,960 tokens. Not "about 21.6 million": the precision here is a gesture. Rounding numbers the platform reports down to the unit means throwing away the only thing that separates a measurement from an impression.
Around that figure stands the rest of the day's skeleton: 1,845 plus 1,094 tool calls, roughly 187.6 and 109.5 minutes of pure swarm time, zero lost results across all runs (the one structural format failure closed by a re-run with revalidation). The model split is 74% cheap workhorse model and 26% expensive: breadth is done by the cheap one, synthesis and verification by the strong one, and not a single expensive agent was spent that day on mechanical reading. Roughly every tenth agent is a pure verifier that produced not one line of content: that is the price of the distrust contour, and we pay it deliberately — what it means for unit economics is computed in Chapter 6.
And there is one hole in this telemetry — my own. The tokens of the orchestrator's main context, the window in which I assigned tasks and picked apart incidents, are not written by the harness: in the log they are marked not_recorded. We could have estimated and added them — the figure would have come out bigger and "more complete." We didn't: we have one rule for everyone, ourselves included — never substitute an estimate for a measurement. Better a hole marked not_recorded in the total than a pretty patch.
What It Looks Like in the Raw
Words about "machine-readable honesty" are cheaper than the thing itself, so here are three entries from the real run log — as they sit in the repository, down to the statuses:
timeline_entries:
- label: "claude:stage1-corpus-consolidation"
run_id: "wf_af5f9b65-f77"
start: "2026-07-25T00:22:00+03:00"
agents: 67
tokens: 8624353
token_status: measured
tool_calls: 727
duration_min: 86.6
- label: "codex:T-07-data-layer"
run_id: "RUN-20260725T082240Z-codex-a7d6e501"
agents: 5 # root + 4 subagent tasks
tokens: null
token_status: not_instrumented
tests: "24/24 PASS"
- label: "antigravity:orchestration"
run_id: null
agents: 13 # self-reported, no run_id
tokens: null
token_status: self_reported
note: "the claimed 3 waves and 13 subagents are not confirmed by telemetry"Three entries — three gradations: measured, not_instrumented, self_reported. A single token_status column does what no report on "multi-agent systems" known to us does: it separates the numbers by degree of provability right in the data, not in a fine-print footnote.
And here is how the factory greets a tool at the door — a fragment of the real kickoff prompt Codex received that morning:
Your superpower is engineering discipline: data schemas, scripts, aggregators, reproducibility. Others write the chapter content; you make sure that not a single number in the prose diverges from the data. […] The counter-aggregator script is the single source of all counters: any figures "about the research itself" in prose come only from its output. Lesson: in the previous run the counters diverged between documents — that must not happen again. […] Build as if a hostile auditor will re-check it.
Note that the prompt does not ask it to "do a good job." It hands over a specific past failure and forbids its repetition. The factory learns not by declarations but by stitching its own scars into the instructions.
And the third exhibit — the very gate that honestly did not open. A line from the Codex release-gate report that day:
status = FAIL, exit code 1
808 errors / 12 categories, 1 warning
among the categories: E_PROSE_NUMBER_UNCLAIMED — 79,
E_SOURCE_ID_OUTSIDE_REGISTRY — 199, E_SCHEMA_REQUIRED — 244Seventy-nine numbers in prose not tied to the claim registry. One hundred ninety-nine source references outside the registry. The gate does not judge the beauty of the text — it checks that every figure has a home, and as long as there were no homes, it kept the gate shut. The edition you are reading passed through a descendant of that gate.
The Passport Office
The day's summary table — five entities brought to a single denominator of honesty:
| Entity | Role | Agents | Tokens | Number status |
|---|---|---|---|---|
| Claude, production day | swarm orchestrator: reading, clean lab, analysis, discovery | 115 subagents | 13,718,066 | measured |
| Claude, consolidation session | reading swarm → synthesis → verification → fixes → control | 67 subagents | 7,891,894 | measured |
| Codex | engineer of the data layer and the release gate | 1 root + 4 tasks | null | not_instrumented |
| Grok | market and social-signal reconnaissance, single voice | 1, no subagents | ~1.1–2.2M input | estimated |
| Antigravity | claimed cross-orchestration | 13 self-reported | none | self_reported, excluded from total |
The total over the rows with measured status: 21,609,960 tokens, 182 subagents, zero lost results. An edition about a research factory was produced by that same factory, and that is not a marketing flourish but a verifiable claim: every run in the table can be opened up down to the level of an individual agent.
The same principle — telemetry instead of promises — also runs on the portal you are reading right now: the "How It's Built" page shows the same kind of log for the site itself. Why a swarm is justified at all on tasks like these, and where it breaks down, is canon in the agent-swarm playbook; this day is its field proof. And what it cost in money — and why the token bill turned out to be the least interesting line of the estimate — is taken apart in Chapter 6.