All research

Research

The Research Factory

How a multi-agent pipeline ran a real factory audit and turned it into a book that caught its own errors — with an open registry of verified claims.

7 chapters, 3 open without registration.

Data you cannot trust

Management heard “deployed.” On the floor, sensors worth 500K each sat in workers' pockets. A model trained on that report does not accelerate production — it accelerates the error.

One field pilot at an industrial enterprise and one production day of assembly — N=1, and the limits of generalisation are spelled out inside. What matters here is not the conclusion but the shop floor: how 19:57:11 of recorded interviews becomes a document package, why the registry header promised 156 pain points while recounting the body gave 88, and how the finished book found an arithmetically impossible table in its own chapter about catching arithmetic discrepancies.

How it happened

I arrived at the enterprise with a ready set of hypotheses. The first thing I found: sensors costing half a million each were sitting in workers' pockets, while management had reported “deployed.” The second: ~21–22% of the headcount lives outside the official books — so any productivity figure “per official headcount” is systematically skewed, long before any AI.

Then the factory proper begins. 19:57:11 is not an estimate of interview length but a sum of timecodes: recounting the primary material instead of trusting the report. The registry promised 156 pain points; recounting the body gave 88 — and the second number is what went to the client. That same rule went into the edition's pipeline: don't trust the header, trust the recount of the body.

On the production day the book assembled itself: 5 tools, 182 subagents, 21,609,960 measured tokens — and one vendor struck from the total for missing telemetry. Then 22 agents checked the result and found an arithmetically impossible table — in the chapter about catching arithmetic discrepancies. Both facts are published: the table and the protocol that caught it.

Three techniques from the edition

Each technique links to an open chapter where it is worked through in full.

What you will take away

One item per chapter — no promises about the future.

  1. 01

    A report is not a deployment

    How to tell a deployment from a report about a deployment.

  2. 02

    A number that survives a recount

    How to count primary material so the number survives a recount on the client's side.

  3. 03

    An antagonist in a clean session

    Why an author cannot review their own work, and what an antagonist in a clean session does.

  4. 04

    What an assembly day is made of

    What a day in which research assembles itself physically consists of.

  5. 05

    A self-check that catches you

    How to build a self-check that catches you rather than flatters you.

  6. 06

    The cost is human

    Why cheaper models barely move the cost of research.

  7. 07

    The window and its review date

    Where the window of opportunity is, and the date on which it gets reviewed.

Alongside the text

Honesty counter

Every load-bearing number in this edition is a claim in an open registry with an independent verification verdict. Corrections are not hidden: both versions are published for each.

The 7 chapters — a map

  1. 0113 minRead

    Sensors in a Pocket

    Sensors worth 500K each sat in workers' pockets while management reported “deployed.” AI on such data isn't useless — it's harmful.

  2. 0214 minRead

    19:57:11

    The registry header promised 156 pain points. Recounting the body gave 88 — and only that number went to the client.

  3. 0310 minSign-up required

    The Author Is Not the Reviewer

    A source called “Gemini Research” looked like a consulting firm. The firm doesn't exist — generative AI does. The author didn't catch it; an antagonist in a clean session did.

  4. 0413 minRead

    The Factory That Assembled This Book

    This book assembled itself in a single day: 5 tools, 182 subagents, 21.6M tokens — and one vendor struck from the total for missing telemetry.

  5. 0511 minSign-up required

    The Book Caught Itself

    22 agents checked the finished book and found an arithmetically impossible table — in the chapter about catching arithmetic discrepancies. Both facts are published.

  6. 0610 minSign-up required

    Labor = 99.4%

    The token bill is arithmetic dust: 99.4% of the research cost is human labor. Cheaper models will change almost nothing.

  7. 0710 minSign-up required

    2027: Factories as a Category

    0 of 9 vendors publish an executable honesty loop. That's not a verdict on the market — it's a window, and it has an expiry date.

What's next

Read an open chapter

Sensors in a Pocket

Open chapters are readable in full with no sign-up — start with the first one.

Companion tool

Acceptance Gate

The rule this edition's registry grew from — “don't trust the header, trust the recount of the body.” Paste your own report: the gate recounts the declared numbers against the body right in your browser, with no network and no LLM.

Verified claims registry

Every claim is an atomically quotable block with a permanent anchor and an independent verification verdict. The “corrected” filter is open — that is where we suggest starting your trust check.

Showing 12 of 49

  1. 01ConfirmedSensors in a Pocket

    Теневой контингент — ~21–22% штата пилотного завода вне официального учёта (аутсорс-контингент + внеучётные работники); реальная численность выше официальной на ~15–20%, поэтому производительность и себестоимость «на официальную численность» систематически искажены

    Research primary material (internal artifact)

  2. 02ConfirmedSensors in a Pocket

    Датчики мониторинга (~500 тыс. ₽/шт) лежали в карманах рабочих «чтобы не получить штраф»; по вертикали доложено «внедрили» — вскрыто просьбой «покажите на экране» (потёмкинская цифровизация, флаг [POTEMKIN?] + Gemba-тесты)

    Research primary material (internal artifact)

  3. 03ConfirmedSensors in a Pocket

    Официальная рентабельность около 1% названа инсайдером «фикцией» (реальная — в несколько раз выше, «два контура учёта») — намеренная фальсификация учёта под налоговую оптимизацию

    Research primary material (internal artifact)

  4. 04ConfirmedSensors in a Pocket

    У коммерческих LLM с поиском и deep-research агентов 3–13% URL-цитат сфабрикованы (нет ни живой страницы, ни архивного снапшота), 5–18% не резолвятся всего (10 моделей; DRBench 53 090 URL + ExpertQA 168 021 URL)

    Source: arxiv.org
  5. 05ConfirmedSensors in a Pocket

    Deep-research агенты галлюцинируют URL статистически значимо чаще search-augmented моделей: pooled 10,7% [10,2–11,2] против 4,8% [4,3–5,2], z=15,15, p<10⁻⁵¹

    Source: arxiv.org
  6. 06ConfirmedSensors in a Pocket

    0 упоминаний руководителя ИТ-направления за 27 интервью — «молчание = данные»: формальная улика оргизоляции ИТ, а не пробел корпуса

    Research primary material (internal artifact)

  7. 07Confirmed19:57:11

    19:57:11 — сумма таймкодов 34 файлов интервью (24 роли-респондента, ~4 месяца поля); кросс-проверка вторым способом: 218 822 слова расшифровок по wc -w

    Research primary material (internal artifact)

  8. 08Corrected19:57:11

    Реестр болей: 88 дедуплицированных записей (PP-FINAL-001…088 без пропусков) — единственная защищаемая цифра тела реестра

    Research primary material (internal artifact)

  9. 09Corrected19:57:11

    Из 17 гипотез 15 строго подтверждены (1 частично — H7, 1 расширена — H15, 0 опровергнуто)

    Research primary material (internal artifact)

  10. 10Confirmed19:57:11

    Флагманский клиентский отчёт — 2 279 строк / ~18 200 слов (≈90 страниц — claimed, согласуется с объёмом слов); клиентский пакет — 8 документов

    Research primary material (internal artifact)

  11. 11Confirmed19:57:11

    Аналитический конвейер пилота: 16 сессий за 2 календарных дня против плановых 14 дней; 5 стадий, 6 ролей агентов, 21 промпт-файл

    Research primary material (internal artifact)

  12. 12Confirmed19:57:11

    Срыв интервью с ИТ-руководителем: интервьюер совместил роли «диагност» и «продавец» и перешёл на личность; дословный финал респондента: «Всё. У меня нет времени. До свидания» → правило говорения ≤30/70 и красные линии интервьюера

    Research primary material (internal artifact)

Author

Vyacheslav Fedoseev

Vyacheslav Fedoseev

AI advisor to C-suite executives. Designs and deploys agent systems; consults leadership teams on AI strategy. Founder of F/ONE — “AI for those who decide”: no hype, numbers and checkable sources.

Who orchestrates the swarm