Chapter 1. Sensors in a Pocket
A humidity sensor costs on the order of five hundred thousand rubles apiece. At the plant where we were running an audit, those sensors sat in the pockets of workers' overalls.
Not in the mix they were supposed to measure — in pockets. The instruments had been damaged during equipment cleaning, and, to avoid a fine, workers took them off and hid them. Up the management vertical, meanwhile, went a report: the system is deployed and working. The chief engineer reported "deployed" — and, by all appearances, believed it himself: on his slide the system existed. It all came to light not through clever analytics or a log audit but through a single request: show it on the screen. No screen on which the sensors displayed anything could be found.
I believed the sensor slide. What changed my mind was not an and not a report — it was the request to show the screen. Since then that rule has been hard-wired into our verbatim, as a mandatory step, not a pious wish.
Potemkin Digitalization
In our methodology this phenomenon has a name — Potemkin digitalization: a state in which digital systems exist in reports and do not exist on the shop floor. It is almost never one person's malice. It is an equilibrium: it is easier for the worker to hide the sensor than to explain the damage; easier for the foreman to confirm "it works" than to investigate; easier for the manager to report an aggregated "deployed" upward. Everyone on their own floor smooths things over just a little — and at the top those smoothings assemble into a fictional enterprise that shares nothing with the real one but the address.
Same plant, another unit, the same pattern in miniature: a process engineer in whose account the unit is fully digitized — and paper on a clothespin sixty seconds of the same recording away from those words; the full scene, timecode and all, is dissected in Chapter 2. The man was not lying: he first described how things are supposed to be, then how they are, and never noticed the transition.
That is why, in the factory's , any about a deployed system gets a flag and a procedure. Here is what it looks like in the working instruction — a fragment of the rule from our methodology, in the exact wording received by both the and the human consultant:
Introduce a verification category "claimed deployed vs actually working": re-check every claim about automation, a sensor, or a system with a line-level worker — a lab technician, a machine operator, a driver — not with a manager. Demand physical proof: a photo of a working screen, a live demonstration. Set the
[POTEMKIN?]flag by analogy with[VERIFY]; keep a "claimed / confirmed by demonstration" column in the coverage matrix; downgrade the confidence of any claim confirmed only by words to Low. For the most sensitive hypotheses — accounting, write-offs, quality — prepare two or three physical Gemba tests in advance, plus unscheduled windows in the visit itinerary.
The Gemba tests in that last sentence are checks made on foot, on site. The "scale test": if weighing data is entered by hand, fraud has a guaranteed door. The "warehouse test": take an item from the accounting system, find it in the yard, time the search — then run it in reverse, from the hardware back to the record. The "unit test": there is an automatic doser — watch whether the operator tops up water "by eye" after the auto-dosing. Not one of these tests requires AI. They require distrust of self-reporting — and that distrust turned out to be the scarcest resource of digital transformation.
It is worth saying how this approach differs from how the market usually measures "digital maturity." The industry's standard instrument is a self-assessment questionnaire: managers receive a survey, score themselves along the axes, and a consultant assembles a radar chart. By construction, that is the institutionalized reported version. The questionnaire does not merely fail to catch the hidden sensor — it issues it a certificate: a checkmark appears in the "IoT monitoring" box, signed by the very manager who had been told "deployed." We are not claiming questionnaires are useless — they are cheap and fine as a warm-up. We are claiming that building investment decisions on them is like running a medical exam on a "how do you feel?" form. The research factory begins where faith in self-reporting ends — anyone's, including our own, and we will come back to that "including our own."
From the Sensor to the Data Loop
One hidden sensor is an anecdote. The problem is that the sensor turned out to be not an exception but a specimen: the deeper the audit descended along the data loop, the wider the reported and observed versions of the enterprise diverged.
The heaviest finding was people. Roughly 21–22 percent of the workforce existed outside the accounting system: outsourced staff on service contracts and workers never entered into the HR loop. One person in five on the site was invisible to any system that computes productivity, unit cost, or payroll. That means all output "per official headcount" is overstated by roughly the same share, and any efficiency calculation built on HR data starts with a twenty percent error — before its first formula.
Next, the economics. The enterprise's official profitability was around one percent, and the very first cross-interviews showed the figure to be a fiction, the product of tax optimization and "two accounting loops"; insiders put the real profitability several times higher. Materials were written off by norms, with deviations smeared across all products at once — meaning the margin of a specific order was fundamentally incomputable: not "hard to calculate" but impossible in principle on this data. Metal was accepted "by theoretical weight" from the tags, without weighing. Stock in one of the shops went negative in the books — the warehouse held a negative quantity of raw material, and the system did not object.
And on top of the distortions — delays. Second-shift data was keyed in by hand the next morning; actual cost closed toward the end of the following month, when the electricity and cement invoices arrived. Any "real-time" analytics on this foundation is real-time in appearance only: it shows the day-before-yesterday's plant, to within the distortions. One of our respondents aptly called management on such data postmortem — decisions are made from the autopsy, not from the pulse.
A separate thread is degradation over time, which no snapshot of the current state can see. We asked not only "how does this work" but "how did this work three to five years ago." Acceptance at shipping: the customer's representative used to reject items right at the truck — now the driver signs the paperwork without looking, and defects surface at the construction site, weeks later and through the courts. Stock control: it used to reconcile — now it has "gone negative." Every such delta is not a fact about today; it is a vector, and it pointed down. An enterprise degrading in its books degrades there at an accelerating rate: the less the data is trusted, the less anyone wants to maintain it.
And it must be said honestly: the field resists measurement actively, not passively. The first item in our field-collection risk register is "Potemkin villages" — the immaculate order staged for the consultant's visit. The mitigations are prosaic: look at twelve months of historical logs before the visit, not the shop window on the day of the visit; leave unscheduled windows in the itinerary and walk into the shop when no one expects it; give people an anonymous channel and an explicit amnesty for disclosing old problems. An audit the enterprise knows about in advance and prepares for measures not the enterprise — it measures the quality of preparation for the audit.
It matters who told us all this. Not a whistleblower, not a disgruntled employee — the managers themselves, in calm interviews, often with a tone of "well, you understand how these things work." The shadow workforce, the two loops, the negative stock — at the plant these are not secrets but background everyone is used to. Nowhere was the data distortion anyone's decision; it was the sum of a thousand small conveniences. That is exactly why it is invisible from the inside and cannot be found with a questionnaire: the questionnaire is filled in by someone who stopped noticing long ago.
Now the chapter's central move. Ask yourself: what happens if you put AI on top of data like this?
The plant, like most mature production sites, already had its own graveyard of pilots: a messenger app for work orders that never took hold; a laser rangefinder rejected because the beam is invisible on long spans in summer; marking guns turned down on price and handling at once. Every failure had a reason, and not one of those reasons was "people are stupid." A new wave of automation that has not studied this graveyard is doomed to repeat it — faster and at greater cost this time. Why AI pilots die on data more often than on models is dissected systematically in our agent-swarm ; here we show the particular case from which the rule grew for us in the first place.
The answer is more unpleasant than it seems. Naive automation on fictitious data is not useless — it is harmful. A dashboard built on two accounting loops will not show "roughly the truth" — it will show a confident lie to the second decimal place. A forecasting model trained on headcount missing a fifth of the workforce will err systematically — and do it beautifully, with charts. AI fixes nothing in the data: it industrializes whatever is in there. An error that was manual and slow becomes automatic and at scale. The pilot report states the conclusion outright: AI on such data will produce erroneous conclusions; first a loop of trustworthy data, then everything else.
Here our expectations were wrong, and it is worth saying so honestly. We went into the audit with the hypothesis that the main work was to find the points where AI would deliver a quick win. The main work turned out to be building a map of where the data could be trusted at all. The data-readiness axis — the one our readiness scanner runs every company through before any talk of models — was born not from the literature but from this plant.
The Mirror: The Machine Lies as Convincingly as the Reporting
It would be convenient to end on the moral "people embellish — the machine is objective." The field data will not allow it, and neither will our own instrument.
A errs not the way a human does — it errs better. The error arrives in marketable condition: with a confident tone, coherent logic, and plausible references. The report "sensors deployed" and a generated report with a nonexistent source are twins: both look more convincing than the truth, because no one checked either one on site.
The scale of the mirror has been measured by external research. A 2026 audit of commercial models and research showed that 3 to 13 percent of the URL citations in their answers are fabricated — the page never existed, not on the live web and not in the archives. What is more, "" agents references twice as often as ordinary models with search: 10.7 percent versus 4.8 — a difference that is not noise but statistics over hundreds of thousands of verified URLs. The instrument positioned as more thorough lies about its sources more often — simply because it generates more of them and verifies them exactly the same way: not at all.
The same studies carry the flip side, important for everything that follows: when agentic self-checking of references is attached to the — a trivial one, via archives and existence checks — the share of broken citations falls by orders of magnitude, to fractions of a percent. Hallucination, then, is not an irremovable defect of the technology but a property of an unverified pipeline. Exactly as Potemkin digitalization is not a property of factories but a property of a vertical in which no one asks to see the screen. The diagnosis is the same, and so is the treatment: built-in, mandatory verification that believes nothing.
Read the symmetry closely. The enterprise reports itself better than it lives. The model cites more confidently than it knows. Both are self-reporting without an external check; only the medium changes. Which means a method that an honest result is obliged to trust no participant in the pipeline: not the client's data — field triangulation and Gemba check that; not its own agents — they are checked by the verification loop laid out in Chapter 3; not even its own pretty numbers — how one of those was executed is the story Chapter 5 tells. The research factory is assembled from precisely this triple distrust. The word "factory" here is not a metaphor of scale but a metaphor of quality control: the product does not leave the shop without inspection, because the machine's defect rate is known.
What does this principle mean at the start of any new project? It fits into one rule, and we will end the chapter with it — no lists, no checklists.
Every fact about an enterprise exists in two versions: reported and observed. The reported version comes free — in a presentation, in a self-assessment questionnaire, in a confident answer at a meeting. The observed one costs money: a trip to the shop floor, the request "show it on the screen," a recount of the registry against the body of the file. The research factory is a machine that systematically pays for the second version. Everything else in this book is the workings of that machine: what twenty hours of field recordings turn into on the pipeline — Chapter 2; who catches the errors of the pipeline itself — Chapter 3; what one of its production days looks like, counted down to the — Chapter 4.