© 2026

· DEEP TECH

The World Is Not Logged. Why real-world data, not intelligence, limits AI

The world records everything, in light, heat and wear. Almost none of it reaches a model. That gap, not intelligence, is what limits AI.

01 / THE CAT IN THE BOX

In 1935 Erwin Schrödinger sealed an imaginary cat in a box with a radioactive atom and a vial of poison. Read naively, quantum mechanics says the cat is alive and dead until someone looks. He meant it as a joke against that reading.

Later physics gave a better answer: the box was never closed. The cat warms the air and scatters light, and its surroundings keep copies of what happened. Wojciech Zurek calls this quantum Darwinism. The world logs itself constantly. Almost none of that log reaches a model.

02 / WHAT A MODEL CAN SEE

In 1960 Rudolf Kálmán defined observability: a system is observable if you can infer its internal state from its outputs. A language model sees one output, the written record of human life. It is rich in arguments and code, and thin on kitchens, factories, soil and bodies.

Claude Shannon measured information as surprise and called it entropy. A model learns only from what it could not predict. Static has maximum entropy and teaches nothing. What counts is surprise that turns out to be structure.

So when a model is confidently wrong about the physical world, I don’t see a stupid system. I see an unobservable one. The state it needed never reached the data.

03 / THE WELL IS RUNNING DRY

Epoch AI projects that frontier models will train on datasets the size of all public human text sometime between 2026 and 2032.

The well is also closing. In one year, from 2023 to 2024, websites used robots.txt to block AI crawlers from 28% of the most important sources in C4, a standard training corpus. And it is filling with model output: by AI detectors’ count, about half of new English articles have been mainly AI-written since 2025.

Scientists still salvage steel from ships sunk before 1945 for radiation detectors, because steel made after the first bomb tests carries traces of fallout. Human text from before late 2022 is becoming the same thing: low-background data.

04 / THE OUROBOROS

Train a model on its own output and it eats its tail. Shumailov and colleagues showed in Nature in 2024 that each generation loses the tails of the distribution first, the rare cases, until a narrow average is left. Alemohammad’s team named it Model Autophagy Disorder, after mad cow disease: without enough fresh real data, quality or diversity decays within a few generations.

The newer research is more precise. Gerstgrasser and colleagues showed that collapse comes from replacing real data. Keep the real data and let synthetic data accumulate beside it, and error stays bounded. Kazdan and colleagues confirmed this in 2025, with a twist: synthetic data helps when real data is scarce and hurts when it is plentiful.

Schrödinger saw the principle in 1944. In What Is Life? he argued that living things stay ordered by drawing order from outside: “What an organism feeds upon is negative entropy.” A model is a structure of the same kind. Close the loop and it runs down.

05 / SYNTHETIC AND ORGANIC

Synthetic data works when something outside the model checks it. AlphaZero mastered chess by playing itself, because the rules decide who won. DeepSeek-R1 learned to reason through reinforcement learning with rewards from answer checkers and compilers, not human examples. The verifier is where the real world enters the loop.

Without a verifier, synthetic data rearranges what the generator already knows. It can add coverage, format and edge cases. It cannot add a fact that never reached the generator. A simulator only knows the physics someone put in it.

My rule: synthetic data multiplies organic data. It does not replace it. Organic data brings the surprise. Synthetic data spreads it around.

06 / THE PRICE OF LOOKING

In 1867 James Clerk Maxwell imagined a demon that sorts fast molecules from slow ones and seems to beat the second law. Leo Szilard showed in 1929 that the demon must measure. Rolf Landauer and Charles Bennett found the bill: erasing a bit of memory costs energy. Information is physical, and so is observation.

Robotics shows the price. Open X-Embodiment pooled over a million real robot trajectories from 22 kinds of robot. HIW-500, a 2026 humanoid dataset, is 500 hours of teleoperation in 12 real homes. A frontier language model reads tens of trillions of tokens.

NVIDIA trained its Cosmos world models on 20 million hours of real video and released them to generate synthetic training data for robots. That is the pattern: expensive real data, distilled into a cheap simulator.

07 / THE NEXT TEN YEARS

Data will be collected, not scraped. The open web was a one-off windfall. What comes next is paid for: licensed archives, sensors, vehicle fleets, teleoperation, and people doing their work on the record.

Provenance becomes valuable. Data that is human, dated and signed will be worth more, the way low-background steel is.

Agents become instruments. An agent that acts in the world and records the result is a sensor. The loop that matters is act, observe, verify, train.

The most valuable AI companies will look like observability companies. They will own a stream of real-world data no one else logs, clean enough to train on and fresh enough to act on.

08 / LOOKING CHANGES THINGS

In physics, a measurement needs no mind. In society it does, because people change when they are recorded. Charles Goodhart put it in 1975: once a measure becomes a target, it stops measuring. So keep raw signals, not summaries. Bennett’s lesson holds here too. The irreversible step is not recording. It is throwing the record away.

09 / INSTRUMENTS FIRST

Brahe measured before Kepler explained. Leeuwenhoek saw microbes two centuries before germ theory. Intelligence follows instruments.

The world is already writing everything down. The next leap in AI belongs to whoever builds the instruments to read it.