· DEEP TECH
System 1 Is Solved. AI agents now fail at spatial reasoning, long horizons and context
Fast, perceptual work is done: agents read screens, click and type as well as people. What breaks now is the slow work of holding a space, a long task and a lot of context in mind.
01 / SYSTEM 1 IS SOLVED
Daniel Kahneman split thinking into System 1, fast and automatic, and System 2, slow and deliberate. Reading a screen, finding the button, filling the form: that is System 1, perception mapped straight to action.
Clicking is the clearest proof. In 2024 OSWorld put agents on a real desktop with 369 tasks. Humans finished about 72 percent. The best model finished about 12. Today the top entries on the verified leaderboard sit above 80 percent, past the human baseline.
Computer use as a research frontier is mostly over. What remains is everything System 2 does.
02 / LONG HORIZONS
OSWorld 2.0, released in June 2026, made the tasks real: 108 workflows that take a person about an hour and a half, around 318 tool calls each instead of 30. At release the best agent completed 20.6 percent, with 54.8 percent partial credit.
The authors are blunt: the failures are not about GUI control or coding. Agents drop constraints stated at the start. They miss information that arrives mid-task. They guess instead of asking. They spend under 7 percent of their budget repairing their own mistakes. The agent can do every step. It cannot hold the whole.
03 / SPACE IS NOT LANGUAGE
In late 2024 the Thinking in Space study showed models video of real rooms and asked how far, which side, what route. Humans scored 79 percent. The best model scored about 45.
Chain-of-thought did not help. Asking the model to draw an explicit cognitive map did. Spatial reasoning does not live in words, so reasoning in words does not fix it. A 2026 re-audit, ReVSI, puts frontier models near 60 percent. Better, still short of a person glancing around a room.
04 / CONTEXT IS THE JOB
A long task is a context problem. Every step adds tokens, and the agent has to decide what to keep, what to drop and what to write down.
Bigger windows do not solve it. Lost in the Middle showed in 2023 that models use information at the start and end of a long input far better than the middle. In 2025 Chroma tested 18 frontier models and found performance falls as input grows, even on simple tasks. More context is not more memory.
What works is structure outside the model: a map for space, durable task state for time. A three-hour task should live in an issue with an owner, a status and a history, not in a context window. That is the argument of The Issue Is the Primitive.
05 / MORAVEC, AGAIN
Hans Moravec noticed in 1988 that what is hard for people is easy for machines, and the reverse. We got chess, then language, then the screen. Holding a three-dimensional world and a three-hour task in mind at once is the old animal skill. That is the frontier now.