· DEEP TECH
Reverse Benchmarking. Measure AI on what the leaders neglect, not what they optimise
Benchmarks make every lab chase the same scores. The useful question runs the other way: where does the best model still disappoint?
01 / THE BEER AND THE COFFEE
Will Guidara, who ran Eleven Madison Park, once took his team to dinner at the best restaurant in New York. Everyone took notes on what it did well. He asked what was disappointing. The answer was the coffee and the beer. So he made one colleague a coffee sommelier and another a beer sommelier.
Rory Sutherland calls this reverse benchmarking. Benchmarking copies what the leader does well, so everyone converges on the same average. Reverse benchmarking finds what the whole field has quietly accepted doing badly, and overinvests there.
02 / BENCHMARKS ARE CONVERGENCE
AI benchmarks are benchmarking in the plain sense. Every lab climbs the same leaderboards, so models converge on the same strengths, and once a measure becomes a target it stops measuring. OSWorld-Verified now sits above its own human baseline. On OSWorld 2.0, with hour-long real workflows, the best agent at release finished about one in five.
Leaderboards also get played. The Leaderboard Illusion found that Meta privately tested 27 model variants on Chatbot Arena before releasing Llama 4 and kept the best score. Two identical checkpoints under different names landed 17 points apart. At that point the leaderboard ranks leaderboard skill.
03 / WHERE THE BEST MODEL DISAPPOINTS
Finishing. The Remote Labor Index takes 240 real freelance projects, from 3D models to data analysis, and asks whether a client would accept the AI’s work. At launch the best agent passed 2.5 percent. Today the best passes about 21 percent. Fast progress, and four in five projects still fail.
Felt speed against real speed. In 2025 METR ran a randomised trial with experienced open-source developers working in their own repositories. They expected AI to make them 24 percent faster, and afterwards believed it had made them 20 percent faster. It made them 19 percent slower.
Uncertainty. Sutherland’s point about trains: passengers mind not knowing more than they mind waiting, which is why a map of your approaching cab beats a faster cab. Agents leave people in exactly that state. They guess instead of asking, and rarely say what they are doing or how sure they are.
04 / RUN IT BACKWARDS
Normal benchmarks run forwards: imagine tasks, score models, hope the ranking predicts reality. A reverse benchmark starts at the other end. Take real outcomes (accepted work, money earned, hours saved) and real failures, and derive the test from them.
Then judge the benchmark on one thing: does it rank systems the way reality does? A simulator should be judged by the agents it produces, not by how real it looks.
That makes failures the raw material. Every agent that stalls, overspends or ships half a job is a test case nobody else has. Like a regression test, it goes into the suite and stays. Where the target moves, a search ranking or a market, the test cannot be memorised.
05 / THE DISAPPOINTING PART
The leaders’ weaknesses are a map. In AI they are finishing, honesty about uncertainty and the real cost of a real job. Few people measure them, which is exactly why they are worth measuring.
Ask what was disappointing. Then build the beer menu.