We want to see how an agent turns a rough idea into something that works in a real, messy codebase.
The Real World Index runs 20 products through five harnesses. That gives us 100 independent attempts at the same fundamental job: understand what is needed, navigate the environment, choose a route, use the tools, recover from trouble, and deliver something coherent.
The products are not the goal. They are the evidence. We judge the outcomes qualitatively so we can study the process and mechanism that produced them—and see which harness behaves most like a capable real-world product builder.
One hundred trips through the whole maze
Starting on 1 October, each product is built five times—once by each harness. Every run begins from a controlled starting point and ends with evidence. The model and harness have to navigate the whole job: understand the brief, inspect an existing codebase, make decisions, use tools, write the product, run it, notice what broke, and try again.
These are not isolated coding puzzles with tidy answers waiting in a test file. They are end-to-end product runs with ambiguity, design judgment, integration work, revisions, and the occasional mystery caused by one missing semicolon.
The product is the visible outcome. The index is about everything the harness did to get there.
What counts as a bet
A handsome landing page cannot sneak through wearing a fake moustache. Each daily bet needs five things:
- A real job somebody recognises from their actual working day.
- A working core loop that accepts meaningful input and produces a useful result.
- A real repository with conventions, constraints, and existing code to understand.
- An inspectable product we can run, click, test, and attempt to upset.
- A complete run record showing the prompts, tools, interventions, failures, and outcome.
That last item matters. The index is only useful if somebody else can understand how the result happened—not just admire the final screenshot from a respectful distance.
Why one hundred
One impressive run is a demo. One disastrous run is also a demo, just with more character. Twenty matched comparisons across five harnesses begin to show a pattern.
The selected catalog products give the experiment range: AI research, contract analysis, document tooling, approvals, security, operations, and data work. Every harness meets the same 20 projects, which lets us compare different mechanisms against recognisably similar work.
By day 100, we should have enough evidence to establish a useful baseline: not “Which model is best?” in the abstract, but which harness process most consistently gets close to a thoughtful, usable real-world product.
The daily run
Choose the product, repository state, constraints, and acceptance checks.
Give the model and harness the brief, tools, and the same rules of engagement.
Let the system inspect, decide, edit, run, recover, and ask for help when it must.
Test the result and capture time, traces, interventions, failures, and surprises.
Every fifth run closes a product set. We put the five outcomes side by side, review how each one was made, and record the qualitative judgment while the evidence is fresh. If a harness only looks good when nobody inspects the path, this is the part where it gets nervous.
What the Real World Index watches
This is a qualitative index, not a test-suite beauty contest. Reviewers use the product to judge the mechanism behind it:
Interpretation
Did the harness understand the actual product intent, including the important things the brief did not spell out?
Approach
How did it explore the repository, form a plan, choose tools, sequence work, and revise its assumptions?
Judgment
Did it make coherent product decisions, or simply produce the most obvious thing that technically fit the prompt?
Recovery
When tools failed or assumptions collapsed, did the harness notice, adapt, and find a sensible route back?
Validation
Did it inspect its own outcome carefully enough to know whether it had built a real product or a confident approximation?
What we publish
Each five-run product set adds a case study to the index. The useful record includes:
- The task: brief, constraints, repository state, and acceptance checks.
- The setup: model, harness, tools, configuration, and instructions.
- The run: trace, tool calls, time, interventions, and failure points.
- The judgment: a qualitative review of the process, the mechanism, and how close the outcome came to a real product.
The result will not be a universal leaderboard carved into stone. It will be a living index of harnesses doing real product work—concrete enough to compare, broad enough to expose patterns, and detailed enough to show why an outcome happened.
After one hundred days, we will have 20 products built five different ways, a clearer view of which mechanisms work in the real world, and at least a few excellent stories about the ones that did not.
The calendar is the live map: 20 catalog products, five harnesses, and 100 runs from 1 October through 8 January.