Measured beats estimated
Every run had to establish its starting numbers before changing anything. The run that worked them out from the data scored full marks on the baseline. The run with the roughest figures scored 6.4 of 15.
Agent evals
Every agent got the same client, the same four simulated weeks, the same eight people to interview and the same rubric a human gets. They worked through the real app in a browser. Nobody scored above 92.4.
Early results: one run per agent, one scenario, rubric v2, as of 2026-10-05. These show the spread, not a ranking. Details are at the bottom.
Out of 100, from the same scorecard people see when they finish.
How it ran: Claude Opus 5.5 using a browser. It read the simulated data through the same endpoints the Systems data screens use, and entered through an owner test link that skips the human check.
How it ran: The Muse macOS app, using the computer. A person rewound its simulated clock and told it to continue; its choices are unchanged.
How it ran: Grok Bot, using the computer.
Each cell is the share of that phase's points an agent earned. Darker is better, outlined is full marks. Hover or tap a cell for the exact points.
Every run had to establish its starting numbers before changing anything. The run that worked them out from the data scored full marks on the baseline. The run with the roughest figures scored 6.4 of 15.
Grok finished in 9 minutes and placed last. Claude Opus 5.5 took 22 minutes, spoke to every person it could, and placed first.
The readout is scored on whether the story fits the person receiving it. Two runs got it right and took all 10 points. One did not and took 4.
The best run left 7.6 points on the table, mostly in how it built the fix. The other two left more than 11.
We do not publish what the agents found or built. That is the part you get to work out.
The engagement is free, takes no sign-up to start, and gives you the same scorecard. No human scores are published yet.