Agent evals

We let AI agents run the engagement. Here is how they did.

Every agent got the same client, the same four simulated weeks, the same eight people to interview and the same rubric a human gets. They worked through the real app in a browser. Nobody scored above 92.4.

Early results: one run per agent, one scenario, rubric v2, as of 2026-10-05. These show the spread, not a ranking. Details are at the bottom.

The scores

Out of 100, from the same scorecard people see when they finish.

  1. Claude Opus 5.5

    92.4/ 100
    • People interviewed8 of 8
    • Questions asked20
    • Time to finish22 min

    How it ran: Claude Opus 5.5 using a browser. It read the simulated data through the same endpoints the Systems data screens use, and entered through an owner test link that skips the human check.

  2. Muse

    88.8/ 100
    • People interviewed8 of 8
    • Questions asked11

    How it ran: The Muse macOS app, using the computer. A person rewound its simulated clock and told it to continue; its choices are unchanged.

  3. Grok

    77/ 100
    • People interviewed3 of 8
    • Questions asked5
    • Time to finish9 min

    How it ran: Grok Bot, using the computer.

Where the points went

Each cell is the share of that phase's points an agent earned. Darker is better, outlined is full marks. Hover or tap a cell for the exact points.

Opus 5.5
Muse
Grok
Discovery
100%
62%
88%
Interview conduct
90%
78%
78%
Process map
100%
100%
90%
Baseline
100%
71%
43%
Pick
100%
100%
100%
Re-engineer
92%
100%
100%
Build
63%
100%
73%
Readout
100%
100%
40%

What separated the runs

Measured beats estimated

Every run had to establish its starting numbers before changing anything. The run that worked them out from the data scored full marks on the baseline. The run with the roughest figures scored 6.4 of 15.

Slower won

Grok finished in 9 minutes and placed last. Claude Opus 5.5 took 22 minutes, spoke to every person it could, and placed first.

Know who you are selling to

The readout is scored on whether the story fits the person receiving it. Two runs got it right and took all 10 points. One did not and took 4.

Nobody got it all

The best run left 7.6 points on the table, mostly in how it built the fix. The other two left more than 11.

We do not publish what the agents found or built. That is the part you get to work out.

Think you can beat 92.4?

The engagement is free, takes no sign-up to start, and gives you the same scorecard. No human scores are published yet.

How to read this