1 models, 171 countable harness × task × trial cells.
Non-comparable provenance: image_digest differs across groups; harness_version_source mixed within group; harness_version_source mixed within group; harness_version_source mixed within group; timeout_s differs across groups. Inspect source rows before publishing.
| Arm (harness × model) | Solved/countable | Solve rate | Wilson 95% CI | Med wall | Total tokens/solve (incl. cache reads) | Uncached in/solve | Output/solve | Cache-read/solve | Token basis | Excluded / unmatched / duplicates |
|---|---|---|---|---|---|---|---|---|---|---|
| cursor × grok-4.5 | 41/45 | 91.1% | 79.3–96.5% | 35.5s | — | — | — | — | self-reported, unknown | 0 / 0 / 0 |
| grokbuild × grok-4.5 | 40/45 | 88.9% | 76.5–95.2% | 64.6s | 607,005 | 59,683 | 13,492 | 533,830 | self-reported | 0 / 0 / 0 |
| pi × grok-4.5 | 34/43 | 79.1% | 64.8–88.6% | 35.4s | — | — | — | — | self-reported, unknown | 2 / 0 / 0 |
| opencode × grok-4.5 | 30/38 | 78.9% | 63.7–88.9% | 21.2s | 277,307 | 43,692 | 9,138 | 224,478 | self-reported | 7 / 0 / 0 |
Timeouts: pi: 1 timeout; finished-basis 81.0%.
Total tokens/solve (incl. cache reads) is computed uniformly from the split fields: uncached input + output + cache read. The vendor `tokens` field is never used. Split fields use each harness's own classification and are NOT cross-comparable. Badge bases: unmetered / self-reported / proxy-measured.
| Harness | grok-4.5 |
|---|---|
| cursor | 91.1% |
| grokbuild | 88.9% |
| opencode | 78.9% |
| pi | 79.1% |