1 models, 207 countable harness × task × trial cells.
Non-comparable provenance: harness_version mixed within group; harness_version mixed within group; harness_version mixed within group; harness_version mixed within group; harness_version mixed within group; harness_version_source mixed within group; harness_version_source mixed within group; harness_version_source mixed within group; harness_version_source mixed within group; harness_version_source mixed within group. Inspect source rows before publishing.
| Arm (harness × model) | Solved/countable | Solve rate @cap | Solve rate finished | Wilson 95% CI | Med wall | Total tokens/solve (incl. cache reads) | Uncached in/solve | Output/solve | Cache-read/solve | Token basis | Excluded / unmatched / duplicates |
|---|---|---|---|---|---|---|---|---|---|---|---|
| pi × deepseek-v4-flash | 29/34 | 85.3% | 87.9% | 69.9–93.6% | 24.5s | — | 34,328 | 29,104 | 1,712,945 | proxy-measured, self-reported | 11 / 0 / 0 |
| claude × deepseek-v4-flash | 31/43 | 72.1% | 72.1% | 57.3–83.3% | 48.4s | — | 39,927 | 56,538 | 2,443,483 | proxy-measured, self-reported | 2 / 0 / 0 |
| opencode × deepseek-v4-flash | 28/44 | 63.6% | 63.6% | 48.9–76.2% | 45.8s | — | 44,649 | 43,508 | 1,904,229 | proxy-measured, self-reported | 1 / 0 / 0 |
| grokbuild × deepseek-v4-flash | 27/45 | 60.0% | 64.3% | 45.5–73.0% | 35.9s | 2,256,739 | 51,637 | 61,784 | 2,143,317 | self-reported | 0 / 0 / 0 |
| codex × deepseek-v4-flash | 24/41 | 58.5% | 66.7% | 43.4–72.2% | 27.2s | — | 43,935 | 63,045 | 1,257,808 | proxy-measured, self-reported | 4 / 0 / 0 |
Total tokens/solve (incl. cache reads) is computed uniformly from the split fields: uncached input + output + cache read. The vendor `tokens` field is never used. Split fields use each harness's own classification and are NOT cross-comparable. Badge bases: unmetered / self-reported / proxy-measured.
| Harness | deepseek-v4-flash |
|---|---|
| claude | 72.1% |
| codex | 58.5% |
| grokbuild | 60.0% |
| opencode | 63.6% |
| pi | 85.3% |