1 models, 129 countable harness × task × trial cells.
Non-comparable provenance: harness_version mixed within group; harness_version mixed within group; harness_version mixed within group; harness_version_source mixed within group; harness_version_source mixed within group; harness_version_source mixed within group. Inspect source rows before publishing.
| Arm (harness × model) | Solved/countable | Solve rate @cap | Solve rate finished | Wilson 95% CI | Med wall | Total tokens/solve (incl. cache reads) | Uncached in/solve | Output/solve | Cache-read/solve | Token basis | Excluded / unmatched / duplicates |
|---|---|---|---|---|---|---|---|---|---|---|---|
| grokbuild × kimi-k3 | 32/45 | 71.1% | 84.2% | 56.6–82.3% | 114.9s | 356,363 | 39,663 | 16,672 | 300,028 | self-reported | 0 / 0 / 0 |
| opencode × kimi-k3 | 31/45 | 68.9% | 81.6% | 54.3–80.5% | 148.5s | — | 44,070 | 22,736 | 498,143 | proxy-measured, self-reported | 0 / 0 / 0 |
| pi × kimi-k3 | 25/39 | 64.1% | 71.4% | 48.4–77.3% | 95.8s | — | 29,633 | 15,229 | 223,478 | proxy-measured, self-reported | 6 / 0 / 0 |
Total tokens/solve (incl. cache reads) is computed uniformly from the split fields: uncached input + output + cache read. The vendor `tokens` field is never used. Split fields use each harness's own classification and are NOT cross-comparable. Badge bases: unmetered / self-reported / proxy-measured.
| Harness | kimi-k3 |
|---|---|
| grokbuild | 71.1% |
| opencode | 68.9% |
| pi | 64.1% |