Coding-agent harness benchmarks
Compares coding-agent harnesses while holding the model and task fixed. Each board is one result-sealed bundle.
Scores are not comparable across bundles: task sets, trial counts, and timeout caps differ.
OpenBench: gpt-5.6 across 7 harnesses
| Harness↕ | Solve rate · Wilson 95%↓ 050100% | Solved↕ | Median wall↕ | Fresh tokens/solve↕ | Uncached input/solve↕ | Output/solve↕ | Cache-read/solve↕ | Cache-write/solve↕ | Telemetry source / basis | Telemetry coverage |
|---|---|---|---|---|---|---|---|---|---|---|
| cursor | 85.7% 72.2%–93.3% | 36/42 | 42.9s | 5,519 | 72 | 5,446 | 407,067 | 45,776 | nativeharness_reported | proxy 0/42 · native 42/42 |
| devin | 81.0% 66.7%–90.0% | 34/42 | 68.4s | — | — | — | — | — | unavailable | proxy 0/42 · native 0/42 |
| grokbuild | 81.0% 66.7%–90.0% | 34/42 | 86.3s | 46,216 | 38,727 | 7,489 | 368,079 | 0 | proxyproxy_measured | proxy 42/42 · native 0/42 |
| opencode | 81.0% 66.7%–90.0% | 34/42 | 60.5s | 42,634 | 36,708 | 5,927 | 313,916 | 0 | nativevendor_split | proxy 0/42 · native 42/42 |
| claude | 76.2% 61.5%–86.5% | 32/42 | 58.8s | 40,255 | 34,455 | 5,800 | 153,488 | 0 | proxyproxy_measured | proxy 42/42 · native 42/42 |
| pi | 76.2% 61.5%–86.5% | 32/42 | 39.7s | 37,602 | 32,916 | 4,686 | 111,408 | 0 | proxyproxy_measured | proxy 42/42 · native 42/42 |
| codex | 73.8% 58.9%–84.7% | 31/42 | 94.6s | 117,107 | 102,174 | 14,933 | 1,164,511 | 0 | proxyproxy_measured | proxy 42/42 · native 0/42 |
OpenBench: deepseek across 5 harnesses
4 caveat(s) from the release page
- Archived first-party matrix published without rerunning models. The source contains 225 rows across 5 harnesses, 15 tasks, and 3 trials.
- The bundled results SHA is sealed. Current-tree task checks report expected post-run definition drift; recorded task digests are retained as run-time fingerprints.
- The unified board ranks only task/trial cells shared by every harness. The historical release page retains its older per-arm countable view, so its denominators differ.
- Harness and host/container version stamps vary in the archive. Token columns use the complete counting-proxy lane for every displayed harness.
| Harness↕ | Solve rate · Wilson 95%↓ 050100% | Solved↕ | Median wall↕ | Fresh tokens/solve↕ | Uncached input/solve↕ | Output/solve↕ | Cache-read/solve↕ | Cache-write/solve↕ | Telemetry source / basis | Telemetry coverage |
|---|---|---|---|---|---|---|---|---|---|---|
| claude | 78.9% 63.7%–88.9% | 30/38 | 47.6s | 83,323 | 35,234 | 48,089 | 2,053,623 | 0 | proxyproxy_measured | proxy 38/38 · native 37/38 |
| pi | 73.7% 58.0%–85.0% | 28/38 | 24.5s | 69,624 | 37,987 | 31,638 | 1,729,696 | 0 | proxyproxy_measured | proxy 38/38 · native 35/38 |
| opencode | 68.4% 52.5%–80.9% | 26/38 | 40.3s | 71,319 | 37,711 | 33,608 | 1,584,020 | 0 | proxyproxy_measured | proxy 38/38 · native 37/38 |
| grokbuild | 65.8% 49.9%–78.8% | 25/38 | 27.7s | 131,327 | 81,795 | 49,532 | 1,745,746 | 0 | proxyproxy_measured | proxy 38/38 · native 0/38 |
| codex | 63.2% 47.3%–76.6% | 24/38 | 27.2s | 92,723 | 40,352 | 52,372 | 1,212,576 | 0 | proxyproxy_measured | proxy 38/38 · native 33/38 |
OpenBench: grok-4.5 across 4 harnesses
4 caveat(s) from the release page
- Archived first-party matrix published without rerunning models. The source contains 180 rows across 4 harnesses, 15 tasks, and 3 trials.
- The bundled results SHA is sealed. Current-tree task checks report expected post-run definition drift; recorded task digests are retained as run-time fingerprints.
- The unified board ranks only task/trial cells shared by every harness. The historical release page retains its older per-arm countable view, so its denominators differ.
- No counting-proxy telemetry was retained for this run. Complete native split telemetry is available only where the board reports 100% native coverage.
| Harness↕ | Solve rate · Wilson 95%↓ 050100% | Solved↕ | Median wall↕ | Fresh tokens/solve↕ | Uncached input/solve↕ | Output/solve↕ | Cache-read/solve↕ | Cache-write/solve↕ | Telemetry source / basis | Telemetry coverage |
|---|---|---|---|---|---|---|---|---|---|---|
| cursor | 89.7% 76.4%–95.9% | 35/39 | 25.9s | 36,707 | 28,600 | 8,106 | 247,504 | 0 | nativeharness_reported | proxy 0/39 · native 39/39 |
| grokbuild | 87.2% 73.3%–94.4% | 34/39 | 42.3s | — | — | — | — | — | unavailable | proxy 0/39 · native 0/39 |
| pi | 79.5% 64.5%–89.2% | 31/39 | 25.3s | — | — | — | — | — | unavailable | proxy 0/39 · native 38/39 |
| opencode | 76.9% 61.7%–87.4% | 30/39 | 21.2s | 55,008 | 45,494 | 9,514 | 250,735 | 0 | nativevendor_split | proxy 0/39 · native 39/39 |
OpenBench showcase: BYO aider vs pi vs opencode (deepseek-v4-flash)
4 caveat(s) from the release page
- Harness mode is not apples-to-apples: aider ran as a one-shot --message invocation, while pi and opencode ran full agentic loops. Token comparison reflects that difference in harness mode, not a pure like-for-like agent loop contest.
- One aider cell excluded: 1 of 9 aider cells was infrastructure-classified and is excluded from solve-rate denominators (aider reported as 8/8). pi and opencode are 9/9.
- Uniform total-token basis: totals include uncached input, output, and cache reads from split fields; the vendor aggregate is never used. Values are 6,192 vs 48,262 vs 111,380 tokens/solve (aider vs pi vs opencode).
- Verifiable bundle: this release includes results.jsonl and provenance.json. Re-check digests with obench verify docs/releases/2026-07-20-aider-showcase.
| Harness↕ | Solve rate · Wilson 95%↓ 050100% | Solved↕ | Median wall↕ | Fresh tokens/solve↕ | Uncached input/solve↕ | Output/solve↕ | Cache-read/solve↕ | Cache-write/solve↕ | Telemetry source / basis | Telemetry coverage |
|---|---|---|---|---|---|---|---|---|---|---|
| aider | 100.0% 67.6%–100.0% | 8/8 | 39.0s | 5,856 | 2,308 | 3,548 | 336 | 0 | proxyproxy_measured | proxy 8/8 · native 0/8 |
| opencode | 100.0% 67.6%–100.0% | 8/8 | 31.7s | 13,669 | 10,328 | 3,342 | 91,600 | 0 | proxyproxy_measured | proxy 8/8 · native 8/8 |
| pi | 100.0% 67.6%–100.0% | 8/8 | 26.4s | 7,742 | 4,928 | 2,814 | 35,456 | 0 | proxyproxy_measured | proxy 8/8 · native 8/8 |
OpenBench: GLM 5.2 quarantine-safe archive
4 caveat(s) from the release page
- Quarantine-safe derived bundle: all 15 terminal-bench/cancel-async-tasks rows were excluded because its load-sensitive checker is binding-quarantined. No model was rerun.
- The bundled results SHA is sealed. Current-tree task checks report expected post-run definition drift; recorded task digests are retained as run-time fingerprints.
- The retained archive lacks image digests, version-source labels, and an explicit per-row timeout field; contemporaneous version-stamp drift is documented in the bundle README.
- The unified board ranks only task/trial cells shared by every harness. The release page also shows the per-arm countable view, so its denominators differ.
| Harness↕ | Solve rate · Wilson 95%↓ 050100% | Solved↕ | Median wall↕ | Fresh tokens/solve↕ | Uncached input/solve↕ | Output/solve↕ | Cache-read/solve↕ | Cache-write/solve↕ | Telemetry source / basis | Telemetry coverage |
|---|---|---|---|---|---|---|---|---|---|---|
| pi | 100.0% 67.6%–100.0% | 8/8 | 693.5s | — | — | — | — | — | unavailable | proxy 0/8 · native 6/8 |
| opencode | 87.5% 52.9%–97.8% | 7/8 | 661.2s | 94,197 | 54,762 | 39,435 | 973,166 | 0 | nativevendor_split | proxy 0/8 · native 8/8 |
| claude | 62.5% 30.6%–86.3% | 5/8 | 569.0s | — | — | — | — | — | unavailable | proxy 0/8 · native 4/8 |
| grokbuild | 62.5% 30.6%–86.3% | 5/8 | 267.6s | — | — | — | — | — | unavailable | proxy 0/8 · native 0/8 |
| codex | 50.0% 21.5%–78.5% | 4/8 | 801.1s | — | — | — | — | — | unavailable | proxy 0/8 · native 4/8 |
No boards match the current filters.
Not ranked (2)
No result-sealed results.jsonl.
- 2026-07-02-m3no results.jsonl (HTML-only release page)
- 2026-07-20-kimi-k3no results.jsonl (HTML-only release page)