OpenBenchleaderboards

Coding-agent harness benchmarks

Compares coding-agent harnesses while holding the model and task fixed. Each board is one result-sealed bundle.

Scores are not comparable across bundles: task sets, trial counts, and timeout caps differ.

OpenBench: gpt-5.6 across 7 harnesses

kind releasedate 2026-07-21denominators matched (task, trial)matched rows 294common task/trials 42results SHA 546f167576aatask-set SHA d949d46e8aae
gpt-5.6-sol
HarnessSolve rate · Wilson 95%
050100%
SolvedMedian wallFresh tokens/solveUncached input/solveOutput/solveCache-read/solveCache-write/solveTelemetry source / basisTelemetry coverage
cursor
85.7%
72.2%–93.3%
36/4242.9s5,519725,446407,06745,776
nativeharness_reported
proxy 0/42 · native 42/42
devin
81.0%
66.7%–90.0%
34/4268.4s
unavailable
proxy 0/42 · native 0/42
grokbuild
81.0%
66.7%–90.0%
34/4286.3s46,21638,7277,489368,0790
proxyproxy_measured
proxy 42/42 · native 0/42
opencode
81.0%
66.7%–90.0%
34/4260.5s42,63436,7085,927313,9160
nativevendor_split
proxy 0/42 · native 42/42
claude
76.2%
61.5%–86.5%
32/4258.8s40,25534,4555,800153,4880
proxyproxy_measured
proxy 42/42 · native 42/42
pi
76.2%
61.5%–86.5%
32/4239.7s37,60232,9164,686111,4080
proxyproxy_measured
proxy 42/42 · native 42/42
codex
73.8%
58.9%–84.7%
31/4294.6s117,107102,17414,9331,164,5110
proxyproxy_measured
proxy 42/42 · native 0/42

OpenBench: deepseek across 5 harnesses

kind releasedate 2026-07-20denominators matched (task, trial)matched rows 190common task/trials 38results SHA ec6b9d58fd3atask-set SHA 1c4805e7fcd0
deepseek-v4-flashcaveats disclosed
4 caveat(s) from the release page
  • Archived first-party matrix published without rerunning models. The source contains 225 rows across 5 harnesses, 15 tasks, and 3 trials.
  • The bundled results SHA is sealed. Current-tree task checks report expected post-run definition drift; recorded task digests are retained as run-time fingerprints.
  • The unified board ranks only task/trial cells shared by every harness. The historical release page retains its older per-arm countable view, so its denominators differ.
  • Harness and host/container version stamps vary in the archive. Token columns use the complete counting-proxy lane for every displayed harness.
HarnessSolve rate · Wilson 95%
050100%
SolvedMedian wallFresh tokens/solveUncached input/solveOutput/solveCache-read/solveCache-write/solveTelemetry source / basisTelemetry coverage
claude
78.9%
63.7%–88.9%
30/3847.6s83,32335,23448,0892,053,6230
proxyproxy_measured
proxy 38/38 · native 37/38
pi
73.7%
58.0%–85.0%
28/3824.5s69,62437,98731,6381,729,6960
proxyproxy_measured
proxy 38/38 · native 35/38
opencode
68.4%
52.5%–80.9%
26/3840.3s71,31937,71133,6081,584,0200
proxyproxy_measured
proxy 38/38 · native 37/38
grokbuild
65.8%
49.9%–78.8%
25/3827.7s131,32781,79549,5321,745,7460
proxyproxy_measured
proxy 38/38 · native 0/38
codex
63.2%
47.3%–76.6%
24/3827.2s92,72340,35252,3721,212,5760
proxyproxy_measured
proxy 38/38 · native 33/38

OpenBench: grok-4.5 across 4 harnesses

kind releasedate 2026-07-20denominators matched (task, trial)matched rows 156common task/trials 39results SHA 3905ac339eeatask-set SHA 1c4805e7fcd0
grok-4.5caveats disclosed
4 caveat(s) from the release page
  • Archived first-party matrix published without rerunning models. The source contains 180 rows across 4 harnesses, 15 tasks, and 3 trials.
  • The bundled results SHA is sealed. Current-tree task checks report expected post-run definition drift; recorded task digests are retained as run-time fingerprints.
  • The unified board ranks only task/trial cells shared by every harness. The historical release page retains its older per-arm countable view, so its denominators differ.
  • No counting-proxy telemetry was retained for this run. Complete native split telemetry is available only where the board reports 100% native coverage.
HarnessSolve rate · Wilson 95%
050100%
SolvedMedian wallFresh tokens/solveUncached input/solveOutput/solveCache-read/solveCache-write/solveTelemetry source / basisTelemetry coverage
cursor
89.7%
76.4%–95.9%
35/3925.9s36,70728,6008,106247,5040
nativeharness_reported
proxy 0/39 · native 39/39
grokbuild
87.2%
73.3%–94.4%
34/3942.3s
unavailable
proxy 0/39 · native 0/39
pi
79.5%
64.5%–89.2%
31/3925.3s
unavailable
proxy 0/39 · native 38/39
opencode
76.9%
61.7%–87.4%
30/3921.2s55,00845,4949,514250,7350
nativevendor_split
proxy 0/39 · native 39/39

OpenBench showcase: BYO aider vs pi vs opencode (deepseek-v4-flash)

kind communitydate 2026-07-20denominators matched (task, trial)matched rows 24common task/trials 8results SHA 677642cc0649task-set SHA 8b757f8abf38
deepseek-v4-flashcaveats disclosed
4 caveat(s) from the release page
  • Harness mode is not apples-to-apples: aider ran as a one-shot --message invocation, while pi and opencode ran full agentic loops. Token comparison reflects that difference in harness mode, not a pure like-for-like agent loop contest.
  • One aider cell excluded: 1 of 9 aider cells was infrastructure-classified and is excluded from solve-rate denominators (aider reported as 8/8). pi and opencode are 9/9.
  • Uniform total-token basis: totals include uncached input, output, and cache reads from split fields; the vendor aggregate is never used. Values are 6,192 vs 48,262 vs 111,380 tokens/solve (aider vs pi vs opencode).
  • Verifiable bundle: this release includes results.jsonl and provenance.json. Re-check digests with obench verify docs/releases/2026-07-20-aider-showcase.
HarnessSolve rate · Wilson 95%
050100%
SolvedMedian wallFresh tokens/solveUncached input/solveOutput/solveCache-read/solveCache-write/solveTelemetry source / basisTelemetry coverage
aider
100.0%
67.6%–100.0%
8/839.0s5,8562,3083,5483360
proxyproxy_measured
proxy 8/8 · native 0/8
opencode
100.0%
67.6%–100.0%
8/831.7s13,66910,3283,34291,6000
proxyproxy_measured
proxy 8/8 · native 8/8
pi
100.0%
67.6%–100.0%
8/826.4s7,7424,9282,81435,4560
proxyproxy_measured
proxy 8/8 · native 8/8

OpenBench: GLM 5.2 quarantine-safe archive

kind releasedate 2026-07-09denominators matched (task, trial)matched rows 40common task/trials 8results SHA 5434f3f8aa66task-set SHA 46075fa27b81
glm-5.2caveats disclosed
4 caveat(s) from the release page
  • Quarantine-safe derived bundle: all 15 terminal-bench/cancel-async-tasks rows were excluded because its load-sensitive checker is binding-quarantined. No model was rerun.
  • The bundled results SHA is sealed. Current-tree task checks report expected post-run definition drift; recorded task digests are retained as run-time fingerprints.
  • The retained archive lacks image digests, version-source labels, and an explicit per-row timeout field; contemporaneous version-stamp drift is documented in the bundle README.
  • The unified board ranks only task/trial cells shared by every harness. The release page also shows the per-arm countable view, so its denominators differ.
HarnessSolve rate · Wilson 95%
050100%
SolvedMedian wallFresh tokens/solveUncached input/solveOutput/solveCache-read/solveCache-write/solveTelemetry source / basisTelemetry coverage
pi
100.0%
67.6%–100.0%
8/8693.5s
unavailable
proxy 0/8 · native 6/8
opencode
87.5%
52.9%–97.8%
7/8661.2s94,19754,76239,435973,1660
nativevendor_split
proxy 0/8 · native 8/8
claude
62.5%
30.6%–86.3%
5/8569.0s
unavailable
proxy 0/8 · native 4/8
grokbuild
62.5%
30.6%–86.3%
5/8267.6s
unavailable
proxy 0/8 · native 0/8
codex
50.0%
21.5%–78.5%
4/8801.1s
unavailable
proxy 0/8 · native 4/8

Not ranked (2)

No result-sealed results.jsonl.

  • 2026-07-02-m3
    no results.jsonl (HTML-only release page)
  • 2026-07-20-kimi-k3
    no results.jsonl (HTML-only release page)

Managed AI gateway benchmarks

Compares request latency, throughput, and reliability across managed AI gateways under separately scheduled cold and warm conditions. Run facts below describe the latest published bundle; every board keeps separate denominators.

Gateway Bench measures request-level transport and serving telemetry. Cold and warm denominators are separate and are never merged.

GPT-4o mini managed gateway request benchmark

date 2026-07-27model match rolling_aliascold blocks 30/30warm blocks 30/30requests 300verified commit 2cd722da4d65experiment e3649dbc6dc4

Request-level benchmark. Cold and warm conditions retain separate denominators from Harness Bench cells.

Gateway leaderboard

OpenBench Composite: absolute TTFT latency, output throughput, and request success on a 0–100 scale. Higher is better; cost is excluded and Direct OpenAI is an unranked reference.

1
OpenRouter
100% request success
92.8composite
2
Cloudflare
100% request success
91.5composite
3
Concentrate
100% request success
90.4composite
4
Vercel
100% request success
89.4composite

Cold requests

Complete blocks: 30/30. TTFT includes DNS, TCP, and TLS setup.

RouteTTFT p50 / p95Stream total p50 / p95Response headers p50 / p95First body byte p50 / p95Throughput tok/s p50 / p95Total / cached / cache-write tokens p50 / p95
Direct OpenAI0.556s / 1.384s0.735s / 1.763s0.232s / 0.887s0.232s / 0.888s92.3 / 123.2total 51.0 / 54.0
cached 0.0 / 0.0
cache write 0.0 / 0.0
Cloudflare0.684s / 1.091s0.812s / 1.215s0.341s / 0.767s0.341s / 0.767s84.6 / 113.7total 51.0 / 53.5
cached 0.0 / 0.0
cache write 0.0 / 0.0
Concentrate0.794s / 0.978s0.930s / 1.803s0.404s / 0.538s0.404s / 0.538s84.9 / 123.0total 51.0 / 54.0
cached 0.0 / 0.0
cache write 0.0 / 0.0
OpenRouter0.501s / 0.703s0.662s / 0.845s0.459s / 0.668s0.470s / 0.668s100.6 / 156.9total 54.0 / 56.5
cached 0.0 / 0.0
cache write — / — (0/30)
Vercel0.842s / 2.528s0.951s / 2.719s0.773s / 2.485s0.783s / 2.490s87.9 / 345.3total 51.0 / 54.0
cached 0.0 / 0.0
cache write 0.0 / 0.0

Warm requests

Complete blocks: 30/30. TTFT begins when the measured request is sent.

RouteTTFT p50 / p95Stream total p50 / p95Response headers p50 / p95First body byte p50 / p95Throughput tok/s p50 / p95Total / cached / cache-write tokens p50 / p95
Direct OpenAI0.504s / 0.977s0.656s / 1.221s0.184s / 0.436s0.184s / 0.437s95.1 / 127.2total 52.0 / 54.5
cached 0.0 / 0.0
cache write 0.0 / 0.0
Cloudflare0.561s / 0.739s0.754s / 0.893s0.260s / 0.410s0.260s / 0.410s96.6 / 126.4total 52.0 / 54.5
cached 0.0 / 0.0
cache write 0.0 / 0.0
Concentrate0.704s / 0.933s0.896s / 1.096s0.361s / 0.436s0.361s / 0.436s85.3 / 131.6total 52.0 / 54.5
cached 0.0 / 0.0
cache write 0.0 / 0.0
OpenRouter0.433s / 0.655s0.587s / 0.764s0.423s / 0.485s0.432s / 0.632s105.9 / 415.5total 54.5 / 57.0
cached 0.0 / 0.0
cache write — / — (0/30)
Vercel0.572s / 1.605s0.726s / 2.458s0.559s / 1.571s0.572s / 1.601s94.4 / 449.2total 52.5 / 55.0
cached 0.0 / 0.0
cache write 0.0 / 0.0

Cold setup

Connection setup phases for cold requests only.

RouteDNS p50 / p95TCP p50 / p95TLS p50 / p95
Direct OpenAI0.002s / 0.003s0.011s / 0.012s0.016s / 0.023s
Cloudflare0.001s / 0.002s0.011s / 0.012s0.017s / 0.023s
Concentrate0.001s / 0.002s0.011s / 0.014s0.018s / 0.023s
OpenRouter0.001s / 0.002s0.011s / 0.012s0.016s / 0.020s
Vercel0.002s / 0.148s0.006s / 0.007s0.024s / 0.028s

Paired request deltas

Every delta is gateway minus Direct OpenAI. These are latency metrics, so positive means slower/worse and negative means faster/better. Medians use complete paired blocks with bootstrap 95% intervals.

Gateway routeConditionΔ response headersΔ TTFT
Cloudflarecold
+0.090s
95% CI +0.072s to +0.120s · paired 30/30
+0.095s
95% CI +0.011s to +0.163s · paired 30/30
Cloudflarewarm
+0.081s
95% CI +0.050s to +0.112s · paired 30/30
+0.037s
95% CI +0.003s to +0.116s · paired 30/30
Concentratecold
+0.169s
95% CI +0.143s to +0.190s · paired 30/30
+0.193s
95% CI +0.097s to +0.235s · paired 30/30
Concentratewarm
+0.168s
95% CI +0.143s to +0.204s · paired 30/30
+0.202s
95% CI +0.142s to +0.230s · paired 30/30
OpenRoutercold
+0.233s
95% CI +0.212s to +0.244s · paired 30/30
-0.088s
95% CI -0.176s to -0.009s · paired 30/30
OpenRouterwarm
+0.230s
95% CI +0.202s to +0.252s · paired 30/30
-0.054s
95% CI -0.109s to -0.018s · paired 30/30
Vercelcold
+0.508s
95% CI +0.440s to +0.596s · paired 30/30
+0.227s
95% CI +0.171s to +0.297s · paired 30/30
Vercelwarm
+0.348s
95% CI +0.328s to +0.465s · paired 30/30
+0.058s
95% CI +0.014s to +0.214s · paired 30/30

OpenBench benchmark results

Digest-verified benchmark releases, methodology, and project information.

Releases

First-party bundles.

Packs

Versioned task and harness packs.

  • openbench/core-smoke@1.0.0 tasks
    Apache-2.0 · data/packs/openbench-core-smoke · 3b1e576398a8
    Tiny polarity-checked smoke tasks (make-it-run, fix-failing-test)

OpenBench benchmark results

Digest-verified benchmark releases, methodology, and project information.

What is being measured

OpenBench runs two benchmark families. They share a task contract and a checker, and no denominators.

Harness Bench

Varies the coding-agent harness — the CLI that wraps a model in a run loop, tool set, and permission policy — while holding the model and task fixed. An arm is (harness, model). A task is solved when its checker.sh exits 0; the harness's own claim of success is never trusted.

Gateway Bench

Measures one model request at a time under separately scheduled cold and warm transport conditions. It reports request success, route verification, transport and stream phase timing, throughput, usage, and per-request cost. It is not a coding-agent outcome benchmark. Gateway Bench requests and Harness Bench cells are never pooled or compared as one denominator.

The Gateway Bench leaderboard uses an absolute summary score: 30% cold TTFT median, 15% cold TTFT p95, 30% warm TTFT median, 15% warm TTFT p95, and 10% warm median output throughput. Cold TTFT includes DNS, TCP, and TLS. Latency scores linearly from 100 at zero to zero at 20 seconds; throughput scores linearly from zero at 5 tok/s to 100 at 200 tok/s. The weighted result is multiplied by request success. Cost is excluded. Direct OpenAI is an unranked reference, and the detailed measurements below the score remain the factual record.

Denominators and intervals

  • Denominators are countable cells. Infrastructure and rate-limit failures are excluded; other failures, including timeouts, stay in the denominator.
  • Harness Bench uses Wilson 95% intervals over matched (task, trial) cells whenever a bundle has two or more arms.
  • Gateway Bench displays complete cold and warm block counts separately. Availability uses a Wilson 95% interval over attempted requests. Phase summaries use successful, route-verified requests and retain metric-specific coverage. Paired deltas use complete gateway/direct blocks and bootstrap 95% intervals.

Efficiency and cost

  • Median wall time is taken among solved cells only.
  • Each Harness Bench arm uses one complete split-token lane across all matched result rows, preferring proxy telemetry and otherwise using native telemetry. Fresh tokens are uncached input plus output; cache reads and cache writes remain separate. Incomplete lanes produce no token metrics and report their row coverage instead.
  • Each per-solve token figure sums traffic from every matched attempt, including failed attempts, then divides by the number of solved cells. It measures attempted traffic required per solve, not the average size of successful attempts alone.
  • Harness $/solve appears only for models with a configured price.
  • Gateway Bench response headers are the time until HTTP response headers. First body byte and semantic TTFT are reported separately; response headers are not labeled TTFB.
  • Gateway Bench measured cost is the frozen-list request estimate. Charged cost is separately reported billing evidence. Each retains its own request coverage, as do total, cached-input, and cache-write token readings.
  • Harness defaults are not clamped.

Comparability

  • Cells from different bundles are never blended. Each board is one bundle; cross-bundle ranking on different task sets is not supported.
  • Every ranked bundle ships results.jsonl plus a provenance digest and is re-verified before it appears here. Digests show tamper-evidence, not absence of cherry-picking.
  • Results cover only the included tasks, trials, model deployments, harness versions, and timeout caps.

Reproducing a board

Every board links its results.jsonl. Re-check a bundle with obench verify <bundle> (harness) or obench gateway probe verify <bundle> (gateway), and rebuild this page with obench site build.

OpenBench benchmark results

Digest-verified benchmark releases, methodology, and project information.

Contact

Want to add a gateway or harness, submit results, report a problem, or share an idea? Reach Matthew through either channel.