OPENBENCH

OpenBench: gpt-5.6 across 7 harnesses

1 models, 311 countable harness × task × trial cells.

gpt-5.6-sol

Provenance note: image_digest differs across groups; harness_version_source mixed within group; harness_version_source mixed within group; harness_version_source mixed within group; harness_version_source mixed within group; harness_version_source mixed within group — a container-rebuild / host-vs-container stamping artifact, not a CLI version change. Pinned CLI versions are identical across arms.

Correctness

Correctness by harness with Wilson 95% confidence intervalscursor86.0%grokbuild82.2%devin80.0%opencode80.0%claude77.8%codex72.7%pi72.7%0%20%40%60%80%100%
Bars show solve rate; whiskers show Wilson 95% confidence intervals.

Efficiency

Total tokens / solve incl. cache reads (log scale) by correctness for each harness0%20%40%60%80%100%130,877183,787258,085362,421508,935cursor86.0% · 446,257grokbuild82.2% · 410,474opencode80.0% · 362,976claude77.8% · 194,532pi72.7% · 149,259Correctness (higher →)Total tokens / solve incl. cache reads (log scale)
Lower-right is better: more correct and leaner. Diamonds identify cursor and opencode; their token totals are self-reported.

Speed

Median wall time by correctness for each harness0%20%40%60%80%100%32.9s50.4s67.8s85.2s102.7scodex72.7% · 95.9sgrokbuild82.2% · 89.4sdevin80.0% · 74.3sopencode80.0% · 61.5sclaude77.8% · 60.8scursor86.0% · 43.3spi72.7% · 39.7sCorrectness (higher →)Median wall time
Lower-right is better: more correct with a lower median wall time among solved runs.
Arm (harness × model)Solved/countableSolve rateWilson 95% CIMed wallTotal tokens/solve (incl. cache reads)Uncached in/solveOutput/solveCache-read/solveToken basisExcluded / unmatched / duplicates
cursor × gpt-5.6-sol37/4386.0%72.7–93.4%43.3s446,257735,738440,446self-reported2 / 0 / 0
grokbuild × gpt-5.6-sol37/4582.2%68.7–90.7%89.4s410,47436,7547,488366,232self-reported0 / 0 / 0
opencode × gpt-5.6-sol36/4580.0%66.2–89.1%61.5s362,97636,5306,148320,299self-reported0 / 0 / 0
devin × gpt-5.6-sol36/4580.0%66.2–89.1%74.3sself-reported0 / 0 / 0
claude × gpt-5.6-sol35/4577.8%63.7–87.5%60.8s194,53233,4615,833155,238self-reported0 / 0 / 0
pi × gpt-5.6-sol32/4472.7%58.2–83.7%39.7s149,25933,1184,733111,408self-reported1 / 0 / 0
codex × gpt-5.6-sol32/4472.7%58.2–83.7%95.9s68,16810,877929,384proxy-measured, self-reported1 / 0 / 0

Timeouts: codex: 1 timeout; finished-basis 74.4%.

Total tokens/solve (incl. cache reads) is computed uniformly from the split fields: uncached input + output + cache read. The vendor `tokens` field is never used. Split fields use each harness's own classification and are NOT cross-comparable. Badge bases: unmetered / self-reported / proxy-measured.

Harness × model correctness

Harnessgpt-5.6-sol
claude77.8%
codex72.7%
cursor86.0%
devin80.0%
grokbuild82.2%
opencode80.0%
pi72.7%

Methodology & limitations

Method

Limitations