OPENBENCH

OpenBench: deepseek across 5 harnesses

1 models, 207 countable harness × task × trial cells.

deepseek-v4-flash

Non-comparable provenance: harness_version mixed within group; harness_version mixed within group; harness_version mixed within group; harness_version mixed within group; harness_version mixed within group; harness_version_source mixed within group; harness_version_source mixed within group; harness_version_source mixed within group; harness_version_source mixed within group; harness_version_source mixed within group. Inspect source rows before publishing.

Correctness

Correctness by harness with Wilson 95% confidence intervalspi85.3%claude72.1%opencode63.6%grokbuild60.0%codex58.5%0%20%40%60%80%100%
Bars show solve rate; whiskers show Wilson 95% confidence intervals.

Efficiency

Total tokens / solve incl. cache reads (log scale) by correctness for each harness0%20%40%60%80%100%713,6431,269,0572,256,7394,013,1127,136,435grokbuild60.0% · 2,256,739Correctness (higher →)Total tokens / solve incl. cache reads (log scale)
Lower-right is better: more correct and leaner. Diamonds identify cursor and opencode; their token totals are self-reported.

Speed

Median wall time by correctness for each harness0%20%40%60%80%100%21.7s29.1s36.5s43.9s51.3sclaude72.1% · 48.4sopencode63.6% · 45.8sgrokbuild60.0% · 35.9scodex58.5% · 27.2spi85.3% · 24.5sCorrectness (higher →)Median wall time
Lower-right is better: more correct with a lower median wall time among solved runs.
Arm (harness × model)Solved/countableSolve rate @capSolve rate finishedWilson 95% CIMed wallTotal tokens/solve (incl. cache reads)Uncached in/solveOutput/solveCache-read/solveToken basisExcluded / unmatched / duplicates
pi × deepseek-v4-flash29/3485.3%87.9%69.9–93.6%24.5s34,32829,1041,712,945proxy-measured, self-reported11 / 0 / 0
claude × deepseek-v4-flash31/4372.1%72.1%57.3–83.3%48.4s39,92756,5382,443,483proxy-measured, self-reported2 / 0 / 0
opencode × deepseek-v4-flash28/4463.6%63.6%48.9–76.2%45.8s44,64943,5081,904,229proxy-measured, self-reported1 / 0 / 0
grokbuild × deepseek-v4-flash27/4560.0%64.3%45.5–73.0%35.9s2,256,73951,63761,7842,143,317self-reported0 / 0 / 0
codex × deepseek-v4-flash24/4158.5%66.7%43.4–72.2%27.2s43,93563,0451,257,808proxy-measured, self-reported4 / 0 / 0

Total tokens/solve (incl. cache reads) is computed uniformly from the split fields: uncached input + output + cache read. The vendor `tokens` field is never used. Split fields use each harness's own classification and are NOT cross-comparable. Badge bases: unmetered / self-reported / proxy-measured.

Harness × model correctness

Harnessdeepseek-v4-flash
claude72.1%
codex58.5%
grokbuild60.0%
opencode63.6%
pi85.3%

Methodology & limitations

Method

Limitations