OPENBENCH

OpenBench: grok-4.5 across 4 harnesses

1 models, 171 countable harness × task × trial cells.

grok-4.5

Non-comparable provenance: image_digest differs across groups; harness_version_source mixed within group; harness_version_source mixed within group; harness_version_source mixed within group; timeout_s differs across groups. Inspect source rows before publishing.

Correctness

Correctness by harness with Wilson 95% confidence intervalscursor91.1%grokbuild88.9%pi79.1%opencode78.9%0%20%40%60%80%100%
Bars show solve rate; whiskers show Wilson 95% confidence intervals.

Efficiency

Total tokens / solve incl. cache reads (log scale) by correctness for each harness0%20%40%60%80%100%252,426321,814410,277523,056666,838grokbuild88.9% · 607,005opencode78.9% · 277,307Correctness (higher →)Total tokens / solve incl. cache reads (log scale)
Lower-right is better: more correct and leaner. Diamonds identify cursor and opencode; their token totals are self-reported.

Speed

Median wall time by correctness for each harness0%20%40%60%80%100%16.0s29.5s42.9s56.4s69.8sgrokbuild88.9% · 64.6scursor91.1% · 35.5spi79.1% · 35.4sopencode78.9% · 21.2sCorrectness (higher →)Median wall time
Lower-right is better: more correct with a lower median wall time among solved runs.
Arm (harness × model)Solved/countableSolve rateWilson 95% CIMed wallTotal tokens/solve (incl. cache reads)Uncached in/solveOutput/solveCache-read/solveToken basisExcluded / unmatched / duplicates
cursor × grok-4.541/4591.1%79.3–96.5%35.5sself-reported, unknown0 / 0 / 0
grokbuild × grok-4.540/4588.9%76.5–95.2%64.6s607,00559,68313,492533,830self-reported0 / 0 / 0
pi × grok-4.534/4379.1%64.8–88.6%35.4sself-reported, unknown2 / 0 / 0
opencode × grok-4.530/3878.9%63.7–88.9%21.2s277,30743,6929,138224,478self-reported7 / 0 / 0

Timeouts: pi: 1 timeout; finished-basis 81.0%.

Total tokens/solve (incl. cache reads) is computed uniformly from the split fields: uncached input + output + cache read. The vendor `tokens` field is never used. Split fields use each harness's own classification and are NOT cross-comparable. Badge bases: unmetered / self-reported / proxy-measured.

Harness × model correctness

Harnessgrok-4.5
cursor91.1%
grokbuild88.9%
opencode78.9%
pi79.1%

Methodology & limitations

Method

Limitations