OPENBENCH

OpenBench: kimi-k3 across 3 harnesses

1 models, 129 countable harness × task × trial cells.

kimi-k3

Non-comparable provenance: harness_version mixed within group; harness_version mixed within group; harness_version mixed within group; harness_version_source mixed within group; harness_version_source mixed within group; harness_version_source mixed within group. Inspect source rows before publishing.

Correctness

Correctness by harness with Wilson 95% confidence intervalsgrokbuild71.1%opencode68.9%pi64.1%0%20%40%60%80%100%
Bars show solve rate; whiskers show Wilson 95% confidence intervals.

Efficiency

Total tokens / solve incl. cache reads (log scale) by correctness for each harness0%20%40%60%80%100%112,692200,398356,363633,7131,126,919grokbuild71.1% · 356,363Correctness (higher →)Total tokens / solve incl. cache reads (log scale)
Lower-right is better: more correct and leaner. Diamonds identify cursor and opencode; their token totals are self-reported.

Speed

Median wall time by correctness for each harness0%20%40%60%80%100%89.4s105.8s122.1s138.5s154.8sopencode68.9% · 148.5sgrokbuild71.1% · 114.9spi64.1% · 95.8sCorrectness (higher →)Median wall time
Lower-right is better: more correct with a lower median wall time among solved runs.
Arm (harness × model)Solved/countableSolve rate @capSolve rate finishedWilson 95% CIMed wallTotal tokens/solve (incl. cache reads)Uncached in/solveOutput/solveCache-read/solveToken basisExcluded / unmatched / duplicates
grokbuild × kimi-k332/4571.1%84.2%56.6–82.3%114.9s356,36339,66316,672300,028self-reported0 / 0 / 0
opencode × kimi-k331/4568.9%81.6%54.3–80.5%148.5s44,07022,736498,143proxy-measured, self-reported0 / 0 / 0
pi × kimi-k325/3964.1%71.4%48.4–77.3%95.8s29,63315,229223,478proxy-measured, self-reported6 / 0 / 0

Total tokens/solve (incl. cache reads) is computed uniformly from the split fields: uncached input + output + cache read. The vendor `tokens` field is never used. Split fields use each harness's own classification and are NOT cross-comparable. Badge bases: unmetered / self-reported / proxy-measured.

Harness × model correctness

Harnesskimi-k3
grokbuild71.1%
opencode68.9%
pi64.1%

Methodology & limitations

Method

Limitations