OPENBENCH

OpenBench: GLM 5.2 quarantine-safe archive

45 retained rows from the archived five-harness matrix.

Archive quarantine

15 rows excluded: all 5 harnesses × 3 trials for terminal-bench/cancel-async-tasks. Its binding DROPPED.md identifies a load-sensitive checker, so those historical scores are not valid evidence.

This bundle was built without rerunning models. The source archive lacks execution-image digests, version-source labels, and an explicit timeout field; see README.md for the contemporaneous version and provenance caveats.

Comparison card

obench 0.1.0 · tasks: terminal-bench/feal-differential-cryptanalysis, terminal-bench/llm-inference-batching-scheduler, terminal-bench/schemelike-metacircular-eval · models: glm-5.2 · trials: 3

ArmSolved/nSolve rateWilson 95% CIMean scoreMean wallTokens/solveToken basis
pi × glm-5.28/8100.0%67.6–100.0%1.000660.2sself-reported
opencode × glm-5.27/977.8%45.3–93.7%0.778654.7s1,216,249self-reported
grokbuild × glm-5.25/955.6%26.7–81.1%0.556627.2sself-reported
claude × glm-5.25/955.6%26.7–81.1%0.556665.5sself-reported
codex × glm-5.24/944.4%18.9–73.3%0.444793.5sself-reported

Tokens/solve uses fresh totals: self-reported tokens when present, else proxy-measured (uncached input + output). Badge: unmetered / self-reported / proxy-measured. Cache-read is not folded into Tokens/solve.

glm-5.2

Correctness

Correctness by harness with Wilson 95% confidence intervalspi100.0%opencode77.8%claude55.6%grokbuild55.6%codex44.4%0%20%40%60%80%100%
Bars show solve rate; whiskers show Wilson 95% confidence intervals.

Efficiency

Total tokens / solve incl. cache reads (log scale) by correctness for each harness0%20%40%60%80%100%384,612683,9471,216,2492,162,8313,846,118opencode77.8% · 1,216,249Correctness (higher →)Total tokens / solve incl. cache reads (log scale)
Lower-right is better: more correct and leaner. Diamonds identify cursor and opencode; their token totals are self-reported.

Speed

Median wall time by correctness for each harness0%20%40%60%80%100%203.6s368.9s534.3s699.7s865.1scodex44.4% · 801.1spi100.0% · 693.5sopencode77.8% · 661.2sclaude55.6% · 569.0sgrokbuild55.6% · 267.6sCorrectness (higher →)Median wall time
Lower-right is better: more correct with a lower median wall time among solved runs.
Arm (harness × model)Solved/countableSolve rate @capSolve rate finishedWilson 95% CIMed wallTotal tokens/solve (incl. cache reads)Uncached in/solveOutput/solveCache-read/solveToken basisExcluded / unmatched / duplicates
pi × glm-5.28/8100.0%100.0%67.6–100.0%693.5sself-reported1 / 0 / 0
opencode × glm-5.27/977.8%100.0%45.3–93.7%661.2s1,216,24964,63546,3711,105,243self-reported0 / 0 / 0
grokbuild × glm-5.25/955.6%83.3%26.7–81.1%267.6sself-reported0 / 0 / 0
claude × glm-5.25/955.6%100.0%26.7–81.1%569.0sself-reported0 / 0 / 0
codex × glm-5.24/944.4%100.0%18.9–73.3%801.1sself-reported0 / 0 / 0

Total tokens/solve (incl. cache reads) is computed uniformly from the split fields: uncached input + output + cache read. The vendor `tokens` field is never used. Split fields use each harness's own classification and are NOT cross-comparable. Badge bases: unmetered / self-reported / proxy-measured.

Harness × model correctness

Harnessglm-5.2
claude55.6%
codex44.4%
grokbuild55.6%
opencode77.8%
pi100.0%

Methodology & limitations

What this card is

A self-contained OpenBench comparison bundle: candidate harness arm(s) highlighted against stock arms on the same result rows.

What `obench verify` proves

`provenance.json` (tamper-evident).

match under the bundle's digest scheme (scheme 2: instruction.md + checker.sh + workspace|workspace.toml + checker_data/; scheme 1 / legacy: same without checker_data/). Missing digests FAIL verification.

What verify does NOT prove

(transcripts are never included in the bundle).

Publish the full matrix whenever possible.