45 retained rows from the archived five-harness matrix.
15 rows excluded: all 5 harnesses × 3 trials for terminal-bench/cancel-async-tasks. Its binding DROPPED.md identifies a load-sensitive checker, so those historical scores are not valid evidence.
This bundle was built without rerunning models. The source archive lacks execution-image digests, version-source labels, and an explicit timeout field; see README.md for the contemporaneous version and provenance caveats.
obench 0.1.0 · tasks: terminal-bench/feal-differential-cryptanalysis, terminal-bench/llm-inference-batching-scheduler, terminal-bench/schemelike-metacircular-eval · models: glm-5.2 · trials: 3
| Arm | Solved/n | Solve rate | Wilson 95% CI | Mean score | Mean wall | Tokens/solve | Token basis |
|---|---|---|---|---|---|---|---|
| pi × glm-5.2 | 8/8 | 100.0% | 67.6–100.0% | 1.000 | 660.2s | — | self-reported |
| opencode × glm-5.2 | 7/9 | 77.8% | 45.3–93.7% | 0.778 | 654.7s | 1,216,249 | self-reported |
| grokbuild × glm-5.2 | 5/9 | 55.6% | 26.7–81.1% | 0.556 | 627.2s | — | self-reported |
| claude × glm-5.2 | 5/9 | 55.6% | 26.7–81.1% | 0.556 | 665.5s | — | self-reported |
| codex × glm-5.2 | 4/9 | 44.4% | 18.9–73.3% | 0.444 | 793.5s | — | self-reported |
Tokens/solve uses fresh totals: self-reported tokens when present, else proxy-measured (uncached input + output). Badge: unmetered / self-reported / proxy-measured. Cache-read is not folded into Tokens/solve.
| Arm (harness × model) | Solved/countable | Solve rate @cap | Solve rate finished | Wilson 95% CI | Med wall | Total tokens/solve (incl. cache reads) | Uncached in/solve | Output/solve | Cache-read/solve | Token basis | Excluded / unmatched / duplicates |
|---|---|---|---|---|---|---|---|---|---|---|---|
| pi × glm-5.2 | 8/8 | 100.0% | 100.0% | 67.6–100.0% | 693.5s | — | — | — | — | self-reported | 1 / 0 / 0 |
| opencode × glm-5.2 | 7/9 | 77.8% | 100.0% | 45.3–93.7% | 661.2s | 1,216,249 | 64,635 | 46,371 | 1,105,243 | self-reported | 0 / 0 / 0 |
| grokbuild × glm-5.2 | 5/9 | 55.6% | 83.3% | 26.7–81.1% | 267.6s | — | — | — | — | self-reported | 0 / 0 / 0 |
| claude × glm-5.2 | 5/9 | 55.6% | 100.0% | 26.7–81.1% | 569.0s | — | — | — | — | self-reported | 0 / 0 / 0 |
| codex × glm-5.2 | 4/9 | 44.4% | 100.0% | 18.9–73.3% | 801.1s | — | — | — | — | self-reported | 0 / 0 / 0 |
Total tokens/solve (incl. cache reads) is computed uniformly from the split fields: uncached input + output + cache read. The vendor `tokens` field is never used. Split fields use each harness's own classification and are NOT cross-comparable. Badge bases: unmetered / self-reported / proxy-measured.
| Harness | glm-5.2 |
|---|---|
| claude | 55.6% |
| codex | 44.4% |
| grokbuild | 55.6% |
| opencode | 77.8% |
| pi | 100.0% |
A self-contained OpenBench comparison bundle: candidate harness arm(s) highlighted against stock arms on the same result rows.
`provenance.json` (tamper-evident).
match under the bundle's digest scheme (scheme 2: instruction.md + checker.sh + workspace|workspace.toml + checker_data/; scheme 1 / legacy: same without checker_data/). Missing digests FAIL verification.
(transcripts are never included in the bundle).
Publish the full matrix whenever possible.