Reliability: runs scoring at least 90 percent
Reliability asks a stricter question than accuracy: of all the runs an arm made, what share reached 90 percent or better?
Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.
Data behind this figure — ZeroShot + MultiTurn · both batteries
| arm | role | column | share of runs at or above 90 % accuracy | runs | cells | min | max | sd | rank | runs in the repository |
|---|---|---|---|---|---|---|---|---|---|---|
| TrueArchitect | TrueArchitect | 9 models | 77.2% | 330 | 34 | 65.0% | 100.0% | 8.2% | 1st of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| CC Bare | bare harness | 6 models | 66.6% | 201 | 22 | 40.0% | 100.0% | 11.5% | 5th of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Codex Bare | bare harness | 3 models | 50.8% | 120 | 12 | 55.0% | 100.0% | 13.6% | 9th of 9 | GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| Cursor Bare | bare harness | 8 models | 65.6% | 320 | 32 | 55.0% | 100.0% | 11.4% | 6th of 9 | Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5 |
| Graphify | indexing tool on Claude Code | 6 models | 68.6% | 201 | 22 | 40.0% | 100.0% | 11.6% | 2nd of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| GitNexus | indexing tool on Claude Code | 6 models | 67.7% | 201 | 22 | 55.0% | 100.0% | 9.2% | 3rd of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| CodeGraph | indexing tool on Claude Code | 6 models | 63.5% | 203 | 22 | 40.0% | 100.0% | 12.9% | 8th of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| CodebaseMemory | indexing tool on Claude Code | 6 models | 65.1% | 201 | 22 | 60.0% | 100.0% | 9.7% | 7th of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Serena | indexing tool on Claude Code | 6 models | 67.6% | 207 | 23 | 50.0% | 100.0% | 10.9% | 4th of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
Figure 3. Reliability asks a stricter question than accuracy: of all the runs an arm made, what share reached 90 percent or better? Each column is that share over the arm's valid, scored runs at one model (or pooled over its roster in the pooled-models view), with TrueArchitect and the bare harnesses as dashed references. The gap between arms is far wider here than in mean accuracy: the index converts a good average into a dependable result.
How this figure is computed
Definition. Within a cell (arm × battery × model × exam), pass-rate90 = 100 · |{r : accuracy(r) ≥ 90}| / |{r scored}|. The threshold is fixed at 90 percent of the battery, that is 27 of 30 on the base battery and 18 of 20 on the hard battery.
Aggregation. Cells are averaged with equal weight across the selected batteries (and models, when pooled). The statistic is bounded in [0, 100] and is a property of the distribution, not of the mean; it is the figure to read when the question is "how often does this work well", rather than "how well does it work on average".
Reading it. Higher is better. Because the threshold is strict, small differences in mean accuracy become large differences here, which is why the figure is reported alongside Figure 2 rather than instead of it.
Provenance
Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.