Figure 3 · E004

Reliability: runs scoring at least 90 percent

Reliability asks a stricter question than accuracy: of all the runs an arm made, what share reached 90 percent or better?

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Reliability: runs scoring at least 90 percenthigher is betterZeroShot + MultiTurn · both batteries
020406080TrueArchitect 77.2%CC Bare 66.6%Codex Bare 50.8%Cursor Bare 65.6%77.2%9 modelsTrueArchitectTrueArchitect1st of 966.6%6 modelsCC Barebare harness5th of 950.8%3 modelsCodex Barebare harness9th of 965.6%8 modelsCursor Barebare harness6th of 968.6%6 modelsGraphifyindexing tool2nd of 967.7%6 modelsGitNexusindexing tool3rd of 963.5%6 modelsCodeGraphindexing tool8th of 965.1%6 modelsCodebaseMemoryindexing tool7th of 967.6%6 modelsSerenaindexing tool4th of 9
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnshare of runs at or above 90 % accuracyrunscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect9 models77.2%3303465.0%100.0%8.2%1st of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
CC Barebare harness6 models66.6%2012240.0%100.0%11.5%5th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Codex Barebare harness3 models50.8%1201255.0%100.0%13.6%9th of 9GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harness8 models65.6%3203255.0%100.0%11.4%6th of 9Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5
Graphifyindexing tool on Claude Code6 models68.6%2012240.0%100.0%11.6%2nd of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
GitNexusindexing tool on Claude Code6 models67.7%2012255.0%100.0%9.2%3rd of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodeGraphindexing tool on Claude Code6 models63.5%2032240.0%100.0%12.9%8th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodebaseMemoryindexing tool on Claude Code6 models65.1%2012260.0%100.0%9.7%7th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Serenaindexing tool on Claude Code6 models67.6%2072350.0%100.0%10.9%4th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5

Figure 3. Reliability asks a stricter question than accuracy: of all the runs an arm made, what share reached 90 percent or better? Each column is that share over the arm's valid, scored runs at one model (or pooled over its roster in the pooled-models view), with TrueArchitect and the bare harnesses as dashed references. The gap between arms is far wider here than in mean accuracy: the index converts a good average into a dependable result.

How this figure is computed

Definition. Within a cell (arm × battery × model × exam), pass-rate90 = 100 · |{r : accuracy(r) ≥ 90}| / |{r scored}|. The threshold is fixed at 90 percent of the battery, that is 27 of 30 on the base battery and 18 of 20 on the hard battery.

Aggregation. Cells are averaged with equal weight across the selected batteries (and models, when pooled). The statistic is bounded in [0, 100] and is a property of the distribution, not of the mean; it is the figure to read when the question is "how often does this work well", rather than "how well does it work on average".

Reading it. Higher is better. Because the threshold is strict, small differences in mean accuracy become large differences here, which is why the figure is reported alongside Figure 2 rather than instead of it.

Provenance

Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.