Figure 2 · E004

Accuracy per run

Accuracy is the share of a run's questions whose answer the grading pipeline marked correct, so every dot is one complete run of one arm at one model.

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Accuracy per runhigher is betterZeroShot + MultiTurn · both batteries
406080100TrueArchitect 93.3%CC Bare 89.4%Codex Bare 85.9%Cursor Bare 89.6%93.3%65.0%100.0%TrueArchitectTrueArchitect1st of 989.4%40.0%100.0%CC Barebare harness7th of 985.9%55.0%100.0%Codex Barebare harness9th of 989.6%55.0%100.0%Cursor Barebare harness6th of 990.0%40.0%100.0%Graphifyindexing tool5th of 990.6%55.0%100.0%GitNexusindexing tool2nd of 988.4%40.0%100.0%CodeGraphindexing tool8th of 990.4%60.0%100.0%CodebaseMemoryindexing tool3rd of 990.1%50.0%100.0%Serenaindexing tool4th of 9
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnaccuracy (% of questions correct)runscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect9 models93.3%3303465.0%100.0%8.2%1st of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
CC Barebare harness6 models89.4%2012240.0%100.0%11.5%7th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Codex Barebare harness3 models85.9%1201255.0%100.0%13.6%9th of 9GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harness8 models89.6%3203255.0%100.0%11.4%6th of 9Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5
Graphifyindexing tool on Claude Code6 models90.0%2012240.0%100.0%11.6%5th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
GitNexusindexing tool on Claude Code6 models90.6%2012255.0%100.0%9.2%2nd of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodeGraphindexing tool on Claude Code6 models88.4%2032240.0%100.0%12.9%8th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodebaseMemoryindexing tool on Claude Code6 models90.4%2012260.0%100.0%9.7%3rd of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Serenaindexing tool on Claude Code6 models90.1%2072350.0%100.0%10.9%4th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5

Figure 2. Accuracy is the share of a run's questions whose answer the grading pipeline marked correct, so every dot is one complete run of one arm at one model. The solid tick is the arm's pooled mean, the thin line its worst-to-best range, and dots are translucent so overlap reads as density; hover a dot for its model, repetition and run directory. TrueArchitect's distribution sits highest and tightest; the lower context use in Figure 1 is not bought with accuracy.

How this figure is computed

Definition. For a scored run r, accuracy(r) = 100 · score(r) / denominator(r), where score is the number of questions with an effective verdict of pass (tier T1 mechanical match or tier T2 judge, human adjudication outranking both) and denominator is the battery size. Did-not-finish runs have no score and are absent from this figure by construction.

Aggregation. The mean tick is the hierarchical mean: cells of battery × model × exam are averaged with equal weight. The range line spans the minimum and maximum run. Every dot is drawn; nothing is summarised away.

Reading it. Higher is better. Compare the position of the means and, separately, the vertical spread: two arms with the same mean and different spread are not equivalent instruments, and the spread is what a user experiences run to run.

Provenance

Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.