Accuracy per run
Accuracy is the share of a run's questions whose answer the grading pipeline marked correct, so every dot is one complete run of one arm at one model.
Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.
Data behind this figure — ZeroShot + MultiTurn · both batteries
| arm | role | column | accuracy (% of questions correct) | runs | cells | min | max | sd | rank | runs in the repository |
|---|---|---|---|---|---|---|---|---|---|---|
| TrueArchitect | TrueArchitect | 9 models | 93.3% | 330 | 34 | 65.0% | 100.0% | 8.2% | 1st of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| CC Bare | bare harness | 6 models | 89.4% | 201 | 22 | 40.0% | 100.0% | 11.5% | 7th of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Codex Bare | bare harness | 3 models | 85.9% | 120 | 12 | 55.0% | 100.0% | 13.6% | 9th of 9 | GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| Cursor Bare | bare harness | 8 models | 89.6% | 320 | 32 | 55.0% | 100.0% | 11.4% | 6th of 9 | Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5 |
| Graphify | indexing tool on Claude Code | 6 models | 90.0% | 201 | 22 | 40.0% | 100.0% | 11.6% | 5th of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| GitNexus | indexing tool on Claude Code | 6 models | 90.6% | 201 | 22 | 55.0% | 100.0% | 9.2% | 2nd of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| CodeGraph | indexing tool on Claude Code | 6 models | 88.4% | 203 | 22 | 40.0% | 100.0% | 12.9% | 8th of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| CodebaseMemory | indexing tool on Claude Code | 6 models | 90.4% | 201 | 22 | 60.0% | 100.0% | 9.7% | 3rd of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Serena | indexing tool on Claude Code | 6 models | 90.1% | 207 | 23 | 50.0% | 100.0% | 10.9% | 4th of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
Figure 2. Accuracy is the share of a run's questions whose answer the grading pipeline marked correct, so every dot is one complete run of one arm at one model. The solid tick is the arm's pooled mean, the thin line its worst-to-best range, and dots are translucent so overlap reads as density; hover a dot for its model, repetition and run directory. TrueArchitect's distribution sits highest and tightest; the lower context use in Figure 1 is not bought with accuracy.
How this figure is computed
Definition. For a scored run r, accuracy(r) = 100 · score(r) / denominator(r), where score is the number of questions with an effective verdict of pass (tier T1 mechanical match or tier T2 judge, human adjudication outranking both) and denominator is the battery size. Did-not-finish runs have no score and are absent from this figure by construction.
Aggregation. The mean tick is the hierarchical mean: cells of battery × model × exam are averaged with equal weight. The range line spans the minimum and maximum run. Every dot is drawn; nothing is summarised away.
Reading it. Higher is better. Compare the position of the means and, separately, the vertical spread: two arms with the same mean and different spread are not equivalent instruments, and the spread is what a user experiences run to run.
Provenance
Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.