Figure 5 · E004

Outcome composition per run

Every valid run becomes one thin column showing how its questions divided between correct (the arm's color), incorrect (dark) and unanswered or unjudged (black), so an arm's whole record is visible at once.

protocols
battery
model
Outcome composition per runhigher is betterZeroShot + MultiTurn · both batteries
025507510093.3%Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5LunaTerraSolTrueArchitectTrueArchitect · 330 runs1st of 989.4%3.9%Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5CC Barebare harness · 201 runs7th of 985.9%7.4%LunaTerraSolCodex Barebare harness · 120 runs9th of 989.6%3.7%Haiku 4.5Sonnet 5Opus 5LunaTerraSolGrok 4.6Comp 2.5Cursor Barebare harness · 320 runs6th of 990.0%3.4%Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5Graphifyindexing tool on Claude Code · 201 runs5th of 990.6%2.8%Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5GitNexusindexing tool on Claude Code · 201 runs2nd of 988.4%5.0%Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5CodeGraphindexing tool on Claude Code · 203 runs8th of 990.4%2.9%Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5CodebaseMemoryindexing tool on Claude Code · 201 runs3rd of 990.1%3.2%Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5Serenaindexing tool on Claude Code · 207 runs4th of 9
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnquestions correct · incorrect · unanswered, per runrunscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect9 models93.3%3303465.0%100.0%8.2%1st of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
CC Barebare harness6 models89.4%2012240.0%100.0%11.5%7th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Codex Barebare harness3 models85.9%1201255.0%100.0%13.6%9th of 9GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harness8 models89.6%3203255.0%100.0%11.4%6th of 9Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5
Graphifyindexing tool on Claude Code6 models90.0%2012240.0%100.0%11.6%5th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
GitNexusindexing tool on Claude Code6 models90.6%2012255.0%100.0%9.2%2nd of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodeGraphindexing tool on Claude Code6 models88.4%2032240.0%100.0%12.9%8th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodebaseMemoryindexing tool on Claude Code6 models90.4%2012260.0%100.0%9.7%3rd of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Serenaindexing tool on Claude Code6 models90.1%2072350.0%100.0%10.9%4th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5

Figure 5. Every valid run becomes one thin column showing how its questions divided between correct (the arm's color), incorrect (dark) and unanswered or unjudged (black), so an arm's whole record is visible at once. Columns are ordered by model then repetition; the number above each group is the arm's mean accuracy and the number inside is its difference from TrueArchitect. The dark bands are the honesty layer: the failures are on the page, not in a footnote.

How this figure is computed

Definition. For a scored run, correct = score, incorrect = the count of effective verdicts of fail, and the residue denominator − correct − incorrect is the unanswered or unjudged portion (an adjudicate or infra verdict, or a question with no verdict). Each is drawn as a share of the denominator.

Aggregation. None: every scored run is one column. The header figure is the hierarchical mean accuracy of the group and the inner figure is that mean minus TrueArchitect's.

Reading it. Higher is better for the colored band. The width of a group is proportional to the number of runs the arm made in the selection, which is itself information: an arm with fewer columns has fewer runs behind its numbers.

Disclosures

Did-not-finish runs have no score and do not appear here; their count per arm is in the run records and in the Limitations page.

Provenance

Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.