Figure 9 · E004

Accuracy by model

The same accuracy as Figure 2, split into one column per model for every arm, so the cross-tier comparison is readable directly.

protocols
battery
model

This figure is always one column per model; the site-wide columns choice does not apply to it.

Accuracy by modelhigher is betterZeroShot + MultiTurn · both batteries
0255075100Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5LunaTerraSolTrueArchitectTrueArchitect1st of 9Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5CC Barebare harness7th of 9LunaTerraSolCodex Barebare harness9th of 9Haiku 4.5Sonnet 5Opus 5LunaTerraSolGrok 4.6Comp 2.5Cursor Barebare harness6th of 9Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5Graphifyindexing tool5th of 9Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5GitNexusindexing tool2nd of 9Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5CodeGraphindexing tool8th of 9Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5CodebaseMemoryindexing tool3rd of 9Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5Serenaindexing tool4th of 9
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnaccuracy (% of questions correct)runscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitectHaiku 4.589.9%39465.0%100.0%8.6%1st of 9Haiku 4.5
Sonnet 596.9%40480.0%100.0%4.8%Sonnet 5
Opus 4.694.8%40480.0%100.0%5.8%Opus 4.6
Opus 4.895.1%40480.0%100.0%5.5%Opus 4.8
Opus 598.8%40495.0%100.0%2.1%Opus 5
Fable 5100.0%112100.0%100.0%0.0%Fable 5
GPT 5.6 Luna87.7%40470.0%100.0%9.6%GPT 5.6 Luna
GPT 5.6 Terra89.7%40470.0%100.0%9.4%GPT 5.6 Terra
GPT 5.6 Sol90.3%40470.0%100.0%9.8%GPT 5.6 Sol
CC Barebare harnessHaiku 4.579.1%36440.0%96.7%16.2%7th of 9Haiku 4.5
Sonnet 587.0%40465.0%100.0%10.1%Sonnet 5
Opus 4.689.4%40475.0%100.0%7.2%Opus 4.6
Opus 4.889.3%30365.0%100.0%10.5%Opus 4.8
Opus 596.0%40485.0%100.0%3.8%Opus 5
Fable 597.8%15390.0%100.0%3.1%Fable 5
Codex Barebare harnessGPT 5.6 Luna85.2%40455.0%100.0%14.1%9th of 9GPT 5.6 Luna
GPT 5.6 Terra87.5%40455.0%100.0%12.2%GPT 5.6 Terra
GPT 5.6 Sol85.0%40460.0%100.0%14.4%GPT 5.6 Sol
Cursor Barebare harnessHaiku 4.587.9%40465.0%100.0%9.8%6th of 9Haiku 4.5
Sonnet 594.1%40480.0%100.0%6.7%Sonnet 5
Opus 597.2%40490.0%100.0%3.7%Opus 5
GPT 5.6 Luna83.7%40460.0%100.0%13.7%GPT 5.6 Luna
GPT 5.6 Terra84.3%40455.0%100.0%14.5%GPT 5.6 Terra
GPT 5.6 Sol85.3%40460.0%100.0%13.6%GPT 5.6 Sol
Grok 4.691.8%40465.0%100.0%9.3%Grok 4.6
Composer 2.592.7%40470.0%100.0%7.9%Composer 2.5
Graphifyindexing tool on Claude CodeHaiku 4.578.1%36440.0%100.0%17.4%5th of 9Haiku 4.5
Sonnet 591.1%40475.0%100.0%8.1%Sonnet 5
Opus 4.689.1%40475.0%100.0%7.1%Opus 4.6
Opus 4.889.3%30370.0%100.0%9.6%Opus 4.8
Opus 597.0%40490.0%100.0%3.3%Opus 5
Fable 596.8%15385.0%100.0%4.9%Fable 5
GitNexusindexing tool on Claude CodeHaiku 4.583.4%36455.0%100.0%11.5%2nd of 9Haiku 4.5
Sonnet 590.3%40475.0%100.0%7.4%Sonnet 5
Opus 4.688.4%40465.0%100.0%9.2%Opus 4.6
Opus 4.890.3%30375.0%100.0%7.9%Opus 4.8
Opus 596.5%40485.0%100.0%4.2%Opus 5
Fable 595.9%15385.0%100.0%5.3%Fable 5
CodeGraphindexing tool on Claude CodeHaiku 4.574.5%38440.0%96.7%17.9%8th of 9Haiku 4.5
Sonnet 590.4%40470.0%100.0%10.0%Sonnet 5
Opus 4.685.9%40465.0%96.7%7.9%Opus 4.6
Opus 4.888.3%30370.0%100.0%10.4%Opus 4.8
Opus 596.1%40485.0%100.0%4.3%Opus 5
Fable 597.0%15390.0%100.0%4.6%Fable 5
CodebaseMemoryindexing tool on Claude CodeHaiku 4.581.0%36460.0%100.0%10.9%3rd of 9Haiku 4.5
Sonnet 592.7%40475.0%100.0%8.1%Sonnet 5
Opus 4.688.2%40470.0%100.0%9.0%Opus 4.6
Opus 4.889.5%30370.0%100.0%8.8%Opus 4.8
Opus 596.0%40485.0%100.0%4.4%Opus 5
Fable 596.4%15385.0%100.0%5.2%Fable 5
Serenaindexing tool on Claude CodeHaiku 4.580.8%38450.0%100.0%14.2%4th of 9Haiku 4.5
Sonnet 588.5%40460.0%100.0%10.9%Sonnet 5
Opus 4.690.0%40475.0%100.0%7.1%Opus 4.6
Opus 4.888.5%30365.0%100.0%10.8%Opus 4.8
Opus 596.5%40485.0%100.0%4.2%Opus 5
Fable 596.1%19485.0%100.0%4.3%Fable 5

Figure 9. The same accuracy as Figure 2, split into one column per model for every arm, so the cross-tier comparison is readable directly. Within an arm the columns run from the fastest, cheapest model to the frontier model of each vendor. The comparison to make is diagonal: TrueArchitect at a fast or everyday model against a bare harness or an indexing tool at a frontier model.

How this figure is computed

Definition and aggregation as in Figure 2, with cells restricted to one model per column.

Reading it. Higher is better. The vendor tiers are declared, not inferred: fast = Haiku 4.5 and GPT 5.6 Luna; everyday = Sonnet 5 and GPT 5.6 Terra; frontier = Opus 5, GPT 5.6 Sol and Composer 2.5.

Provenance

Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.