Figure 10 · E004

Pass rate by question category and difficulty

Every question carries one or more categories and a difficulty; each cell is the pass rate of an arm on the questions in that category (or at that difficulty), pooled over the selected models — by default the Anthropic models, which every arm ran.

protocol
battery
model

Each cell pools the arm's Anthropic models with equal weight per model, so every arm is compared over the same vendor's models; an arm that ran none of them shows a gap.

Pass rate by question category and difficultyhigher is betterZeroShot · both batteries · Anthropic models
absence-proofaggregationcallscollisioncross-stackdbdispatchendpointsfield-grainimpactmulti-hopnameless-entryreality-checktypesdifficulty 1difficulty 2difficulty 3difficulty 4TrueArchitectTrueArchitect828496929999991006798911009798100989891CC Barebare harness766089569696849463828889849498948877Cursor Barebare harness8370986299100951008097981009398100989887Graphifyindexing tool on Claude Code7551896396989097668992928897100959080GitNexusindexing tool on Claude Code844990619799989971889196919899999083CodeGraphindexing tool on Claude Code55549270939787974687921008496100959176CodebaseMemoryindexing tool on Claude Code804791609698901007094871009099100969481Serenaindexing tool on Claude Code81629165959793957186879185989995908136 pts+36 ptscolor = difference from bare Claude Code in the same column · number = pass rate %
Data behind this figure ZeroShot · both batteries · Anthropic models
armcolumnpass rateattemptscorrectmodels
ta-ask-fz-009absence-proof82.0%100825
ta-ask-fz-009aggregation84.0%100845
ta-ask-fz-009calls96.0%4484296
ta-ask-fz-009collision92.0%100925
ta-ask-fz-009cross-stack98.5%4484416
ta-ask-fz-009db99.2%4484446
ta-ask-fz-009dispatch99.0%100995
ta-ask-fz-009endpoints99.8%5045036
ta-ask-fz-009field-grain67.0%100675
ta-ask-fz-009impact98.1%3363296
ta-ask-fz-009multi-hop91.0%100915
ta-ask-fz-009nameless-entry100.0%1501505
ta-ask-fz-009reality-check97.2%2502435
ta-ask-fz-009types98.0%2802746
ta-ask-fz-009difficulty 1100.0%5605606
ta-ask-fz-009difficulty 298.0%5605486
ta-ask-fz-009difficulty 397.5%5605456
ta-ask-fz-009difficulty 490.8%10009085
coldabsence-proof76.0%100765
coldaggregation60.0%100605
coldcalls89.3%3603185
coldcollision56.0%100565
coldcross-stack96.3%3603455
colddb96.3%3603455
colddispatch84.0%100845
coldendpoints93.8%4053775
coldfield-grain63.0%100635
coldimpact81.7%2702165
coldmulti-hop88.0%100885
coldnameless-entry88.7%1501335
coldreality-check84.4%2502115
coldtypes94.4%2252115
colddifficulty 198.4%4504425
colddifficulty 294.0%4504205
colddifficulty 388.0%4503915
colddifficulty 477.1%10007715
cursorabsence-proof83.3%60503
cursoraggregation70.0%60423
cursorcalls97.5%2402343
cursorcollision61.7%60373
cursorcross-stack98.8%2402373
cursordb100.0%2402403
cursordispatch95.0%60573
cursorendpoints99.6%2702693
cursorfield-grain80.0%60483
cursorimpact97.2%1801753
cursormulti-hop98.3%60593
cursornameless-entry100.0%90903
cursorreality-check92.7%1501393
cursortypes98.0%1501473
cursordifficulty 1100.0%3003003
cursordifficulty 298.0%3002943
cursordifficulty 398.0%3002943
cursordifficulty 487.0%6005223
graphifyabsence-proof75.0%100755
graphifyaggregation51.0%100515
graphifycalls88.8%3603155
graphifycollision63.0%100635
graphifycross-stack95.5%3603425
graphifydb98.0%3603525
graphifydispatch90.0%100905
graphifyendpoints96.9%4053915
graphifyfield-grain66.0%100665
graphifyimpact89.0%2702375
graphifymulti-hop92.0%100925
graphifynameless-entry92.0%1501385
graphifyreality-check88.0%2502205
graphifytypes97.2%2252185
graphifydifficulty 199.6%4504485
graphifydifficulty 295.4%4504275
graphifydifficulty 390.0%4504005
graphifydifficulty 479.5%10007955
gitnexusabsence-proof84.0%100845
gitnexusaggregation49.0%100495
gitnexuscalls89.5%3603185
gitnexuscollision61.0%100615
gitnexuscross-stack96.5%3603475
gitnexusdb98.8%3603555
gitnexusdispatch98.0%100985
gitnexusendpoints98.9%4054015
gitnexusfield-grain71.0%100715
gitnexusimpact88.3%2702355
gitnexusmulti-hop91.0%100915
gitnexusnameless-entry96.0%1501445
gitnexusreality-check91.2%2502285
gitnexustypes98.4%2252215
gitnexusdifficulty 199.4%4504475
gitnexusdifficulty 298.8%4504445
gitnexusdifficulty 390.2%4504025
gitnexusdifficulty 482.6%10008265
codegraphabsence-proof55.0%100555
codegraphaggregation54.0%100545
codegraphcalls92.0%3603285
codegraphcollision70.0%100705
codegraphcross-stack93.3%3603335
codegraphdb96.5%3603465
codegraphdispatch87.0%100875
codegraphendpoints97.3%4053935
codegraphfield-grain46.0%100465
codegraphimpact86.7%2702305
codegraphmulti-hop92.0%100925
codegraphnameless-entry100.0%1501505
codegraphreality-check83.6%2502095
codegraphtypes96.0%2252155
codegraphdifficulty 199.8%4504495
codegraphdifficulty 295.2%4504265
codegraphdifficulty 390.8%4504045
codegraphdifficulty 476.3%10007635
cbmabsence-proof80.0%100805
cbmaggregation47.0%100475
cbmcalls90.8%3603235
cbmcollision60.0%100605
cbmcross-stack95.5%3603425
cbmdb98.3%3603535
cbmdispatch90.0%100905
cbmendpoints99.8%4054045
cbmfield-grain70.0%100705
cbmimpact93.7%2702515
cbmmulti-hop87.0%100875
cbmnameless-entry100.0%1501505
cbmreality-check90.4%2502265
cbmtypes98.8%2252225
cbmdifficulty 199.8%4504495
cbmdifficulty 295.6%4504285
cbmdifficulty 394.2%4504215
cbmdifficulty 481.0%10008105
serenaabsence-proof80.8%108856
serenaaggregation62.1%108676
serenacalls90.5%3603235
serenacollision64.6%108706
serenacross-stack95.3%3603415
serenadb96.8%3603475
serenadispatch92.5%108996
serenaendpoints95.1%4053835
serenafield-grain70.8%108736
serenaimpact85.7%2702285
serenamulti-hop87.1%108946
serenanameless-entry91.1%1621466
serenareality-check84.8%2702266
serenatypes98.0%2252205
serenadifficulty 199.0%4504455
serenadifficulty 295.0%4504255
serenadifficulty 389.8%4504005
serenadifficulty 480.7%10808606

Figure 10. Every question carries one or more categories and a difficulty; each cell is the pass rate of an arm on the questions in that category (or at that difficulty), pooled over the selected models — by default the Anthropic models, which every arm ran. The number in each cell is the pass rate; the color is that cell's difference from bare Claude Code in the same column — blue above, orange below, neutral at the baseline — so a row reads as where an arm beats or trails the null hypothesis. The categories where an index should matter most, structural questions such as callers, impact and cross-stack tracing, are where the separation is widest.

How this figure is computed

Definition. For an arm, model, exam, battery and question, the per-question record gives pass (attempts marked correct) and n (attempts). For a category c, the pass rate at one model is Σ pass / Σ n over the questions carrying c; the cell value is the equal-weight mean over the arm's models. Difficulty cells are formed the same way over the questions at that difficulty.

Aggregation. One question may carry several categories and then contributes to each; category rates are therefore not additive across categories. The model control decides which models a cell pools, with equal weight per model. The figure opens on the Anthropic roster because every arm ran those models, so the cells compare like with like; "OpenAI models" does the same for the GPT roster (TrueArchitect, Codex and Cursor); a single model is the sharpest read. With "all" a cell pools the arm's OWN roster: TrueArchitect ran GPT 5.6 Luna, Terra and Sol as well as the Claude models, while bare Claude Code and the indexing tools ran Claude models only, so a pooled TrueArchitect cell averages in three models its comparators never ran. The multi-hop column is the clearest case: two questions (h06, h07) on which TrueArchitect at every Claude model matches or beats bare Claude Code, while the GPT models score far lower and pull the all-models cell from 91 down to 69; field-grain flips the same way. The comparison pages carry this grid restricted to their own arms and models.

Reading it. Higher is better. Read a row to see an arm's profile and a column to see which arms handle a question class. Color encodes polarity, not magnitude: each cell is colored by its distance from bare Claude Code's pass rate in the same column (the column mean if that arm is absent from the selection), on a symmetric scale whose half-width is the largest difference in the figure and never less than ten points; the number is always the pass rate itself. A two-hue diverging scale with a neutral midpoint was chosen over a green-to-red one because the latter is unreadable under red–green color vision deficiency.

Provenance

Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.