Comparison · E004

TrueArchitect & Anthropic

Every arm that ran at an Anthropic model — bare Claude Code, Cursor Agent at Claude models, and the five indexing tools installed on Claude Code — beside TrueArchitect at the same models.

Switch the columns to per model and each comparison is like for like: Haiku against Haiku, Sonnet against Sonnet, Opus against Opus.

Arms: TrueArchitect · CC Bare · Cursor Bare · Graphify · GitNexus · CodeGraph · CodebaseMemory · Serena. Models: Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5.1920 valid scored runs. Figures open with one column per arm (pooled models); the Columns control switches every figure to one column per model, for like-for-like reads at a fixed model, and the choice is remembered site-wide. The defaults pool the ZeroShot and MultiTurn protocols over both batteries (each protocol × battery × model cell one vote). The arm and model set is fixed to this comparison.

Figure 1 Context tokens per run

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Context tokens per runlower is betterZeroShot + MultiTurn · both batteries
02.5M5.0M7.5M10.0MTrueArchitect 3.96MCC Bare 8.59MCursor Bare 10.44M3.96M6 modelsTrueArchitectTrueArchitect1st of 88.59M6 modelsCC Barebare harness3rd of 810.44M3 modelsCursor Barebare harness8th of 88.96M6 modelsGraphifyindexing tool4th of 89.28M6 modelsGitNexusindexing tool6th of 86.92M6 modelsCodeGraphindexing tool2nd of 89.11M6 modelsCodebaseMemoryindexing tool5th of 89.64M6 modelsSerenaindexing tool7th of 8
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumncontext tokens per run (input + cache read)runscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect6 models3.96M21022786k20.85M3.76M1st of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CC Barebare harness6 models8.59M201223.36M29.61M4.69M3rd of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Cursor Barebare harness3 models10.44M120122.32M23.84M4.41M8th of 8Haiku 4.5 · Sonnet 5 · Opus 5
Graphifyindexing tool on Claude Code6 models8.96M201223.58M27.70M4.57M4th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
GitNexusindexing tool on Claude Code6 models9.28M201224.74M31.16M4.74M6th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodeGraphindexing tool on Claude Code6 models6.92M203222.75M29.62M4.32M2nd of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodebaseMemoryindexing tool on Claude Code6 models9.11M201224.07M25.47M4.84M5th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Serenaindexing tool on Claude Code6 models9.64M207234.13M30.98M5.29M7th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5

Context tokens are the tokens the model read to answer a whole run: the uncached input plus every cache read, summed over every API call in the run, all threads included. Each column is the mean over the valid, scored runs of one arm at one model, every arm split by model because token counts are model facts; the dashed lines mark TrueArchitect's and each bare harness's pooled value wherever pooling is licensed. Full figure, definition and disclosures →

Figure 11 Cost per correct answer

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Cost per correct answerlower is betterZeroShot + MultiTurn · both batteries
$0.00$0.10$0.20$0.30TrueArchitect $0.136CC Bare $0.310Cursor Bare $0.207$0.1366 modelsTrueArchitectTrueArchitect1st of 8$0.3106 modelsCC Barebare harness5th of 8$0.2073 modelsCursor Barebare harness2nd of 8$0.3426 modelsGraphifyindexing tool8th of 8$0.3116 modelsGitNexusindexing tool6th of 8$0.2836 modelsCodeGraphindexing tool4th of 8$0.2746 modelsCodebaseMemoryindexing tool3rd of 8$0.3386 modelsSerenaindexing tool7th of 8
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnUSD per correct answerrunscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect6 models$0.13621022$0.026$0.424$0.0821st of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CC Barebare harness6 models$0.31020122$0.071$0.939$0.1865th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Cursor Barebare harness3 models$0.20712012$0.014$0.619$0.1432nd of 8Haiku 4.5 · Sonnet 5 · Opus 5
Graphifyindexing tool on Claude Code6 models$0.34220122$0.065$0.901$0.1998th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
GitNexusindexing tool on Claude Code6 models$0.31120122$0.070$0.677$0.1536th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodeGraphindexing tool on Claude Code6 models$0.28320322$0.067$0.750$0.1444th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodebaseMemoryindexing tool on Claude Code6 models$0.27420122$0.072$0.623$0.1313rd of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Serenaindexing tool on Claude Code6 models$0.33820723$0.071$0.898$0.1897th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5

Cost per correct answer divides what a run cost by the number of questions it answered correctly: the vendor-reported figure where the harness reported one, otherwise the benchmark's estimate from the published rate tables over the run's four token columns. Each column is the mean over valid, scored runs at one model (or pooled over the arm's roster in the pooled-models view); TrueArchitect and the bare harnesses are the dashed references, and the data table names each row's cost basis. Full figure, definition and disclosures →

Figure 5 Outcome composition per run

protocols
battery
model
Outcome composition per runhigher is betterZeroShot + MultiTurn · both batteries
025507510095.6%Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5TrueArchitectTrueArchitect · 210 runs1st of 889.4%6.1%Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5CC Barebare harness · 201 runs7th of 893.1%2.5%Haiku 4.5Sonnet 5Opus 5Cursor Barebare harness · 120 runs2nd of 890.0%5.6%Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5Graphifyindexing tool on Claude Code · 201 runs6th of 890.6%5.0%Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5GitNexusindexing tool on Claude Code · 201 runs3rd of 888.4%7.2%Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5CodeGraphindexing tool on Claude Code · 203 runs8th of 890.4%5.1%Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5CodebaseMemoryindexing tool on Claude Code · 201 runs4th of 890.1%5.4%Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5Serenaindexing tool on Claude Code · 207 runs5th of 8
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnquestions correct · incorrect · unanswered, per runrunscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect6 models95.6%2102265.0%100.0%6.3%1st of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CC Barebare harness6 models89.4%2012240.0%100.0%11.5%7th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Cursor Barebare harness3 models93.1%1201265.0%100.0%8.1%2nd of 8Haiku 4.5 · Sonnet 5 · Opus 5
Graphifyindexing tool on Claude Code6 models90.0%2012240.0%100.0%11.6%6th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
GitNexusindexing tool on Claude Code6 models90.6%2012255.0%100.0%9.2%3rd of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodeGraphindexing tool on Claude Code6 models88.4%2032240.0%100.0%12.9%8th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodebaseMemoryindexing tool on Claude Code6 models90.4%2012260.0%100.0%9.7%4th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Serenaindexing tool on Claude Code6 models90.1%2072350.0%100.0%10.9%5th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5

Every valid run becomes one thin column showing how its questions divided between correct (the arm's color), incorrect (dark) and unanswered or unjudged (black), so an arm's whole record is visible at once. Columns are ordered by model then repetition; the number above each group is the arm's mean accuracy and the number inside is its difference from TrueArchitect. Full figure, definition and disclosures →

Figure 2 Accuracy per run

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Accuracy per runhigher is betterZeroShot + MultiTurn · both batteries
406080100TrueArchitect 95.6%CC Bare 89.4%Cursor Bare 93.1%95.6%65.0%100.0%TrueArchitectTrueArchitect1st of 889.4%40.0%100.0%CC Barebare harness7th of 893.1%65.0%100.0%Cursor Barebare harness2nd of 890.0%40.0%100.0%Graphifyindexing tool6th of 890.6%55.0%100.0%GitNexusindexing tool3rd of 888.4%40.0%100.0%CodeGraphindexing tool8th of 890.4%60.0%100.0%CodebaseMemoryindexing tool4th of 890.1%50.0%100.0%Serenaindexing tool5th of 8
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnaccuracy (% of questions correct)runscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect6 models95.6%2102265.0%100.0%6.3%1st of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CC Barebare harness6 models89.4%2012240.0%100.0%11.5%7th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Cursor Barebare harness3 models93.1%1201265.0%100.0%8.1%2nd of 8Haiku 4.5 · Sonnet 5 · Opus 5
Graphifyindexing tool on Claude Code6 models90.0%2012240.0%100.0%11.6%6th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
GitNexusindexing tool on Claude Code6 models90.6%2012255.0%100.0%9.2%3rd of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodeGraphindexing tool on Claude Code6 models88.4%2032240.0%100.0%12.9%8th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodebaseMemoryindexing tool on Claude Code6 models90.4%2012260.0%100.0%9.7%4th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Serenaindexing tool on Claude Code6 models90.1%2072350.0%100.0%10.9%5th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5

Accuracy is the share of a run's questions whose answer the grading pipeline marked correct, so every dot is one complete run of one arm at one model. The solid tick is the arm's pooled mean, the thin line its worst-to-best range, and dots are translucent so overlap reads as density; hover a dot for its model, repetition and run directory. Full figure, definition and disclosures →

Figure 12 Cost per run

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Cost per runlower is betterZeroShot + MultiTurn · both batteries
$0.00$5.0$10$15$20$25TrueArchitect $3.06CC Bare $6.18Cursor Bare $4.49$3.06$0.759$7.27TrueArchitectTrueArchitect1st of 8$6.18$2.14$15.0CC Barebare harness5th of 8$4.49$0.409$12.9Cursor Barebare harness2nd of 8$7.14$1.94$24.1Graphifyindexing tool8th of 8$6.49$2.06$13.6GitNexusindexing tool6th of 8$5.81$1.16$14.4CodeGraphindexing tool4th of 8$5.70$2.10$12.0CodebaseMemoryindexing tool3rd of 8$6.80$2.12$13.9Serenaindexing tool7th of 8
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnUSD per runrunscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect6 models$3.0621022$0.759$7.27$1.481st of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CC Barebare harness6 models$6.1820122$2.14$15.0$2.665th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Cursor Barebare harness3 models$4.4912012$0.409$12.9$2.932nd of 8Haiku 4.5 · Sonnet 5 · Opus 5
Graphifyindexing tool on Claude Code6 models$7.1420122$1.94$24.1$3.898th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
GitNexusindexing tool on Claude Code6 models$6.4920122$2.06$13.6$2.496th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodeGraphindexing tool on Claude Code6 models$5.8120322$1.16$14.4$2.494th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodebaseMemoryindexing tool on Claude Code6 models$5.7020122$2.10$12.0$2.023rd of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Serenaindexing tool on Claude Code6 models$6.8020723$2.12$13.9$2.827th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5

Every dot is what one run cost (vendor-reported where available, otherwise estimated from the rate tables), so the spread of an arm's price is visible, not just its mean. The solid tick is the pooled mean and the thin line the range; hover a dot for its model, repetition and run directory. Full figure, definition and disclosures →

Figure 13 Tool result tokens per run

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Tool result tokens per runlower is betterZeroShot + MultiTurn · both batteries
0200k400k600kTrueArchitect 93kCC Bare 125kCursor Bare 104k93k21k292kTrueArchitectTrueArchitect3rd of 8125k3k536kCC Barebare harness6th of 8104k19k617kCursor Barebare harness4th of 8130k10k517kGraphifyindexing tool8th of 889k10k527kGitNexusindexing tool2nd of 8128k7k414kCodeGraphindexing tool7th of 882k10k345kCodebaseMemoryindexing tool1st of 8109k9k445kSerenaindexing tool5th of 8
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumntool result tokens per run (estimated)runscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect6 models93k2102221k292k61k3rd of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CC Barebare harness6 models125k201223k536k147k6th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Cursor Barebare harness3 models104k1201219k617k86k4th of 8Haiku 4.5 · Sonnet 5 · Opus 5
Graphifyindexing tool on Claude Code6 models130k2012210k517k144k8th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
GitNexusindexing tool on Claude Code6 models89k2012210k527k83k2nd of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodeGraphindexing tool on Claude Code6 models128k203227k414k111k7th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodebaseMemoryindexing tool on Claude Code6 models82k2012210k345k83k1st of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Serenaindexing tool on Claude Code6 models109k207239k445k111k5th of 8Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5

Tool result tokens are the volume of text the arm's tools pushed back into the model's context over a run — file contents, search hits, shell output, or an index query's result — estimated per call from the result size. Every dot is one run, split per model for every arm; the solid tick is the mean and the thin line the range. Full figure, definition and disclosures →

Figure 10 Pass rate by question category and difficulty

protocol
battery
model

Each cell pools the arm's Anthropic models with equal weight per model, so every arm is compared over the same vendor's models; an arm that ran none of them shows a gap.

Pass rate by question category and difficultyhigher is betterZeroShot · both batteries · Anthropic models
absence-proofaggregationcallscollisioncross-stackdbdispatchendpointsfield-grainimpactmulti-hopnameless-entryreality-checktypesdifficulty 1difficulty 2difficulty 3difficulty 4TrueArchitectTrueArchitect828496929999991006798911009798100989891CC Barebare harness766089569696849463828889849498948877Cursor Barebare harness8370986299100951008097981009398100989887Graphifyindexing tool on Claude Code7551896396989097668992928897100959080GitNexusindexing tool on Claude Code844990619799989971889196919899999083CodeGraphindexing tool on Claude Code55549270939787974687921008496100959176CodebaseMemoryindexing tool on Claude Code804791609698901007094871009099100969481Serenaindexing tool on Claude Code81629165959793957186879185989995908136 pts+36 ptscolor = difference from bare Claude Code in the same column · number = pass rate %
Data behind this figure ZeroShot · both batteries · Anthropic models
armcolumnpass rateattemptscorrectmodels
ta-ask-fz-009absence-proof82.0%100825
ta-ask-fz-009aggregation84.0%100845
ta-ask-fz-009calls96.0%4484296
ta-ask-fz-009collision92.0%100925
ta-ask-fz-009cross-stack98.5%4484416
ta-ask-fz-009db99.2%4484446
ta-ask-fz-009dispatch99.0%100995
ta-ask-fz-009endpoints99.8%5045036
ta-ask-fz-009field-grain67.0%100675
ta-ask-fz-009impact98.1%3363296
ta-ask-fz-009multi-hop91.0%100915
ta-ask-fz-009nameless-entry100.0%1501505
ta-ask-fz-009reality-check97.2%2502435
ta-ask-fz-009types98.0%2802746
ta-ask-fz-009difficulty 1100.0%5605606
ta-ask-fz-009difficulty 298.0%5605486
ta-ask-fz-009difficulty 397.5%5605456
ta-ask-fz-009difficulty 490.8%10009085
coldabsence-proof76.0%100765
coldaggregation60.0%100605
coldcalls89.3%3603185
coldcollision56.0%100565
coldcross-stack96.3%3603455
colddb96.3%3603455
colddispatch84.0%100845
coldendpoints93.8%4053775
coldfield-grain63.0%100635
coldimpact81.7%2702165
coldmulti-hop88.0%100885
coldnameless-entry88.7%1501335
coldreality-check84.4%2502115
coldtypes94.4%2252115
colddifficulty 198.4%4504425
colddifficulty 294.0%4504205
colddifficulty 388.0%4503915
colddifficulty 477.1%10007715
cursorabsence-proof83.3%60503
cursoraggregation70.0%60423
cursorcalls97.5%2402343
cursorcollision61.7%60373
cursorcross-stack98.8%2402373
cursordb100.0%2402403
cursordispatch95.0%60573
cursorendpoints99.6%2702693
cursorfield-grain80.0%60483
cursorimpact97.2%1801753
cursormulti-hop98.3%60593
cursornameless-entry100.0%90903
cursorreality-check92.7%1501393
cursortypes98.0%1501473
cursordifficulty 1100.0%3003003
cursordifficulty 298.0%3002943
cursordifficulty 398.0%3002943
cursordifficulty 487.0%6005223
graphifyabsence-proof75.0%100755
graphifyaggregation51.0%100515
graphifycalls88.8%3603155
graphifycollision63.0%100635
graphifycross-stack95.5%3603425
graphifydb98.0%3603525
graphifydispatch90.0%100905
graphifyendpoints96.9%4053915
graphifyfield-grain66.0%100665
graphifyimpact89.0%2702375
graphifymulti-hop92.0%100925
graphifynameless-entry92.0%1501385
graphifyreality-check88.0%2502205
graphifytypes97.2%2252185
graphifydifficulty 199.6%4504485
graphifydifficulty 295.4%4504275
graphifydifficulty 390.0%4504005
graphifydifficulty 479.5%10007955
gitnexusabsence-proof84.0%100845
gitnexusaggregation49.0%100495
gitnexuscalls89.5%3603185
gitnexuscollision61.0%100615
gitnexuscross-stack96.5%3603475
gitnexusdb98.8%3603555
gitnexusdispatch98.0%100985
gitnexusendpoints98.9%4054015
gitnexusfield-grain71.0%100715
gitnexusimpact88.3%2702355
gitnexusmulti-hop91.0%100915
gitnexusnameless-entry96.0%1501445
gitnexusreality-check91.2%2502285
gitnexustypes98.4%2252215
gitnexusdifficulty 199.4%4504475
gitnexusdifficulty 298.8%4504445
gitnexusdifficulty 390.2%4504025
gitnexusdifficulty 482.6%10008265
codegraphabsence-proof55.0%100555
codegraphaggregation54.0%100545
codegraphcalls92.0%3603285
codegraphcollision70.0%100705
codegraphcross-stack93.3%3603335
codegraphdb96.5%3603465
codegraphdispatch87.0%100875
codegraphendpoints97.3%4053935
codegraphfield-grain46.0%100465
codegraphimpact86.7%2702305
codegraphmulti-hop92.0%100925
codegraphnameless-entry100.0%1501505
codegraphreality-check83.6%2502095
codegraphtypes96.0%2252155
codegraphdifficulty 199.8%4504495
codegraphdifficulty 295.2%4504265
codegraphdifficulty 390.8%4504045
codegraphdifficulty 476.3%10007635
cbmabsence-proof80.0%100805
cbmaggregation47.0%100475
cbmcalls90.8%3603235
cbmcollision60.0%100605
cbmcross-stack95.5%3603425
cbmdb98.3%3603535
cbmdispatch90.0%100905
cbmendpoints99.8%4054045
cbmfield-grain70.0%100705
cbmimpact93.7%2702515
cbmmulti-hop87.0%100875
cbmnameless-entry100.0%1501505
cbmreality-check90.4%2502265
cbmtypes98.8%2252225
cbmdifficulty 199.8%4504495
cbmdifficulty 295.6%4504285
cbmdifficulty 394.2%4504215
cbmdifficulty 481.0%10008105
serenaabsence-proof80.8%108856
serenaaggregation62.1%108676
serenacalls90.5%3603235
serenacollision64.6%108706
serenacross-stack95.3%3603415
serenadb96.8%3603475
serenadispatch92.5%108996
serenaendpoints95.1%4053835
serenafield-grain70.8%108736
serenaimpact85.7%2702285
serenamulti-hop87.1%108946
serenanameless-entry91.1%1621466
serenareality-check84.8%2702266
serenatypes98.0%2252205
serenadifficulty 199.0%4504455
serenadifficulty 295.0%4504255
serenadifficulty 389.8%4504005
serenadifficulty 480.7%10808606

Every question carries one or more categories and a difficulty; each cell is the pass rate of an arm on the questions in that category (or at that difficulty), pooled over the selected models — by default the Anthropic models, which every arm ran. The number in each cell is the pass rate; the color is that cell's difference from bare Claude Code in the same column — blue above, orange below, neutral at the baseline — so a row reads as where an arm beats or trails the null hypothesis. Full figure, definition and disclosures →