Comparison · E004

TrueArchitect vs Harnesses

The null hypotheses of each ecosystem — Claude Code with nothing added, OpenAI's Codex CLI, and Cursor Agent through its SDK — beside TrueArchitect across every model any of them ran.

Every column is one model, so each read is harness against harness at a fixed model: the model is held constant and only the tooling around it changes.

Arms: TrueArchitect · CC Bare · Codex Bare · Cursor Bare. Models: Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5.1242 valid scored runs. Figures open with one column per arm (pooled models); the Columns control switches every figure to one column per model, for like-for-like reads at a fixed model, and the choice is remembered site-wide. The defaults pool the ZeroShot and MultiTurn protocols over both batteries (each protocol × battery × model cell one vote). The arm and model set is fixed to this comparison.

Figure 1 Context tokens per run

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Context tokens per runlower is betterZeroShot + MultiTurn · both batteries
05.0M10.0M15.0MCC Bare 8.59MCodex Bare 4.55M8.75MHaiku 4.53.37MSonnet 52.63MOpus 4.62.86MOpus 4.83.24MOpus 51.90MFable 54.35MLuna2.85MTerra2.58MSolTrueArchitectTrueArchitect8.59M6 modelsCC Barebare harness2nd of 24.55M3 modelsCodex Barebare harness1st of 213.83MHaiku 4.59.22MSonnet 58.27MOpus 56.18MLuna3.30MTerra2.47MSol6.58MGrok 4.69.38MComp 2.5Cursor Barebare harness

Raw token counts are never pooled across vendors (tokenizers differ), so arms whose selected models span vendors are shown one column per model; ★ marks the best value at each model.

Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumncontext tokens per run (input + cache read)runscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitectHaiku 4.58.75M3942.33M20.85M6.00MHaiku 4.5
Sonnet 53.37M4041.54M7.77M1.87MSonnet 5
Opus 4.62.63M404786k7.38M1.96MOpus 4.6
Opus 4.82.86M404884k7.66M2.13MOpus 4.8
Opus 53.24M4041.34M5.90M1.52MOpus 5
Fable 51.90M112793k3.17M1.01MFable 5
GPT 5.6 Luna4.35M404774k14.24M3.95MGPT 5.6 Luna
GPT 5.6 Terra2.85M404603k8.14M2.44MGPT 5.6 Terra
GPT 5.6 Sol2.58M404664k7.91M2.10MGPT 5.6 Sol
CC Barebare harness6 models8.59M201223.36M29.61M4.69M2nd of 2Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Codex Barebare harness3 models4.55M120122.15M10.61M2.03M1st of 2GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harnessHaiku 4.513.83M4042.32M23.84M5.11MHaiku 4.5
Sonnet 59.22M4045.15M16.79M3.26MSonnet 5
Opus 58.27M4046.57M16.43M2.11MOpus 5
GPT 5.6 Luna6.18M4044.18M10.70M1.46MGPT 5.6 Luna
GPT 5.6 Terra3.30M4041.23M5.89M1.26MGPT 5.6 Terra
GPT 5.6 Sol2.47M404985k5.16M1.28MGPT 5.6 Sol
Grok 4.66.58M4044.29M9.61M1.85MGrok 4.6
Composer 2.59.38M4045.02M14.81M3.53MComposer 2.5

Context tokens are the tokens the model read to answer a whole run: the uncached input plus every cache read, summed over every API call in the run, all threads included. Each column is the mean over the valid, scored runs of one arm at one model, every arm split by model because token counts are model facts; the dashed lines mark TrueArchitect's and each bare harness's pooled value wherever pooling is licensed. Full figure, definition and disclosures →

Figure 11 Cost per correct answer

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Cost per correct answerlower is betterZeroShot + MultiTurn · both batteries
$0.00$0.10$0.20$0.30TrueArchitect $0.115CC Bare $0.310Codex Bare $0.093Cursor Bare $0.172$0.1159 modelsTrueArchitectTrueArchitect2nd of 4$0.3106 modelsCC Barebare harness4th of 4$0.0933 modelsCodex Barebare harness1st of 4$0.1728 modelsCursor Barebare harness3rd of 4
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnUSD per correct answerrunscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect9 models$0.11533034$0.004$0.424$0.0862nd of 4Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
CC Barebare harness6 models$0.31020122$0.071$0.939$0.1864th of 4Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Codex Barebare harness3 models$0.09312012$0.004$0.240$0.0761st of 4GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harness8 models$0.17232032$0.005$0.652$0.1533rd of 4Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5

Cost per correct answer divides what a run cost by the number of questions it answered correctly: the vendor-reported figure where the harness reported one, otherwise the benchmark's estimate from the published rate tables over the run's four token columns. Each column is the mean over valid, scored runs at one model (or pooled over the arm's roster in the pooled-models view); TrueArchitect and the bare harnesses are the dashed references, and the data table names each row's cost basis. Full figure, definition and disclosures →

Figure 5 Outcome composition per run

protocols
battery
model
Outcome composition per runhigher is betterZeroShot + MultiTurn · both batteries
025507510093.3%Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5LunaTerraSolTrueArchitectTrueArchitect · 330 runs1st of 489.4%3.9%Haiku 4.5Sonnet 5Opus 4.6Opus 4.8Opus 5Fable 5CC Barebare harness · 201 runs3rd of 485.9%7.4%LunaTerraSolCodex Barebare harness · 120 runs4th of 489.6%3.7%Haiku 4.5Sonnet 5Opus 5LunaTerraSolGrok 4.6Comp 2.5Cursor Barebare harness · 320 runs2nd of 4
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnquestions correct · incorrect · unanswered, per runrunscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect9 models93.3%3303465.0%100.0%8.2%1st of 4Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
CC Barebare harness6 models89.4%2012240.0%100.0%11.5%3rd of 4Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Codex Barebare harness3 models85.9%1201255.0%100.0%13.6%4th of 4GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harness8 models89.6%3203255.0%100.0%11.4%2nd of 4Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5

Every valid run becomes one thin column showing how its questions divided between correct (the arm's color), incorrect (dark) and unanswered or unjudged (black), so an arm's whole record is visible at once. Columns are ordered by model then repetition; the number above each group is the arm's mean accuracy and the number inside is its difference from TrueArchitect. Full figure, definition and disclosures →

Figure 2 Accuracy per run

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Accuracy per runhigher is betterZeroShot + MultiTurn · both batteries
406080100TrueArchitect 93.3%CC Bare 89.4%Codex Bare 85.9%Cursor Bare 89.6%93.3%65.0%100.0%TrueArchitectTrueArchitect1st of 489.4%40.0%100.0%CC Barebare harness3rd of 485.9%55.0%100.0%Codex Barebare harness4th of 489.6%55.0%100.0%Cursor Barebare harness2nd of 4
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnaccuracy (% of questions correct)runscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect9 models93.3%3303465.0%100.0%8.2%1st of 4Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
CC Barebare harness6 models89.4%2012240.0%100.0%11.5%3rd of 4Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Codex Barebare harness3 models85.9%1201255.0%100.0%13.6%4th of 4GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harness8 models89.6%3203255.0%100.0%11.4%2nd of 4Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5

Accuracy is the share of a run's questions whose answer the grading pipeline marked correct, so every dot is one complete run of one arm at one model. The solid tick is the arm's pooled mean, the thin line its worst-to-best range, and dots are translucent so overlap reads as density; hover a dot for its model, repetition and run directory. Full figure, definition and disclosures →

Figure 12 Cost per run

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Cost per runlower is betterZeroShot + MultiTurn · both batteries
$0.00$5.0$10$15TrueArchitect $2.51CC Bare $6.18Codex Bare $1.80Cursor Bare $3.66$2.51$0.097$7.27TrueArchitectTrueArchitect2nd of 4$6.18$2.14$15.0CC Barebare harness4th of 4$1.80$0.110$4.08Codex Barebare harness1st of 4$3.66$0.141$12.9Cursor Barebare harness3rd of 4
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnUSD per runrunscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect9 models$2.5133034$0.097$7.27$1.612nd of 4Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
CC Barebare harness6 models$6.1820122$2.14$15.0$2.664th of 4Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Codex Barebare harness3 models$1.8012012$0.110$4.08$1.281st of 4GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harness8 models$3.6632032$0.141$12.9$3.053rd of 4Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5

Every dot is what one run cost (vendor-reported where available, otherwise estimated from the rate tables), so the spread of an arm's price is visible, not just its mean. The solid tick is the pooled mean and the thin line the range; hover a dot for its model, repetition and run directory. Full figure, definition and disclosures →

Figure 13 Tool result tokens per run

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Tool result tokens per runlower is betterZeroShot + MultiTurn · both batteries
0500k1.0MCC Bare 125kCodex Bare 240k170k59k292kHaiku 4.580k21k151kSonnet 576k26k132kOpus 4.666k23k101kOpus 4.897k42k162kOpus 548k21k77kFable 5269k119k413kLuna186k77k269kTerra178k97k276kSolTrueArchitectTrueArchitect125k3k536kCC Barebare harness1st of 2240k35k496kCodex Barebare harness2nd of 2142k42k617kHaiku 4.574k19k156kSonnet 596k25k370kOpus 5301k44k883kLuna453k50k1.29MTerra718k52k1.22MSol210k66k585kGrok 4.6248k89k600kComp 2.5Cursor Barebare harness

Raw token counts are never pooled across vendors (tokenizers differ), so arms whose selected models span vendors are shown one column per model; ★ marks the best value at each model.

Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumntool result tokens per run (estimated)runscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitectHaiku 4.5170k39459k292k83kHaiku 4.5
Sonnet 580k40421k151k43kSonnet 5
Opus 4.676k40426k132k34kOpus 4.6
Opus 4.866k40423k101k25kOpus 4.8
Opus 597k40442k162k40kOpus 5
Fable 548k11221k77k22kFable 5
GPT 5.6 Luna269k404119k413k82kGPT 5.6 Luna
GPT 5.6 Terra186k40477k269k57kGPT 5.6 Terra
GPT 5.6 Sol178k40497k276k48kGPT 5.6 Sol
CC Barebare harness6 models125k201223k536k147k1st of 2Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Codex Barebare harness3 models240k1201235k496k156k2nd of 2GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harnessHaiku 4.5142k40442k617k112kHaiku 4.5
Sonnet 574k40419k156k45kSonnet 5
Opus 596k40425k370k74kOpus 5
GPT 5.6 Luna301k40444k883k222kGPT 5.6 Luna
GPT 5.6 Terra453k40450k1.29M280kGPT 5.6 Terra
GPT 5.6 Sol718k40452k1.22M368kGPT 5.6 Sol
Grok 4.6210k40466k585k132kGrok 4.6
Composer 2.5248k40489k600k146kComposer 2.5

Tool result tokens are the volume of text the arm's tools pushed back into the model's context over a run — file contents, search hits, shell output, or an index query's result — estimated per call from the result size. Every dot is one run, split per model for every arm; the solid tick is the mean and the thin line the range. Full figure, definition and disclosures →

Figure 10 Pass rate by question category and difficulty

protocol
battery
model

Each cell pools the arm's Anthropic models with equal weight per model, so every arm is compared over the same vendor's models; an arm that ran none of them shows a gap.

Pass rate by question category and difficultyhigher is betterZeroShot · both batteries · Anthropic models
absence-proofaggregationcallscollisioncross-stackdbdispatchendpointsfield-grainimpactmulti-hopnameless-entryreality-checktypesdifficulty 1difficulty 2difficulty 3difficulty 4TrueArchitectTrueArchitect828496929999991006798911009798100989891CC Barebare harness766089569696849463828889849498948877Cursor Barebare harness837098629910095100809798100939810098988736 pts+36 ptscolor = difference from bare Claude Code in the same column · number = pass rate %
Data behind this figure ZeroShot · both batteries · Anthropic models
armcolumnpass rateattemptscorrectmodels
ta-ask-fz-009absence-proof82.0%100825
ta-ask-fz-009aggregation84.0%100845
ta-ask-fz-009calls96.0%4484296
ta-ask-fz-009collision92.0%100925
ta-ask-fz-009cross-stack98.5%4484416
ta-ask-fz-009db99.2%4484446
ta-ask-fz-009dispatch99.0%100995
ta-ask-fz-009endpoints99.8%5045036
ta-ask-fz-009field-grain67.0%100675
ta-ask-fz-009impact98.1%3363296
ta-ask-fz-009multi-hop91.0%100915
ta-ask-fz-009nameless-entry100.0%1501505
ta-ask-fz-009reality-check97.2%2502435
ta-ask-fz-009types98.0%2802746
ta-ask-fz-009difficulty 1100.0%5605606
ta-ask-fz-009difficulty 298.0%5605486
ta-ask-fz-009difficulty 397.5%5605456
ta-ask-fz-009difficulty 490.8%10009085
coldabsence-proof76.0%100765
coldaggregation60.0%100605
coldcalls89.3%3603185
coldcollision56.0%100565
coldcross-stack96.3%3603455
colddb96.3%3603455
colddispatch84.0%100845
coldendpoints93.8%4053775
coldfield-grain63.0%100635
coldimpact81.7%2702165
coldmulti-hop88.0%100885
coldnameless-entry88.7%1501335
coldreality-check84.4%2502115
coldtypes94.4%2252115
colddifficulty 198.4%4504425
colddifficulty 294.0%4504205
colddifficulty 388.0%4503915
colddifficulty 477.1%10007715
cursorabsence-proof83.3%60503
cursoraggregation70.0%60423
cursorcalls97.5%2402343
cursorcollision61.7%60373
cursorcross-stack98.8%2402373
cursordb100.0%2402403
cursordispatch95.0%60573
cursorendpoints99.6%2702693
cursorfield-grain80.0%60483
cursorimpact97.2%1801753
cursormulti-hop98.3%60593
cursornameless-entry100.0%90903
cursorreality-check92.7%1501393
cursortypes98.0%1501473
cursordifficulty 1100.0%3003003
cursordifficulty 298.0%3002943
cursordifficulty 398.0%3002943
cursordifficulty 487.0%6005223

Every question carries one or more categories and a difficulty; each cell is the pass rate of an arm on the questions in that category (or at that difficulty), pooled over the selected models — by default the Anthropic models, which every arm ran. The number in each cell is the pass rate; the color is that cell's difference from bare Claude Code in the same column — blue above, orange below, neutral at the baseline — so a row reads as where an arm beats or trails the null hypothesis. Full figure, definition and disclosures →