TrueArchitect vs Harnesses
The null hypotheses of each ecosystem — Claude Code with nothing added, OpenAI's Codex CLI, and Cursor Agent through its SDK — beside TrueArchitect across every model any of them ran.
Every column is one model, so each read is harness against harness at a fixed model: the model is held constant and only the tooling around it changes.
Arms: TrueArchitect · CC Bare · Codex Bare · Cursor Bare. Models: Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5.1242 valid scored runs. Figures open with one column per arm (pooled models); the Columns control switches every figure to one column per model, for like-for-like reads at a fixed model, and the choice is remembered site-wide. The defaults pool the ZeroShot and MultiTurn protocols over both batteries (each protocol × battery × model cell one vote). The arm and model set is fixed to this comparison.
Figure 1 Context tokens per run
Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.
Raw token counts are never pooled across vendors (tokenizers differ), so arms whose selected models span vendors are shown one column per model; ★ marks the best value at each model.
Data behind this figure — ZeroShot + MultiTurn · both batteries
| arm | role | column | context tokens per run (input + cache read) | runs | cells | min | max | sd | rank | runs in the repository |
|---|---|---|---|---|---|---|---|---|---|---|
| TrueArchitect | TrueArchitect | Haiku 4.5 ★ | 8.75M | 39 | 4 | 2.33M | 20.85M | 6.00M | — | Haiku 4.5 |
| Sonnet 5 ★ | 3.37M | 40 | 4 | 1.54M | 7.77M | 1.87M | Sonnet 5 | |||
| Opus 4.6 ★ | 2.63M | 40 | 4 | 786k | 7.38M | 1.96M | Opus 4.6 | |||
| Opus 4.8 ★ | 2.86M | 40 | 4 | 884k | 7.66M | 2.13M | Opus 4.8 | |||
| Opus 5 ★ | 3.24M | 40 | 4 | 1.34M | 5.90M | 1.52M | Opus 5 | |||
| Fable 5 ★ | 1.90M | 11 | 2 | 793k | 3.17M | 1.01M | Fable 5 | |||
| GPT 5.6 Luna ★ | 4.35M | 40 | 4 | 774k | 14.24M | 3.95M | GPT 5.6 Luna | |||
| GPT 5.6 Terra ★ | 2.85M | 40 | 4 | 603k | 8.14M | 2.44M | GPT 5.6 Terra | |||
| GPT 5.6 Sol | 2.58M | 40 | 4 | 664k | 7.91M | 2.10M | GPT 5.6 Sol | |||
| CC Bare | bare harness | 6 models | 8.59M | 201 | 22 | 3.36M | 29.61M | 4.69M | 2nd of 2 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Codex Bare | bare harness | 3 models | 4.55M | 120 | 12 | 2.15M | 10.61M | 2.03M | 1st of 2 | GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| Cursor Bare | bare harness | Haiku 4.5 | 13.83M | 40 | 4 | 2.32M | 23.84M | 5.11M | — | Haiku 4.5 |
| Sonnet 5 | 9.22M | 40 | 4 | 5.15M | 16.79M | 3.26M | Sonnet 5 | |||
| Opus 5 | 8.27M | 40 | 4 | 6.57M | 16.43M | 2.11M | Opus 5 | |||
| GPT 5.6 Luna | 6.18M | 40 | 4 | 4.18M | 10.70M | 1.46M | GPT 5.6 Luna | |||
| GPT 5.6 Terra | 3.30M | 40 | 4 | 1.23M | 5.89M | 1.26M | GPT 5.6 Terra | |||
| GPT 5.6 Sol ★ | 2.47M | 40 | 4 | 985k | 5.16M | 1.28M | GPT 5.6 Sol | |||
| Grok 4.6 ★ | 6.58M | 40 | 4 | 4.29M | 9.61M | 1.85M | Grok 4.6 | |||
| Composer 2.5 ★ | 9.38M | 40 | 4 | 5.02M | 14.81M | 3.53M | Composer 2.5 |
Context tokens are the tokens the model read to answer a whole run: the uncached input plus every cache read, summed over every API call in the run, all threads included. Each column is the mean over the valid, scored runs of one arm at one model, every arm split by model because token counts are model facts; the dashed lines mark TrueArchitect's and each bare harness's pooled value wherever pooling is licensed. Full figure, definition and disclosures →
Figure 11 Cost per correct answer
Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.
Data behind this figure — ZeroShot + MultiTurn · both batteries
| arm | role | column | USD per correct answer | runs | cells | min | max | sd | rank | runs in the repository |
|---|---|---|---|---|---|---|---|---|---|---|
| TrueArchitect | TrueArchitect | 9 models | $0.115 | 330 | 34 | $0.004 | $0.424 | $0.086 | 2nd of 4 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| CC Bare | bare harness | 6 models | $0.310 | 201 | 22 | $0.071 | $0.939 | $0.186 | 4th of 4 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Codex Bare | bare harness | 3 models | $0.093 | 120 | 12 | $0.004 | $0.240 | $0.076 | 1st of 4 | GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| Cursor Bare | bare harness | 8 models | $0.172 | 320 | 32 | $0.005 | $0.652 | $0.153 | 3rd of 4 | Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5 |
Cost per correct answer divides what a run cost by the number of questions it answered correctly: the vendor-reported figure where the harness reported one, otherwise the benchmark's estimate from the published rate tables over the run's four token columns. Each column is the mean over valid, scored runs at one model (or pooled over the arm's roster in the pooled-models view); TrueArchitect and the bare harnesses are the dashed references, and the data table names each row's cost basis. Full figure, definition and disclosures →
Figure 5 Outcome composition per run
Data behind this figure — ZeroShot + MultiTurn · both batteries
| arm | role | column | questions correct · incorrect · unanswered, per run | runs | cells | min | max | sd | rank | runs in the repository |
|---|---|---|---|---|---|---|---|---|---|---|
| TrueArchitect | TrueArchitect | 9 models | 93.3% | 330 | 34 | 65.0% | 100.0% | 8.2% | 1st of 4 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| CC Bare | bare harness | 6 models | 89.4% | 201 | 22 | 40.0% | 100.0% | 11.5% | 3rd of 4 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Codex Bare | bare harness | 3 models | 85.9% | 120 | 12 | 55.0% | 100.0% | 13.6% | 4th of 4 | GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| Cursor Bare | bare harness | 8 models | 89.6% | 320 | 32 | 55.0% | 100.0% | 11.4% | 2nd of 4 | Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5 |
Every valid run becomes one thin column showing how its questions divided between correct (the arm's color), incorrect (dark) and unanswered or unjudged (black), so an arm's whole record is visible at once. Columns are ordered by model then repetition; the number above each group is the arm's mean accuracy and the number inside is its difference from TrueArchitect. Full figure, definition and disclosures →
Figure 2 Accuracy per run
Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.
Data behind this figure — ZeroShot + MultiTurn · both batteries
| arm | role | column | accuracy (% of questions correct) | runs | cells | min | max | sd | rank | runs in the repository |
|---|---|---|---|---|---|---|---|---|---|---|
| TrueArchitect | TrueArchitect | 9 models | 93.3% | 330 | 34 | 65.0% | 100.0% | 8.2% | 1st of 4 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| CC Bare | bare harness | 6 models | 89.4% | 201 | 22 | 40.0% | 100.0% | 11.5% | 3rd of 4 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Codex Bare | bare harness | 3 models | 85.9% | 120 | 12 | 55.0% | 100.0% | 13.6% | 4th of 4 | GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| Cursor Bare | bare harness | 8 models | 89.6% | 320 | 32 | 55.0% | 100.0% | 11.4% | 2nd of 4 | Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5 |
Accuracy is the share of a run's questions whose answer the grading pipeline marked correct, so every dot is one complete run of one arm at one model. The solid tick is the arm's pooled mean, the thin line its worst-to-best range, and dots are translucent so overlap reads as density; hover a dot for its model, repetition and run directory. Full figure, definition and disclosures →
Figure 12 Cost per run
Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.
Data behind this figure — ZeroShot + MultiTurn · both batteries
| arm | role | column | USD per run | runs | cells | min | max | sd | rank | runs in the repository |
|---|---|---|---|---|---|---|---|---|---|---|
| TrueArchitect | TrueArchitect | 9 models | $2.51 | 330 | 34 | $0.097 | $7.27 | $1.61 | 2nd of 4 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| CC Bare | bare harness | 6 models | $6.18 | 201 | 22 | $2.14 | $15.0 | $2.66 | 4th of 4 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Codex Bare | bare harness | 3 models | $1.80 | 120 | 12 | $0.110 | $4.08 | $1.28 | 1st of 4 | GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| Cursor Bare | bare harness | 8 models | $3.66 | 320 | 32 | $0.141 | $12.9 | $3.05 | 3rd of 4 | Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5 |
Every dot is what one run cost (vendor-reported where available, otherwise estimated from the rate tables), so the spread of an arm's price is visible, not just its mean. The solid tick is the pooled mean and the thin line the range; hover a dot for its model, repetition and run directory. Full figure, definition and disclosures →
Figure 13 Tool result tokens per run
Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.
Raw token counts are never pooled across vendors (tokenizers differ), so arms whose selected models span vendors are shown one column per model; ★ marks the best value at each model.
Data behind this figure — ZeroShot + MultiTurn · both batteries
| arm | role | column | tool result tokens per run (estimated) | runs | cells | min | max | sd | rank | runs in the repository |
|---|---|---|---|---|---|---|---|---|---|---|
| TrueArchitect | TrueArchitect | Haiku 4.5 | 170k | 39 | 4 | 59k | 292k | 83k | — | Haiku 4.5 |
| Sonnet 5 | 80k | 40 | 4 | 21k | 151k | 43k | Sonnet 5 | |||
| Opus 4.6 ★ | 76k | 40 | 4 | 26k | 132k | 34k | Opus 4.6 | |||
| Opus 4.8 ★ | 66k | 40 | 4 | 23k | 101k | 25k | Opus 4.8 | |||
| Opus 5 | 97k | 40 | 4 | 42k | 162k | 40k | Opus 5 | |||
| Fable 5 ★ | 48k | 11 | 2 | 21k | 77k | 22k | Fable 5 | |||
| GPT 5.6 Luna ★ | 269k | 40 | 4 | 119k | 413k | 82k | GPT 5.6 Luna | |||
| GPT 5.6 Terra ★ | 186k | 40 | 4 | 77k | 269k | 57k | GPT 5.6 Terra | |||
| GPT 5.6 Sol ★ | 178k | 40 | 4 | 97k | 276k | 48k | GPT 5.6 Sol | |||
| CC Bare | bare harness | 6 models | 125k | 201 | 22 | 3k | 536k | 147k | 1st of 2 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Codex Bare | bare harness | 3 models | 240k | 120 | 12 | 35k | 496k | 156k | 2nd of 2 | GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| Cursor Bare | bare harness | Haiku 4.5 ★ | 142k | 40 | 4 | 42k | 617k | 112k | — | Haiku 4.5 |
| Sonnet 5 ★ | 74k | 40 | 4 | 19k | 156k | 45k | Sonnet 5 | |||
| Opus 5 ★ | 96k | 40 | 4 | 25k | 370k | 74k | Opus 5 | |||
| GPT 5.6 Luna | 301k | 40 | 4 | 44k | 883k | 222k | GPT 5.6 Luna | |||
| GPT 5.6 Terra | 453k | 40 | 4 | 50k | 1.29M | 280k | GPT 5.6 Terra | |||
| GPT 5.6 Sol | 718k | 40 | 4 | 52k | 1.22M | 368k | GPT 5.6 Sol | |||
| Grok 4.6 ★ | 210k | 40 | 4 | 66k | 585k | 132k | Grok 4.6 | |||
| Composer 2.5 ★ | 248k | 40 | 4 | 89k | 600k | 146k | Composer 2.5 |
Tool result tokens are the volume of text the arm's tools pushed back into the model's context over a run — file contents, search hits, shell output, or an index query's result — estimated per call from the result size. Every dot is one run, split per model for every arm; the solid tick is the mean and the thin line the range. Full figure, definition and disclosures →
Figure 10 Pass rate by question category and difficulty
Each cell pools the arm's Anthropic models with equal weight per model, so every arm is compared over the same vendor's models; an arm that ran none of them shows a gap.
Data behind this figure — ZeroShot · both batteries · Anthropic models
| arm | column | pass rate | attempts | correct | models |
|---|---|---|---|---|---|
| ta-ask-fz-009 | absence-proof | 82.0% | 100 | 82 | 5 |
| ta-ask-fz-009 | aggregation | 84.0% | 100 | 84 | 5 |
| ta-ask-fz-009 | calls | 96.0% | 448 | 429 | 6 |
| ta-ask-fz-009 | collision | 92.0% | 100 | 92 | 5 |
| ta-ask-fz-009 | cross-stack | 98.5% | 448 | 441 | 6 |
| ta-ask-fz-009 | db | 99.2% | 448 | 444 | 6 |
| ta-ask-fz-009 | dispatch | 99.0% | 100 | 99 | 5 |
| ta-ask-fz-009 | endpoints | 99.8% | 504 | 503 | 6 |
| ta-ask-fz-009 | field-grain | 67.0% | 100 | 67 | 5 |
| ta-ask-fz-009 | impact | 98.1% | 336 | 329 | 6 |
| ta-ask-fz-009 | multi-hop | 91.0% | 100 | 91 | 5 |
| ta-ask-fz-009 | nameless-entry | 100.0% | 150 | 150 | 5 |
| ta-ask-fz-009 | reality-check | 97.2% | 250 | 243 | 5 |
| ta-ask-fz-009 | types | 98.0% | 280 | 274 | 6 |
| ta-ask-fz-009 | difficulty 1 | 100.0% | 560 | 560 | 6 |
| ta-ask-fz-009 | difficulty 2 | 98.0% | 560 | 548 | 6 |
| ta-ask-fz-009 | difficulty 3 | 97.5% | 560 | 545 | 6 |
| ta-ask-fz-009 | difficulty 4 | 90.8% | 1000 | 908 | 5 |
| cold | absence-proof | 76.0% | 100 | 76 | 5 |
| cold | aggregation | 60.0% | 100 | 60 | 5 |
| cold | calls | 89.3% | 360 | 318 | 5 |
| cold | collision | 56.0% | 100 | 56 | 5 |
| cold | cross-stack | 96.3% | 360 | 345 | 5 |
| cold | db | 96.3% | 360 | 345 | 5 |
| cold | dispatch | 84.0% | 100 | 84 | 5 |
| cold | endpoints | 93.8% | 405 | 377 | 5 |
| cold | field-grain | 63.0% | 100 | 63 | 5 |
| cold | impact | 81.7% | 270 | 216 | 5 |
| cold | multi-hop | 88.0% | 100 | 88 | 5 |
| cold | nameless-entry | 88.7% | 150 | 133 | 5 |
| cold | reality-check | 84.4% | 250 | 211 | 5 |
| cold | types | 94.4% | 225 | 211 | 5 |
| cold | difficulty 1 | 98.4% | 450 | 442 | 5 |
| cold | difficulty 2 | 94.0% | 450 | 420 | 5 |
| cold | difficulty 3 | 88.0% | 450 | 391 | 5 |
| cold | difficulty 4 | 77.1% | 1000 | 771 | 5 |
| cursor | absence-proof | 83.3% | 60 | 50 | 3 |
| cursor | aggregation | 70.0% | 60 | 42 | 3 |
| cursor | calls | 97.5% | 240 | 234 | 3 |
| cursor | collision | 61.7% | 60 | 37 | 3 |
| cursor | cross-stack | 98.8% | 240 | 237 | 3 |
| cursor | db | 100.0% | 240 | 240 | 3 |
| cursor | dispatch | 95.0% | 60 | 57 | 3 |
| cursor | endpoints | 99.6% | 270 | 269 | 3 |
| cursor | field-grain | 80.0% | 60 | 48 | 3 |
| cursor | impact | 97.2% | 180 | 175 | 3 |
| cursor | multi-hop | 98.3% | 60 | 59 | 3 |
| cursor | nameless-entry | 100.0% | 90 | 90 | 3 |
| cursor | reality-check | 92.7% | 150 | 139 | 3 |
| cursor | types | 98.0% | 150 | 147 | 3 |
| cursor | difficulty 1 | 100.0% | 300 | 300 | 3 |
| cursor | difficulty 2 | 98.0% | 300 | 294 | 3 |
| cursor | difficulty 3 | 98.0% | 300 | 294 | 3 |
| cursor | difficulty 4 | 87.0% | 600 | 522 | 3 |
Every question carries one or more categories and a difficulty; each cell is the pass rate of an arm on the questions in that category (or at that difficulty), pooled over the selected models — by default the Anthropic models, which every arm ran. The number in each cell is the pass rate; the color is that cell's difference from bare Claude Code in the same column — blue above, orange below, neutral at the baseline — so a row reads as where an arm beats or trails the null hypothesis. Full figure, definition and disclosures →