Comparison · E004

TrueArchitect & OpenAI

TrueArchitect at GPT 5.6 Luna, Terra and Sol beside the two harnesses that run those models natively: OpenAI's own Codex CLI, and Cursor Agent at the same GPT models.

This is the vendor-native comparison: a bare harness built around a model family, against the same family given TrueArchitect's codebase index.

Arms: TrueArchitect · Codex Bare · Cursor Bare. Models: GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol.480 valid scored runs. Figures open with one column per arm (pooled models); the Columns control switches every figure to one column per model, for like-for-like reads at a fixed model, and the choice is remembered site-wide. The defaults pool the ZeroShot and MultiTurn protocols over both batteries (each protocol × battery × model cell one vote). The arm and model set is fixed to this comparison.

Figure 1 Context tokens per run

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Context tokens per runlower is betterZeroShot + MultiTurn · both batteries
02.0M4.0MTrueArchitect 3.26MCodex Bare 4.55MCursor Bare 3.98M3.26M3 modelsTrueArchitectTrueArchitect1st of 34.55M3 modelsCodex Barebare harness3rd of 33.98M3 modelsCursor Barebare harness2nd of 3
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumncontext tokens per run (input + cache read)runscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect3 models3.26M12012603k14.24M3.02M1st of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Codex Barebare harness3 models4.55M120122.15M10.61M2.03M3rd of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harness3 models3.98M12012985k10.70M2.08M2nd of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol

Context tokens are the tokens the model read to answer a whole run: the uncached input plus every cache read, summed over every API call in the run, all threads included. Each column is the mean over the valid, scored runs of one arm at one model, every arm split by model because token counts are model facts; the dashed lines mark TrueArchitect's and each bare harness's pooled value wherever pooling is licensed. Full figure, definition and disclosures →

Figure 11 Cost per correct answer

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Cost per correct answerlower is betterZeroShot + MultiTurn · both batteries
$0.00$0.03$0.05$0.07$0.10TrueArchitect $0.076Codex Bare $0.093Cursor Bare $0.079$0.0763 modelsTrueArchitectTrueArchitect1st of 3$0.0933 modelsCodex Barebare harness3rd of 3$0.0793 modelsCursor Barebare harness2nd of 3
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnUSD per correct answerrunscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect3 models$0.07612012$0.004$0.344$0.0811st of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Codex Barebare harness3 models$0.09312012$0.004$0.240$0.0763rd of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harness3 models$0.07912012$0.005$0.250$0.0652nd of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol

Cost per correct answer divides what a run cost by the number of questions it answered correctly: the vendor-reported figure where the harness reported one, otherwise the benchmark's estimate from the published rate tables over the run's four token columns. Each column is the mean over valid, scored runs at one model (or pooled over the arm's roster in the pooled-models view); TrueArchitect and the bare harnesses are the dashed references, and the data table names each row's cost basis. Full figure, definition and disclosures →

Figure 5 Outcome composition per run

protocols
battery
model
Outcome composition per runhigher is betterZeroShot + MultiTurn · both batteries
025507510089.2%LunaTerraSolTrueArchitectTrueArchitect · 120 runs1st of 385.9%3.3%LunaTerraSolCodex Barebare harness · 120 runs2nd of 384.4%4.8%LunaTerraSolCursor Barebare harness · 120 runs3rd of 3
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnquestions correct · incorrect · unanswered, per runrunscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect3 models89.2%1201270.0%100.0%9.6%1st of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Codex Barebare harness3 models85.9%1201255.0%100.0%13.6%2nd of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harness3 models84.4%1201255.0%100.0%13.8%3rd of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol

Every valid run becomes one thin column showing how its questions divided between correct (the arm's color), incorrect (dark) and unanswered or unjudged (black), so an arm's whole record is visible at once. Columns are ordered by model then repetition; the number above each group is the arm's mean accuracy and the number inside is its difference from TrueArchitect. Full figure, definition and disclosures →

Figure 2 Accuracy per run

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Accuracy per runhigher is betterZeroShot + MultiTurn · both batteries
60708090100TrueArchitect 89.2%Codex Bare 85.9%Cursor Bare 84.4%89.2%70.0%100.0%TrueArchitectTrueArchitect1st of 385.9%55.0%100.0%Codex Barebare harness2nd of 384.4%55.0%100.0%Cursor Barebare harness3rd of 3
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnaccuracy (% of questions correct)runscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect3 models89.2%1201270.0%100.0%9.6%1st of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Codex Barebare harness3 models85.9%1201255.0%100.0%13.6%2nd of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harness3 models84.4%1201255.0%100.0%13.8%3rd of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol

Accuracy is the share of a run's questions whose answer the grading pipeline marked correct, so every dot is one complete run of one arm at one model. The solid tick is the arm's pooled mean, the thin line its worst-to-best range, and dots are translucent so overlap reads as density; hover a dot for its model, repetition and run directory. Full figure, definition and disclosures →

Figure 12 Cost per run

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Cost per runlower is betterZeroShot + MultiTurn · both batteries
$0.00$2.0$4.0$6.0TrueArchitect $1.49Codex Bare $1.80Cursor Bare $1.60$1.49$0.097$5.96TrueArchitectTrueArchitect1st of 3$1.80$0.110$4.08Codex Barebare harness3rd of 3$1.60$0.141$5.10Cursor Barebare harness2nd of 3
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnUSD per runrunscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect3 models$1.4912012$0.097$5.96$1.361st of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Codex Barebare harness3 models$1.8012012$0.110$4.08$1.283rd of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harness3 models$1.6012012$0.141$5.10$1.312nd of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol

Every dot is what one run cost (vendor-reported where available, otherwise estimated from the rate tables), so the spread of an arm's price is visible, not just its mean. The solid tick is the pooled mean and the thin line the range; hover a dot for its model, repetition and run directory. Full figure, definition and disclosures →

Figure 13 Tool result tokens per run

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Tool result tokens per runlower is betterZeroShot + MultiTurn · both batteries
0500k1.0MTrueArchitect 211kCodex Bare 240kCursor Bare 490k211k77k413kTrueArchitectTrueArchitect1st of 3240k35k496kCodex Barebare harness2nd of 3490k44k1.29MCursor Barebare harness3rd of 3
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumntool result tokens per run (estimated)runscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect3 models211k1201277k413k75k1st of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Codex Barebare harness3 models240k1201235k496k156k2nd of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harness3 models490k1201244k1.29M341k3rd of 3GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol

Tool result tokens are the volume of text the arm's tools pushed back into the model's context over a run — file contents, search hits, shell output, or an index query's result — estimated per call from the result size. Every dot is one run, split per model for every arm; the solid tick is the mean and the thin line the range. Full figure, definition and disclosures →

Figure 10 Pass rate by question category and difficulty

protocol
battery
model

Each cell is the arm's pass rate at Anthropic models alone.

Pass rate by question category and difficultyhigher is betterZeroShot · both batteries · Anthropic models
absence-proofaggregationcollisiondispatchfield-grainmulti-hopnameless-entryreality-checkdifficulty 4TrueArchitectTrueArchitect7580889853331008278Codex Barebare harness9050529853501008375Cursor Barebare harness904552885322100817024 pts+24 ptscolor = difference from the column mean in the same column · number = pass rate %
Data behind this figure ZeroShot · both batteries · Anthropic models
armcolumnpass rateattemptscorrectmodels
ta-ask-fz-009absence-proof75.0%60453
ta-ask-fz-009aggregation80.0%60483
ta-ask-fz-009collision88.3%60533
ta-ask-fz-009dispatch98.3%60593
ta-ask-fz-009field-grain53.3%60323
ta-ask-fz-009multi-hop33.3%60203
ta-ask-fz-009nameless-entry100.0%90903
ta-ask-fz-009reality-check82.0%1501233
ta-ask-fz-009difficulty 478.3%6004703
codexabsence-proof90.0%60543
codexaggregation50.0%60303
codexcollision51.7%60313
codexdispatch98.3%60593
codexfield-grain53.3%60323
codexmulti-hop50.0%60303
codexnameless-entry100.0%90903
codexreality-check82.7%1501243
codexdifficulty 475.0%6004503
cursorabsence-proof90.0%60543
cursoraggregation45.0%60273
cursorcollision51.7%60313
cursordispatch88.3%60533
cursorfield-grain53.3%60323
cursormulti-hop21.7%60133
cursornameless-entry100.0%90903
cursorreality-check81.3%1501223
cursordifficulty 470.3%6004223

Every question carries one or more categories and a difficulty; each cell is the pass rate of an arm on the questions in that category (or at that difficulty), pooled over the selected models — by default the Anthropic models, which every arm ran. The number in each cell is the pass rate; the color is that cell's difference from bare Claude Code in the same column — blue above, orange below, neutral at the baseline — so a row reads as where an arm beats or trails the null hypothesis. Full figure, definition and disclosures →