TTFX GOLDRUN · DURING-MORTEM · WRITTEN MID-RUN, NOT AFTER
every dollar of self-inflicted cost_
This run's harness was improved while it ran, and improvements cost money.
None of that waste is edited out: it is in the ledgers, the totals, and the projections you see on the
dashboard. This page lists it, so the conservatism is checkable, not claimed.
cost-heavy changes pushed mid-run
- The render gate (mid-run): "kept" was retightened from compiles to
compiles AND emits ANSI-styled animation. Already-paid unstyled effects were dropped and re-earned at
full price (ds lost 7 in one stroke). Every re-earn dollar stays in the totals.
- Reasoning-budget truncation, three variants: deepseek (0 chars at 8k),
grok (output pinned at exactly its 16k cap), sonnet-5 (Anthropic max_tokens covers thinking+text — files cut
mid-fence). ~$7 of paid, unusable responses before each was diagnosed. All on the ledgers.
- grok paid ~$1.90 asking a model to write a file the harness generates for free
(effects/mod.rs). Harness bug, our cost, kept.
- Duplicate-process era: a bad kill pattern briefly ran two copies of some models —
double-bought cores. Kept.
- Restart re-cores: fable bought its core 3x, sonnet 2x, across harness fixes. Kept.
- Fat-chunk era: early asks shipped 22–33k prompt tokens (60-line chunks) before
rechunking to the sidecar's ~700-char geometry. Those expensive asks are in the ledgers.
why the projections are conservative
- Projections use the blended rate — total spend (all waste above included)
divided by kept effects, scaled to 37. Marginal rates are far lower: sonnet-5's last effects cost ~$0.12
against a blended ~$0.60. A clean run hits the projected number with a wide margin.
- The "without leCore" column is measured, and measured LOW: real cold+warm
corpus sends, priced by provider bills — and the measurements found no cache discount for
Anthropic-by-default or DeepSeek-via-OR (warm = 1.0x cold). We still credit the full-context arm with the same
keep-rate we achieved, though several models cannot even fit the 396k corpus in their windows.
- DHH's side is his published successes only — $23–550 per model. His failed runs
(DSV4 Flash, GPT Luna) cost real money too; we count none of that against him.
- All measurement overhead is ours: the $31.76 cold/warm corpus measurement and
every probe came out of our account and none of it is billed against the comparison.
- The quality bar rose mid-run and totals were not reset. The strictest gate is
applied retroactively to spending accrued under weaker gates.
receipts: per-model
ledgers ·
run logs ·
generated source ·
account drain sampled from the provider every 60s (header of the dashboard). swap 'sol' for any model tag.