Frontier economics

Pacing the frontier makes frontier lab economics viable.

Stippled portraits of Dario Amodei and Sam Altman, with a red insect and scattered red and blue dots.

Dario Amodei wants frontier labs to slow capability development together without sacrificing commercial advantage. Labs must fund new capability while recovering what they spent on earlier models and failed experiments. In the cost-conscious market, inference innovation expands margins by lowering the cost of an adequate answer. In the quality-conscious market, margins grow mainly through capital investment in better models.

The financial case for pacing is more time to earn from each generation, with fewer expensive attempts to replace it. Inference innovation improves the margins on existing models. Amodei’s proposal also includes restrictions on China’s access to chips, unauthorized distillation, and model-weight theft.

01

Amodei wants more safety time without sacrificing commercial advantage.

He proposes government coordination and capability checkpoints. At Anthropic, he commits to giving independent reviewers access and publication rights. [1]

He also calls for restrictions on Chinese access to compute, unauthorized distillation, and model-weight theft. He identifies distillation as a way for lagging companies to catch up at a fraction of the cost of independent development. [1]

02

Astra scores higher than Fable 5 at one-fifth the benchmark cost.

On September 13, GPT-6 Astra at high effort scores 51.05 at $1.72 per benchmark task; Fable 5 with fallback scores 49.70 at $8.75. [2]

These figures use Intelligence Index v4.3. They measure benchmark performance and task cost; actual customer switching and provider profitability remain unknown.

Historical price decline · Through 3 September 2026 · Index v4.1.1

Estimated time for the minimum price of comparable capability to halve: 46 days.

Scale

Each line tracks the lowest available price in a flagship model’s five-point score band, starting at that flagship’s launch. Prices are divided by their starting values, so 0.5 means half the original price.

Twelve model price histories, normalized to their starting prices. The five older histories fall to between 3% and 7% of their starting prices over 98 to 294 days; the average fitted price half-life is about 46 days.
This earlier snapshot estimates how quickly comparable benchmark capability became cheaper. It does not measure changes in a lab’s revenue or profit. The September 13 comparisons above use a newer benchmark; their raw scores and task costs cannot be compared with this history. Dots mark price events, straight lines connect them, and open circles mark the end of each history. The dotted black curve is the average decline rate. Artificial Analysis data via CatalystNeuro ↗
Method, model histories & data

The estimate uses five releases with at least 90 days of history. Each release contributes its full history, from 98 to 294 days; the cutoff does not shorten any line. Newer releases are shown with all available points, including the September 1–3 launches. GPT-5 is excluded from the estimate.

We fit an exponential decline to each older release’s daily price history, using log prices and a starting value of 1. We average the five decline rates and convert that rate to a half-life: 45.82 days. The displayed straight segments interpolate between price events.

Three releases in the estimate share the 55–60 score band, so their histories overlap. A band can include models scoring below the flagship. Requiring a model to match or beat the flagship’s exact score gives a half-life of 50 days; this is a sensitivity check, not an uncertainty interval.

The starting price can already be below the flagship’s own price. Before August 19, most prices are later observations dated back to launch, with some known price cuts added. Earlier prices and availability were not fully recorded. Scores use Artificial Analysis v4.1.1: the September 3 snapshot, plus Astra’s launch-day measurements from the September 4 morning archive. GPT-5’s price remains in the market comparison data.

Full histories in this snapshot
ReleaseScore bandDaysIn fit
GPT-5.135–40294Yes
GPT-5.450–55182Yes
GPT-5.555–60133Yes
Opus 4.755–60140Yes
Opus 4.855–6098Yes
GPT-5.6 Sol60–6556No
Fable 560–6586No
Opus 560–6541No
Fable 5.165–702No
Gemini 3.8 Flash55–601No
Muse Spark 1.360–651No
GPT-6 Astra60–650No

Download the plotted data (JSON) · Benchmark methodology ↗

The current comparisons come from the September 13 Pareto-frontier snapshot, covering 41 configurations on Index v4.3. We checked the four configurations cited here against Artificial Analysis’s embedded data. The cost frontier contains 13 configurations from 4 model families. These positions identify benchmark tradeoffs, not profitable businesses. Download that snapshot (JSON).

Fable 5.1 at max effort with fallback has the highest score in this snapshot: 53.37 at $7.63 per task. Astra max scores 52.81 at $3.26. Fable costs 2.3 times as much for 0.56 more index points. The benchmark does not establish how much customers would pay for that score difference. Fable 5 uses its reported fallback configuration.

03

Cost-conscious margins grow through inference innovation; quality-conscious margins grow mainly through capital investment.

Cost-conscious

Several models already do the job well enough, so customers choose by total cost per successful task. [4] Providers expand margins through inference innovations that reduce serving costs. Better memory management, for example, lets the same hardware serve more requests. [5]

Quality-conscious

Customers pay more for results that cheaper models cannot deliver. Labs expand margins mainly by investing capital in better models. The investment funds research, compute, and training; its return depends on producing results worth the higher price.

The resulting earnings must recover spending on successful models as well as failed models and experiments.

Public OpenRouter usage · 14 August–12 September 2026 · Index v4.3

OpenRouter token dollars by Intelligence Index.

Among matched models, the two largest score bands are 40–45 ($60.1M) and 50–55 ($55.9M). Substantial usage remains below the frontier alongside demand for the highest-scoring models. Each model is placed at its highest measured score; estimated scores are marked separately.

Dollar-weighted histogram of OpenRouter public token usage, valued at listed prices. The 40–45 Intelligence Index band totals $60.1 million and the 50–55 band $55.9 million. Another $34.7 million has no matched score. Dark green indicates measured scores, light green estimated scores, and gray unscored models.
Measured scoreAA-estimated scoreUnscored
Token dollars are prompt and completion counts valued at listed prices, not observed bills. Scores cover 87.7% of the $283.0M priced total; $34.7M is shown separately as unscored. Another 6.7% of source tokens lack prices and sit outside that dollar total. Customer motives and provider margins are not measured. Sources: OpenRouter ↗ and Artificial Analysis ↗.
Calculation, coverage & data

We use all 631 routes in OpenRouter’s public monthly rankings snapshot, covering 14 August through 12 September 2026. For each route, value equals prompt tokens times its listed input price, plus completion tokens times its listed output price. The September 13 catalog supplies prices. We match exact canonical model versions and retain each billing variant’s price before combining model totals. Free variants contribute $0; missing prices remain unknown.

Scores use Artificial Analysis Intelligence Index v4.3, retrieved September 13. Each exact release receives its highest measured configuration score. When only AA-estimated scores exist, we use the highest estimate and show it in light green. The scored dollar total covers 79 model releases with measured scores and 93 with estimated scores. OpenRouter does not disclose the reasoning settings of these requests: the horizontal axis describes available model capability, not the score of every request.

Bars use fixed five-point intervals starting at zero, with the lower endpoint included and the upper endpoint excluded. Unscored models retain their known dollar value in the gray bar; Tencent Hy4 preview accounts for $28.2M of that bar. We do not substitute scores across unverified model updates, including the April and August DeepSeek V4 Pro releases or the August and September Qwen3.8 Max releases.

These are current base-price valuations of public usage. Actual bills also depend on caching, historical prices, provider routing, long-context tiers, negotiated discounts, and request or media fees. The public snapshot reports zero cache-detail fields throughout; it does not establish that caching was absent. The chart does not measure all OpenRouter traffic or all AI spending.

As a sensitivity check, charging 80% of input tokens at the listed cache-read rate where available reduces the priced total to $100.8M. The 40–45 and 50–55 bands remain the largest. This is an explicit pricing scenario, not observed billing or an uncertainty interval.

Token value by score band
Index bandValue at list pricesShare of priced total
0–5$0.0M0.0%
5–10$1.1M0.4%
10–15$1.9M0.7%
15–20$3.4M1.2%
20–25$9.1M3.2%
25–30$9.6M3.4%
30–35$44.6M15.8%
35–40$38.9M13.8%
40–45$60.1M21.2%
45–50$23.7M8.4%
50–55$55.9M19.7%
Unscored$34.7M12.3%

Download the calculation and model-level data (JSON) · CSV · Graph (SVG)

Source: OpenRouter Rankings, accessed September 13, 2026; usage through September 12. Rankings data are licensed under CC BY 4.0. OpenRouter data documentation · Artificial Analysis benchmark methodology. The download preserves source URLs, retrieval times, checksums, matching decisions, and rows without prices or scores.

Frontier labs can’t afford to keep fighting for first place, and they can’t afford to be second. Staying ahead requires another round of capital investment. Falling behind threatens the premium needed to repay the last one.

04

Pacing gives frontier investments longer to pay back.

With longer gaps between frontier investments, labs have more time to earn from their existing models. At unchanged prices, inference improvements turn lower serving costs into higher margins.

One option in the AI Futures proposal would allocate 70% of compute to serving customers, 25% to safety, and 5% to developing new capabilities. [3]

The financial test is whether the extra earnings from existing models and lower development spending exceed added safety costs and the earnings sacrificed by delaying better models.