Two numbers from Vercel's September AI Gateway Production Index only make sense together. In August, open-weight models processed 56% of every token routed through the gateway — the first month they held a majority. They accounted for 14% of what teams paid.
That gap is the whole story of enterprise inference right now. Volume has migrated to models teams can run cheaply; money has not followed it. Nine months ago open weights handled fewer than one token in ten. They have gained share every month since April, climbing from 13% to 56%, and the frontier labs have lost almost none of the revenue.
Key takeaways
- Open-weight models ran 56% of August gateway tokens but drew only 14% of spend, after sitting at 7% of volume in December 2025.
- Average price per token fell 23.2% in August, a third consecutive monthly decline, with the median high-volume team paying 7.6% less.
- Anthropic kept 64 cents of every dollar spent on the gateway even as its flagship Fable 5 lost two-thirds of its spend share to the cheaper Opus 5.
What cheaper tokens are actually buying
The price collapse is the clearest signal in the report. Average cost per token dropped 23.2% during August — the third straight monthly decline and the sharpest since April. Among teams that pushed more than ten million tokens in both July and August, the median one paid 7.6% less per token, compared with a 2.9% decline the month before. The rate of decline itself is accelerating.
Vercel's reading is that capability has finally caught up at the cheap end, so routine production work no longer justifies frontier pricing. Teams get more inference from a fixed budget and reserve the expensive tier for the jobs that repay it. That is a structural change in how budgets are allocated, not a temporary discount.
Why losing its flagship cost Anthropic nothing
Fable is the most capable model Anthropic sells; Opus sits one tier below at roughly half the price. When the US export control on Fable 5 was lifted and access resumed on July 1, its share of gateway spend jumped to 13.2%. Opus 5 arrived at the end of that same month, and by August Fable 5 was down to 4.9% while Opus 5 had climbed to 22.5%.
Nine in ten teams running Fable cut back, and more of them landed on Opus 5 than on anything else. The verdict was about value, not quality — double the price did not buy a proportional result on the workloads these teams were running. Crucially, the traffic stepped down a tier rather than crossing to a competitor. Anthropic has taken at least 61 cents of every gateway dollar every month since December, 64 cents of it in August, and has held the top two spots by spend throughout that stretch even as the models occupying them changed.
Google's Gemini 3 Flash lost 95% of its volume
The counterexample is severe. Gemini 3 Flash has shed 95% of its share of gateway tokens since May, and more than three-quarters of the volume it lost went to other labs entirely — Google's own full-size Flash successors recovered under a tenth of it.
Where it went is instructive. Roughly half moved to cheaper options, led by GPT-5.6 Luna at less than half the price per token. Most of the rest went the other direction, to Claude Opus 5 and Sonnet 5 at roughly nine and three times Gemini 3 Flash's price. Teams were not chasing a price point; they were chasing fit, and they were willing to move in either direction to find it. Google's overall share of gateway token volume fell from 30% to 5% over the period, with Gemini 3 Flash alone responsible for 22 of the 25 points lost.
The same pattern shows up on the way up. Within five days of Z.ai shipping GLM-5.3-Flash, it was running three times GLM-5.2's daily volume, and by August 31 it carried two-thirds of Z.ai's tokens. Continuity of model profile, not brand loyalty, is what keeps traffic in place — a dynamic also visible in how quickly routing layers have reorganized around Chinese open-weight releases.
Astra outspent Fable 5.1 two to one at the top
The report's special section covers the premium tier, where two flagships launched two days apart at the same price. Anthropic shipped Fable 5.1 on September 1; OpenAI put GPT-6 Astra on the gateway on September 3 at two and a half times the price of GPT-5.6 Sol.
Astra took one in every three dollars spent on OpenAI models within two days, and its share has stayed between 28% and 39% since. Measured over each model's first twelve days, Astra collected 7.7% of all gateway spend against Fable 5.1's 3.7%, and was used by twice as many teams. Inside OpenAI's own lineup the split is stark: Astra and Sol handled 27% of tokens but 71% of spending from September 4 to 16, while Luna and Nano moved more than twice the tokens for about a ninth of the cost.
Elsewhere in the data, Google's Nano Banana led image spend for the first time at 50% against GPT Image's 44%, even though GPT Image generated more images, 46% to 39%. Veo rose to second in video spend at 20%, up from 15% in July, behind Seedance. The share of videos generated by xAI's Grok Imagine has more than halved since June, sliding from 42% to 31% to 19%.
How much weight the figures carry
These are anonymized aggregates from traffic that Vercel routes, not an industry census, and the caveats matter. Spend is estimated from labs' published list prices, so negotiated enterprise rates are invisible. Token counts include reasoning and cached-input tokens. The open-weight classification is broader than in earlier editions of the report, which inflates comparisons against the oldest figures. Prior months get revised as methodology changes.
Even discounted, the direction holds. If open weights carry the majority of production volume while frontier models keep the majority of revenue, the two markets have split: one competing on price for routine work, the other on capability for the minority of tasks that justify it. Next month's index will show whether Astra's early lead over Fable 5.1 was a launch effect or a durable position.
FAQ
What is the AI Gateway Production Index?
It is a monthly report from Vercel based on anonymized, aggregate traffic routed through its AI Gateway, which sits between production applications and AI labs. The September 2026 edition covers data collected through August 2026, with the special report on GPT-6 Astra extending into mid-September.
Do open-weight models now dominate AI spending?
No. They ran 56% of gateway tokens in August but accounted for just 14% of spend. Frontier closed-weight models still take the large majority of the money because they cost far more per token and handle the workloads teams are willing to pay a premium for.
Why did Anthropic keep its revenue share after Fable 5 declined?
Because the workloads moved down a tier rather than to another lab. Fable 5 fell from 13.2% to 4.9% of gateway spend in August while Opus 5, priced at roughly half, rose to 22.5%. Anthropic ended the month with 64% of all gateway spend.






