Fireworks AI has released Ember-1, a reasoning model built on Moonshot AI's Kimi K3 that the company says reaches the base model's quality while emitting roughly 40% fewer tokens. The model went live on September 24 as a research preview on Fireworks' serverless platform, and its posted rate card is identical to K3's: $3 per million uncached input tokens, $0.30 per million cached, and $15 per million output. Nothing about the price per token changed; the bill falls only because there are fewer tokens on it.
Key takeaways
- Ember-1 lists at $3 per million input and $15 per million output tokens β the same public rate as the Kimi K3 base model it was trained from β with a 1,048,576-token context window.
- Across seven public benchmarks and two customers' production traffic, Fireworks reports it shortened K3's reasoning by 35–50% without losing accuracy.
- In live A/B tests on two customers' coding workloads, Ember-1 used about 35% fewer tokens per task at comparable quality, and one of the two has moved it into production.
Why reasoning traces became the expensive part
Reasoning models like K3 spend the majority of their generated tokens — sometimes more than 90% — on internal deliberation rather than on the answer a user reads.
Agentic loops turn that cost into compound interest. Because each turn feeds the whole prior transcript back in, a trace written on turn one is paid for again on every turn after it, and the context window Fireworks describes grows roughly quadratically in turn count.
The obvious workaround does not hold up. Turning K3's reasoning-effort dial down cut tokens but surrendered too much accuracy, according to the company's announcement, so Fireworks set out to train economy in rather than configure it in.
How Fireworks trained the reasoning down
Getting there took more than 50 training experiments and over 200 evaluations on the company's serverless training product, plus new algorithms for compressing chain-of-thought output without losing accuracy. Fireworks says it used its own data, not customer data.
The team deliberately preserved one class of reasoning: self-reflection. Revisiting an assumption or tracing an outcome back to an earlier decision helps a model recover from its own mistakes, so the goal was to keep that while cutting unproductive loops. The training mix was kept wide — maths and coding through to search, tool use and software engineering — a hedge against the obvious failure mode, a model terse only on benchmark-shaped problems.
Across seven public benchmarks and two customers' production traffic, Fireworks reports that K3's reasoning could be shortened by 35–50% with no sacrifice in accuracy. The trained behavior also shows restraint on unsuccessful attempts, cutting short exactly the prolonged loops that run up the largest bills.
What the benchmarks and A/B tests show
Fireworks evaluated Ember-1 on its Specialized Intelligence Index, introduced days earlier, including Doximity's Bedside Bench — 500 physician-validated clinical cases across 10 specialty categories. On cost per task there, the company says Ember-1 set a new Pareto frontier against open and closed models alike, including GPT-5.6 Sol, GPT-6 Astra and Claude Opus 5.
The comparison that matters is against K3 itself at three reasoning-effort settings. On every benchmark with more than 50 samples, Fireworks places Ember-1 on or near the cost-quality frontier, which is another way of saying the cheapest way to run K3 is no longer to tell it to think less.
The production evidence is narrower but more concrete. In live A/B tests on two customers' coding workloads, Ember-1 used roughly 35% fewer tokens per task at comparable quality, with task completion, success scores and failure rates holding or improving. One of the two has since moved it into production and plans to retire the base model. Internally, Fireworks reports its own developers did not notice the switch.
What to watch next
The caveats are structural rather than technical. Ember-1 ships as a two-week research preview that becomes permanent only on community demand, and it is served by a single provider on OpenRouter. Because the savings are volume-based, they shrink on workloads where reasoning is already a small share of output.
K3 has meanwhile become a base that third parties retrain and resell — a shift already visible in Moonshot AI's revenue trajectory. Fireworks is opening training on Ember-1 so enterprises can specialize it further.
FAQ
Is Ember-1 cheaper per token than Kimi K3?
No. Ember-1's posted pricing matches K3's at $3 per million uncached input tokens, $0.30 per million cached input and $15 per million output. The cost reduction comes entirely from generating fewer reasoning tokens for the same task, not from a lower rate.
Is Ember-1 open source?
Fireworks has not released Ember-1's weights. It is offered as a hosted serving option alongside the base Kimi K3 model on Fireworks' serverless platform and through OpenRouter, initially as a research preview.
How much of the 40% figure is independently verified?
The token-reduction numbers come from Fireworks' own benchmark runs and from A/B tests on two unnamed customers' production traffic. No third party has published an independent replication, so the 35–50% range should be read as vendor-reported.






