AI Newsway

Ember-1 Costs the Same Per Token as Kimi K3 and Uses 40% Fewer of Them

Fireworks Research says training the model to shorten its reasoning traces β€” rather than turning the reasoning-effort dial down β€” is what preserved accuracy across seven benchmarks and two customers' production coding traffic.

|5 min read0
AI Summary
Fireworks AI released Ember-1 on September 24, a reasoning model trained from Moonshot AI's Kimi K3 that matches the base model's quality using roughly 40% fewer tokens. Its per-token price is unchanged, so savings come purely from shorter reasoning traces. Fireworks reports 35-50% reasoning reduction across seven benchmarks and about 35% fewer tokens per task in two customers' production coding A/B tests. It ships as a two-week research preview, with the figures vendor-reported.
Server racks in a computer room installation β€” Ember-1 targets the inference bill that long reasoning traces run up in multi-turn agentic workloads.
Server racks in a computer room installation β€” Ember-1 targets the inference bill that long reasoning traces run up in multi-turn agentic workloads.

Fireworks AI has released Ember-1, a reasoning model built on Moonshot AI's Kimi K3 that the company says reaches the base model's quality while emitting roughly 40% fewer tokens. The model went live on September 24 as a research preview on Fireworks' serverless platform, and its posted rate card is identical to K3's: $3 per million uncached input tokens, $0.30 per million cached, and $15 per million output. Nothing about the price per token changed; the bill falls only because there are fewer tokens on it.

Key takeaways

  • Ember-1 lists at $3 per million input and $15 per million output tokens β€” the same public rate as the Kimi K3 base model it was trained from β€” with a 1,048,576-token context window.
  • Across seven public benchmarks and two customers' production traffic, Fireworks reports it shortened K3's reasoning by 35–50% without losing accuracy.
  • In live A/B tests on two customers' coding workloads, Ember-1 used about 35% fewer tokens per task at comparable quality, and one of the two has moved it into production.

Why reasoning traces became the expensive part

Reasoning models like K3 spend the majority of their generated tokens — sometimes more than 90% — on internal deliberation rather than on the answer a user reads.

Agentic loops turn that cost into compound interest. Because each turn feeds the whole prior transcript back in, a trace written on turn one is paid for again on every turn after it, and the context window Fireworks describes grows roughly quadratically in turn count.

The obvious workaround does not hold up. Turning K3's reasoning-effort dial down cut tokens but surrendered too much accuracy, according to the company's announcement, so Fireworks set out to train economy in rather than configure it in.

How Fireworks trained the reasoning down

Getting there took more than 50 training experiments and over 200 evaluations on the company's serverless training product, plus new algorithms for compressing chain-of-thought output without losing accuracy. Fireworks says it used its own data, not customer data.

The team deliberately preserved one class of reasoning: self-reflection. Revisiting an assumption or tracing an outcome back to an earlier decision helps a model recover from its own mistakes, so the goal was to keep that while cutting unproductive loops. The training mix was kept wide — maths and coding through to search, tool use and software engineering — a hedge against the obvious failure mode, a model terse only on benchmark-shaped problems.

Across seven public benchmarks and two customers' production traffic, Fireworks reports that K3's reasoning could be shortened by 35–50% with no sacrifice in accuracy. The trained behavior also shows restraint on unsuccessful attempts, cutting short exactly the prolonged loops that run up the largest bills.

What the benchmarks and A/B tests show

Fireworks evaluated Ember-1 on its Specialized Intelligence Index, introduced days earlier, including Doximity's Bedside Bench — 500 physician-validated clinical cases across 10 specialty categories. On cost per task there, the company says Ember-1 set a new Pareto frontier against open and closed models alike, including GPT-5.6 Sol, GPT-6 Astra and Claude Opus 5.

The comparison that matters is against K3 itself at three reasoning-effort settings. On every benchmark with more than 50 samples, Fireworks places Ember-1 on or near the cost-quality frontier, which is another way of saying the cheapest way to run K3 is no longer to tell it to think less.

The production evidence is narrower but more concrete. In live A/B tests on two customers' coding workloads, Ember-1 used roughly 35% fewer tokens per task at comparable quality, with task completion, success scores and failure rates holding or improving. One of the two has since moved it into production and plans to retire the base model. Internally, Fireworks reports its own developers did not notice the switch.

What to watch next

The caveats are structural rather than technical. Ember-1 ships as a two-week research preview that becomes permanent only on community demand, and it is served by a single provider on OpenRouter. Because the savings are volume-based, they shrink on workloads where reasoning is already a small share of output.

K3 has meanwhile become a base that third parties retrain and resell — a shift already visible in Moonshot AI's revenue trajectory. Fireworks is opening training on Ember-1 so enterprises can specialize it further.

FAQ

Is Ember-1 cheaper per token than Kimi K3?

No. Ember-1's posted pricing matches K3's at $3 per million uncached input tokens, $0.30 per million cached input and $15 per million output. The cost reduction comes entirely from generating fewer reasoning tokens for the same task, not from a lower rate.

Is Ember-1 open source?

Fireworks has not released Ember-1's weights. It is offered as a hosted serving option alongside the base Kimi K3 model on Fireworks' serverless platform and through OpenRouter, initially as a research preview.

How much of the 40% figure is independently verified?

The token-reduction numbers come from Fireworks' own benchmark runs and from A/B tests on two unnamed customers' production traffic. No third party has published an independent replication, so the 35–50% range should be read as vendor-reported.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

GLM-5.3 Scores 60 on Artificial Analysis Index at Half the Usual Output Price
LLM & Chatbots

GLM-5.3 Scores 60 on Artificial Analysis Index at Half the Usual Output Price

Artificial Analysis scored Z.ai's GLM-5.3 at 60 on its Intelligence Index, well above the 35 median, at $4.40 per million output tokens. The catch is verbosity.

Seung Jung40 days ago
AWS Becomes the First Cloud to Carry OpenAI's Gated Cyber Models
SaaS & Cloud

AWS Becomes the First Cloud to Carry OpenAI's Gated Cyber Models

Daybreak Red and Blue are now sold through Amazon Bedrock, moving OpenAI's gated cyber models into enterprise cloud procurement and AWS governance.

Seung Jung44 days ago
Grok Bot Lets AI Agents Run Their Own Group Chat β€” and Sign Into Your Accounts
LLM & Chatbots

Grok Bot Lets AI Agents Run Their Own Group Chat β€” and Sign Into Your Accounts

SpaceXAI opened a Grok Bot beta where multiple agents coordinate in group chats, assign ownership to each other, and sign into a user's own accounts.

Seung Jung44 days ago
ChatGPT Ads Reach the UK, Japan, Korea, Brazil and Mexico
LLM & Chatbots

ChatGPT Ads Reach the UK, Japan, Korea, Brazil and Mexico

OpenAI turned on ChatGPT sponsored placements in five more countries on 11 August, limited to logged-in adults on the Free and Go tiers.

Seung Jung44 days ago
Qwen3.8's 27B Open Model Is the Release That Actually Matters
LLM & Chatbots

Qwen3.8's 27B Open Model Is the Release That Actually Matters

Alibaba's Qwen3.8 open weights pair a 2.4T mixture-of-experts model with a 27B dense multimodal model sized for a single 24GB consumer GPU.

Seung Jung44 days ago
Ox Alpha Unmasked: Z.ai Ships GLM-5.3-Flash With MIT Weights and 15-Cent Pricing
LLM & Chatbots

Ox Alpha Unmasked: Z.ai Ships GLM-5.3-Flash With MIT Weights and 15-Cent Pricing

Z.ai confirmed the stealth Ox Alpha model is GLM-5.3-Flash: 320B parameters, 18B active, MIT-licensed weights and sub-dollar output pricing.

Seung Jung31 days ago