Fireworks AI launches Ember-1, saying token usage drops about 40% versus Kimi K3

Fireworks AI launches Ember-1, saying token usage drops about 40% versus Kimi K3

N
News Editor
2026-09-28 07:32:48
Fireworks AI has introduced Ember-1, an inference-optimized model built through post-training on Moonshot AI’s open-weight Kimi K3. The company said the model is designed to preserve task accuracy while producing shorter reasoning traces, cutting token usage by about 40% compared with the original Kimi K3. Ember-1 is currently available only as a research preview through Fireworks’ serverless API. Its weights, training code, and specific algorithms have not been released, and self-hosting is not supported at this stage. Fireworks said the goal is to keep K3’s coding capability while lowering customer costs. In internal testing, the company found that more than 90% of tokens generated by K3 in multi-turn agent workloads were spent on internal reasoning, creating significant cost overhead. In production A/B tests covering programming workloads from two customers, Fireworks said Ember-1 reduced tokens per task by about 35% at similar quality levels. In one test, output tokens fell from 49,300 to 29,900 per task, reasoning tokens dropped 71.3%, total tokens declined 39%, and task scores were broadly unchanged.

Fireworks AI has released Ember-1, an inference-optimized model built through post-training on Moonshot AI’s open-weight Kimi K3.

The company said Ember-1 learns to produce shorter reasoning traces while maintaining task accuracy. Compared with the original Kimi K3, Fireworks said the model cuts token usage by about 40%.

Research preview only

For now, Ember-1 can only be deployed through Fireworks’ serverless API as a research preview. The model’s weights, training code, and specific algorithms have not been made public, and self-hosting is not available.

Fireworks said the release is meant to give customers K3’s coding ability at a lower cost.

Internal and production testing

According to Fireworks, internal tests showed that in multi-turn agent workloads, more than 90% of the tokens generated by K3 were used for internal reasoning, making costs significant.

In production A/B tests, programming workloads from two customers showed that Ember-1 reduced tokens per task by about 35% while keeping quality at a similar level. In one test, output tokens per task fell from 49,300 to 29,900, reasoning tokens dropped 71.3%, total tokens fell 39%, and task scores were broadly flat.

The item was reported by MarkTechPost and aggregated by Techub.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.