Reflection unveils 501B-parameter Beam, trailing the latest top Chinese open-source models

Reflection unveils 501B-parameter Beam, trailing the latest top Chinese open-source models

N
News Editor
2026-10-06 13:04:18
Nvidia-backed AI startup Reflection has introduced Beam, its first open-weight model. Built on a mixture-of-experts, or MoE, architecture, Beam has 501 billion total parameters and activates 23 billion parameters per inference step. The model is aimed at coding, reasoning and agent-based workloads. Based on metrics released by Reflection, Beam performs close to GLM-5.2 overall, but it does not yet match the newest leading open-source models from China. On DeepSWE, Beam scored 44.4, compared with 44.0 for GLM-5.2 and 51.0 for Qwen 3.8-Max. GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash posted 61.0, 68.0 and 74.2, respectively. Reflection also disclosed the scale of Beam’s training. Pretraining used 6,144 Nvidia GB300 GPUs and processed 23.8 trillion tokens in under four weeks. That was followed by four continuous weeks of reinforcement learning on 10,500 GB300 GPUs, producing more than 100 million rollouts. The company said this ranks among the largest reinforcement learning training runs in the public record, to its knowledge. Beam remains in final safety testing, with only a small number of users getting early access. Reflection said it plans to release the full weights, a technical report and a model card later this month under the Apache 2.0 license.

Reflection, an AI startup backed by Nvidia, has released Beam, its first open-weight model. Beam uses a mixture-of-experts architecture with 501 billion total parameters, while activating 23 billion parameters at a time. The model is designed for coding, reasoning and agent tasks.

Benchmark results put Beam near GLM-5.2

By Reflection’s own published results, Beam performs close to GLM-5.2 overall, though it still trails the latest leading Chinese open-source models.

On DeepSWE, Beam scored 44.4. GLM-5.2 scored 44.0, while Qwen 3.8-Max reached 51.0. GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash posted 61.0, 68.0 and 74.2, respectively.

Large-scale pretraining and reinforcement learning

Reflection also disclosed the scale of the training run. During pretraining, Beam used 6,144 Nvidia GB300 GPUs and processed 23.8 trillion tokens in less than four weeks. The company then ran reinforcement learning for another four straight weeks on 10,500 GB300 GPUs, generating more than 100 million rollouts.

Reflection said that, to its knowledge, this was one of the largest reinforcement learning training runs ever disclosed in the public record.

Full weights to be released later this month

Beam is still in its final safety testing phase, and only a small group of users can access it early. Reflection said it plans to publish the full weights, a technical report and a model card later this month under the Apache 2.0 license.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.