TL;DR
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Reflection has introduced Beam, a 501-billion-parameter sparse Mixture-of-Experts model with 23 billion active parameters, aimed at coding, reasoning and agentic tasks. The company says Beam is undergoing final red-teaming and evaluations, with weights and technical materials planned for release later this month; independent verification of its reported benchmark results is not yet available in the supplied material.
Reflection has introduced Beam, a sparse Mixture-of-Experts model with 501 billion total parameters and 23 billion active, designed for coding, reasoning and agentic workloads. The company says it is completing red-teaming and evaluations and plans to release the weights, technical report, model card and developer artifacts later this month, making Beam’s claims open to outside inspection.
Reflection describes Beam as its first open-weight model. The company says pretraining used 23.8 trillion tokens drawn from curated web material and proprietary licensed datasets. It says the model matched or outperformed available open base models of similar size, though the announcement does not provide details here that allow readers to independently assess that comparison.
For reinforcement learning, Reflection reports generating more than 100 million rollouts over four weeks on 10,500 NVIDIA GB300 GPUs. It says training and grading used about 1.3 billion sandboxes and that it sourced one million coding, agentic and STEM environments. Reflection calls the effort one of the largest reinforcement-learning runs by an open lab to date; that characterization is the company’s claim.
Reflection says Beam is competitive with larger open models such as GLM 5.2 on coding and agentic tasks and approaches Qwen 3.8-Max on those tasks. It also acknowledges that Kimi K3 remains ahead on raw capability. For advanced reasoning, the company reports scores comparable to GLM 5.2 while using three to four times less inference compute. Its compute estimates use active parameters and generated tokens, and exclude prompt prefill, context-dependent attention operations and serving overhead.
Lower Compute for Coding Tasks
Beam’s reported results place inference efficiency at the center of its pitch. If the company’s comparisons hold up, a model that performs strongly on coding and agentic work while using fewer active parameters or less generation compute could lower the cost of deploying such systems. That may matter to enterprises running coding assistants or agents repeatedly, where per-task compute can affect operating costs.
The practical value will depend on more than headline parameter counts. Model access, hardware requirements, response quality on real workloads and serving costs will shape whether Beam is useful in production. The planned release of weights and documentation should give developers a way to examine those factors, but the announcement alone does not establish them.
As an affiliate, we earn on qualifying purchases.
Training at Reinforcement Learning Scale
Beam combines large-scale pretraining with reinforcement learning, a method that uses feedback from tasks or environments to refine model behavior. Reflection says it used asynchronous policy-gradient training, where multiple training processes can generate and learn from rollouts at different times. The company reports developing methods to keep learning stable despite stale samples and numerical differences between training and inference systems.
Reflection says some rollouts were more than a day old, corresponding to 107 model weight versions behind the current policy, while training remained stable. It also says Beam was trained with a length penalty intended to reward successful solutions while discouraging unnecessary tokens. These are descriptions of the company’s training approach; the supplied report excerpt does not include the full technical documentation needed to evaluate them.
“Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.”
— Reflection
Benchmarks Await Independent Review
The supplied announcement reports benchmark comparisons, but it does not provide independent confirmation of Beam’s performance. Reflection says final red-teaming and evaluations are still underway. The planned technical report and model card may add evaluation methods and safety information; until they are available, it is unclear how the results will hold up under outside testing or across different deployments.
Reflection’s inference-compute comparisons are estimates rather than measured serving costs. They exclude prompt prefill, some attention costs and serving overhead, and rely on generated-token counts and active parameters. Actual costs and latency may vary with workloads, hardware and implementation. The announcement also does not specify the final license terms, access conditions or exact release date.
Release Materials Later This Month
Reflection says it plans to publish Beam’s weights, technical report, model card and developer artifacts later this month, after its current red-teaming and evaluation work. The company is accepting sign-ups for early access. The announcement does not give a precise publication date, and it remains unclear whether all materials will be available at once or under what license.
Once the materials are released, developers and researchers will be able to examine the model and test the company’s benchmark and efficiency claims. Those results, along with details on access and deployment, will help clarify Beam’s practical standing among open-weight models.
Key Questions
What is Beam?
Beam is Reflection’s first open-weight model, a sparse Mixture-of-Experts system intended for coding, reasoning and agentic workloads. Reflection reports 501 billion total parameters and 23 billion active parameters.
When will Beam’s weights be released?
Reflection says it plans to release the weights and supporting materials later this month. It has not announced an exact date in the supplied report.
How did Reflection train Beam?
Reflection says it pretrained Beam on 23.8 trillion tokens and generated more than 100 million reinforcement-learning rollouts over four weeks using 10,500 NVIDIA GB300 GPUs.
How does Beam compare with other models?
Reflection says Beam is competitive with GLM 5.2 and approaches Qwen 3.8-Max on coding and agentic tasks, while Kimi K3 remains ahead on raw capability. These are the company’s comparisons and await outside review.
Are Beam’s efficiency claims independently verified?
The supplied announcement does not include independent verification. Reflection describes its compute figures as estimates that exclude prompt prefill, some attention costs and serving overhead.
Source: hn
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
