Quick Facts
- Beam is a sparse Mixture-of-Experts model with 501B total parameters and 23B active per token, released as open source.
- The reinforcement learning phase used 10,500 NVIDIA GB300 GPUs over four weeks, generating more than 100 million rollouts.
- Beam scores 80.9 on SWE-Bench Verified and 97.8 on AIME 2026, while using one-third to one-fourth the hardware of comparable open models.
Reflection AI on Tuesday released Beam, a 501-billion-parameter open-source model designed for coding, reasoning, and agentic workloads. The New York-based company, founded in 2024 by former Google DeepMind researchers Misha Laskin and Ioannis Antonoglou, built Beam end-to-end from scratch using a three-phase training process.
Beam is a sparse Mixture-of-Experts model. It holds 501 billion total parameters but activates only 23 billion per token, reducing the active compute footprint while maintaining a large model capacity. Midtraining extended the model's effective context length to one million tokens.
Training at Scale
Phase one involved pretraining Beam Base on 23.8 trillion tokens using a cluster of 6,144 NVIDIA GB300 NVL72 GPUs. That base model was ready in under four weeks. Reflection applied custom per-language filters to remove low-quality code from the training dataset.
The reinforcement learning phase was the most hardware-intensive. Reflection ran over 100 million rollouts across 10,500 NVIDIA GB300 GPUs over four weeks, using a maximum context length of 256,000 tokens and nearly one million environments. The company trained at 92.3% goodput. Reflection says this is one of the largest-scale RL runs conducted by any open lab to date.
Reflection also trained Beam with a controllable length penalty. Users can set a reasoning effort parameter: lower values favor shorter responses, higher values allow longer reasoning chains for harder tasks.
Benchmark Results
On benchmarks Reflection published, Beam scores 80.9 on SWE-Bench Verified, 80.1 on Terminal-Bench v2.1, and 97.8 on AIME 2026. Moonshot AI's Kimi K3 leads on raw capability, scoring 88.3 on Terminal-Bench v2.1 versus Beam's 80.1. Reflection acknowledges that gap.
Where Reflection focuses its efficiency argument is hardware cost. The company says Beam matches or beats GLM-5.2, a model with roughly 250 billion more parameters, on certain tasks while using one-third to one-fourth the hardware. Reflection also says Beam approaches the performance of Qwen 3.8-Max, a model with over two trillion parameters.
Artificial Analysis, which received early access from Reflection, said Beam shows early signs of being one of the most token-efficient open models it has evaluated. No independent benchmark results have been published yet.
The Business Case
Reflection sells end-to-end AI systems to governments, large enterprises, and public sector institutions that want full control over their AI stack. CEO Laskin framed the open-source release around sovereignty. "The only way to own intelligence is, by definition, if it's open," he said.
Laskin has been direct about the competitive threat from Chinese open-source models. "DeepSeek and Qwen and all these models are our wake-up call because if we don't do anything about it, then effectively, the global standard of intelligence will be built by someone else," he said in October 2025. He also argued that enterprises using Chinese models inherit legal liability tied to training data gathered under less restrictive copyright laws.
Reflection has raised $4.66 billion in total funding. A $2.5 billion round in March 2026 valued the company at $27.5 billion. In June 2026, the company signed a compute agreement with SpaceX worth up to $6.3 billion. The company had roughly 320 employees as of this year.
Read more: Reflection AI debuts open-source Beam model with 501B parameters
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
