Reflection unveils Beam, a 501B open weight model it says rivals GLM 5.2 with 3 to 4 times less compute
Reflection AI introduced Beam, a 501 billion parameter open weight model with 23 billion active parameters. It claims GLM 5.2 level reasoning at 3 to 4 times less compute, with Apache 2.0 weights due in October.

The race for the best open weight AI model has a new American contender. On Monday, 5 October 2026, the Brooklyn based startup Reflection AI introduced Beam, its first open weight model. Beam is a sparse mixture of experts model with 501 billion parameters in total, of which 23 billion are active for each token. Reflection built it for coding, reasoning and agent work, and it claims that Beam is competitive with larger Chinese open models such as GLM 5.2 while using a fraction of the compute when it answers.
The weights are not public yet. Reflection says Beam is going through final red teaming and evaluations, and that the weights, a technical report, a model card and developer tools will follow later this month under the Apache 2.0 license. Until then, a select group of users gets early access through a waitlist. That makes this a preview, but a very detailed one: the launch post is one of the most detailed descriptions of a large training run we have seen from a Western lab.

What Reflection actually announced
Here are the facts that are on the record in Reflection's blog post and in the reporting by TechCrunch.
- Size and design. Beam is a text only mixture of experts model with 501 billion total and 23 billion active parameters. It has 52 layers and combines local and global attention. Reflection says its midtraining stage extends the effective context length to 1 million tokens.
- Pretraining. The base model saw 23.8 trillion tokens from the web, public sources and proprietary licensed datasets. Pretraining ran end to end in under four weeks on a cluster of 6,144 Nvidia GB300 NVL72 GPUs.
- Reinforcement learning. After pretraining, Reflection ran what it calls a high compute RL campaign: more than 100 million rollouts on 10,500 Nvidia GB300 GPUs over four weeks, with rollouts of up to 256,000 tokens.
- Focus. The model targets coding, agentic tasks and reasoning. Reflection calls it a "workhorse model" for enterprises, the public sector and developers.
- Release. Weights, documentation and the full stack for running, evaluating and fine tuning Beam are promised for October under Apache 2.0, one of the most permissive open source licenses. Reflection says it will launch with distribution partners, including hyperscalers and neoclouds according to TechCrunch, and with integrations into common open source libraries and harnesses.
- Series. Beam is described as the first model in a series. Reflection says it is already training the next one.
Who is Reflection?
Reflection is not a household name, but it is one of the best funded AI startups in the United States. According to TechCrunch, it was founded in 2024 by two former Google DeepMind researchers and has raised roughly 4.7 billion dollars, citing PitchBook data. Backers include Nvidia, Sequoia Capital and Lightspeed Venture Partners, and its last round valued the company at 25 billion dollars before the new money.
The company has also been buying compute on a large scale. This summer, Reflection signed deals worth more than 7 billion dollars in total with SpaceX and Nebius for access to Nvidia GB300 chips through 2029. That explains how a two year old startup can run a 10,500 GPU reinforcement learning job for a month.
Reflection's business pitch is something it calls "AI factories": institutions such as banks, trading firms or governments would train Reflection's models on their own data and run them locally. TechCrunch reports that Reflection is already testing a sovereign AI factory partnership with Shinsegae Group in South Korea. Open weights are the core of that idea. A customer that wants full control over its model cannot rent a closed API.
Smaller than the Chinese giants
The most interesting design choice is the size. Beam is big, but it is clearly smaller than the open models it wants to compete with.

According to TechCrunch, Z.ai's GLM 5.2 has roughly 744 billion parameters with 40 billion active. Mistral Large 4, which Mistral presented as a preview on 6 October and which we covered in our article on Le Chonk, has 1 trillion parameters with 49 billion active. Beam activates only 23 billion per token.
Active parameters matter because they decide how much work the hardware does for every token the model generates. Fewer active parameters mean lower cost per token and faster answers, as long as quality holds up. Reflection estimates the compute of each model as two times the active parameters times the number of generated tokens, and on that basis it claims that Beam reaches scores comparable to GLM 5.2 on advanced reasoning benchmarks while using 3 to 4 times less inference compute. The company admits that this is an approximation that ignores prompt processing, attention costs and serving overhead.
The benchmarks, read honestly
Reflection published a long benchmark table with eight models: Beam, Inkling from Mira Murati's Thinking Machines Lab, Nvidia's Nemotron 3 Ultra, GLM 5.2 and GLM 5.3 from Z.ai, Kimi K3 from Moonshot AI, Qwen 3.8 Max from Alibaba and DeepSeek V4.1 Flash. Many cells are marked "not reported", and the scores for the other models come from Artificial Analysis and DataCurve, not from the same harness. All numbers are vendor claims and have not been independently verified.

The table shows a clear pattern. Beam is far ahead of the other Western open models. On Terminal Bench 2.1 it scores 80.1 percent, compared with 63.8 for Inkling and 56.4 for Nemotron 3 Ultra. On SWE Bench Verified it reaches 80.9 percent against 77.6 for Inkling and 70.7 for Nemotron. On SWE Bench Pro v1 it posts 65.5 percent, ahead of GLM 5.2 at 62.1 and just behind Qwen 3.8 Max at 67.7.
But Beam does not lead the open field as a whole. On Terminal Bench 2.1, GLM 5.3, Kimi K3, Qwen 3.8 Max and DeepSeek V4.1 Flash all score between 86.6 and 90.6 percent. On DeepSWE v1.1, Beam's 44.4 percent is roughly level with GLM 5.2 at 44.0, but well behind GLM 5.3 at 61.0, Kimi K3 at 68.0 and DeepSeek V4.1 Flash at 74.2. Reflection says so itself: frontier open models like Kimi K3 "remain ahead on raw capability", and Beam's advantage is efficiency.

The reasoning results look similar. On Humanity's Last Exam without tools, Beam scores 36.2 percent, ahead of Inkling and Nemotron, but behind all of the Chinese models in the table. On GPQA Diamond it reaches 90.5 percent, where the range for the whole table is 87.0 to 93.5. On AIME 2026 it scores 97.8 percent. For agents, Beam reaches 78.7 percent on MCP Atlas, 77.4 percent on BrowseComp with context management and 38.0 percent on the tau3 banking test, which is ahead of GLM 5.2 and Kimi K3 there, but behind Qwen 3.8 Max at 55.2.
So the honest summary is this: Beam is the strongest Western open weight model in Reflection's own comparison, roughly at the level of GLM 5.2 on many tasks, and behind the newest Chinese releases on raw scores. Whether its efficiency makes up for that gap will only become clear once independent testers can run the weights.
A very large reinforcement learning run
Most of Reflection's post is about how Beam was trained, and this is where it gets unusually specific.

Reflection built a pool of nearly one million training environments for software engineering, terminal use, competitive coding, science, web search, tool use and general knowledge work. Most of them came from synthetic data pipelines, with additional vendor data and open source material. Tasks were filtered so that they were neither always solved nor impossible for the model, and broken or "hackable" tasks were removed. Training and grading used about 1.3 billion sandboxes, and the platform supported up to 170,000 sandboxes at the same time.
For comparison, Reflection says Inkling was trained on 30 million rollouts and MiMo on 753,000. Beam's reasoning expert alone used 80 million of the more than 100 million rollouts. Reflection calls it "one of the largest scale RL runs conducted by any open lab to date" and says capabilities kept improving with more RL compute, "with no sign of a plateau."
A few engineering details stand out:
- Asynchronous training. Rollouts were generated while the trainer kept learning. Every token is tagged with the model version that produced it, and Reflection says learning stayed stable even with samples that were more than a day old, 107 weight versions behind the current policy.
- Fast updates. New weights reached the inference fleet in a median of about 12 seconds.
- Resilience. During the run, 71 inference incidents were handled without stopping the training job. Capacity came back in a median of eight minutes.
- Efficient reasoning. A controllable length penalty rewarded correct answers while discouraging unnecessary tokens. Users get a reasoning effort setting to trade speed against quality.
Reflection also reports an interesting side effect. During a phase that trained on reasoning, software engineering and terminal tasks, Beam got better at web browsing even though no browsing tasks were in the mix. With web access, it learned on its own to query other language models and to use OCR services to read documents.
Data and safety
On the pretraining side, Reflection says it removes about 95 percent of raw internet tokens through parsing, deduplication and quality filtering. At the same time, it claims its own classifiers kept roughly 1.8 trillion high quality tokens that conventional filters would have thrown away, including 87 percent of its curated web code tokens. The pretraining run executed nine semi automatic rewinds after gradient spikes or suspected silent data corruption, and its goodput, the share of time spent on training steps that ended up in the final model, reached 92.3 percent towards the end.
For safety and alignment, Reflection trained a second model from the same pretrained checkpoint with its own pipeline focused on behavior principles, and merged both teachers into Beam through on policy distillation. The principles come in three tiers: hard rules such as following the safety policy and keeping its identity as an AI agent, qualities such as making accurate claims and admitting uncertainty, and a default style that is direct and efficient. Reflection says it will publish its safety evaluation results in the technical report and open source the safety tests it built internally.
Why it matters
For the last two years, the strongest open weight models have mostly come from China: DeepSeek, Qwen, Kimi and GLM. American labs such as OpenAI, Anthropic and Google kept their best models closed. That left companies and governments that want to run a top model on their own hardware with an uncomfortable choice, especially in regulated industries where the origin of a model matters.
Beam is a serious attempt to change that from the United States, and it arrives on the same day as Mistral Large 4 from Europe. If Reflection delivers the Apache 2.0 weights this month, developers get a model that is smaller and cheaper to serve than most of its rivals, with a license that allows almost any use. For enterprises, the efficiency claim may matter more than the last few benchmark points. Agents run for hours and generate huge numbers of tokens, so a model that needs a third of the compute per answer can change the economics of a whole product.
It also matters for Nvidia. Reflection is backed by Nvidia and trained everything on GB300 chips, and the "AI factory" idea of every institution running its own model is exactly the future that sells more GPUs.
Dany's take
I like how much Reflection put on the table. Most launch posts give you three charts and a waitlist. This one explains the data pipeline, the RL infrastructure, the failures during training and even where the model is weaker than the competition. That is the right attitude for a company that wants to build trust in open models.
But I would not get carried away yet. The benchmarks are self reported, many cells are empty, and Beam is behind the newest Chinese models on several of the hardest coding and reasoning tests. The 3 to 4 times efficiency claim is based on an estimate, not on measured serving costs. And until the weights can actually be downloaded, Beam is a promise.
What I will watch: whether the weights really ship under Apache 2.0 in October, what independent testers find on coding agents, and what it costs per million tokens at the first hosting partners. If those three line up, Beam could become the default open model for Western companies that do not want to depend on a closed API or on a model from China.

Sources
- Reflection: Introducing Beam (5 October 2026)
- TechCrunch: Reflection debuts Beam, an open weight AI model to rival Chinese models at lower compute cost (5 October 2026)
- Mistral: Mistral Large 4 (6 October 2026, for the comparison of model sizes)
Source: reflection.ai