Reflection AI Beam: Benchmarks, Open Weights, Access
Reflection AI's Beam is a 501B open-weight model, 23B active, with Apache 2.0 weights due in October 2026. Its own benchmarks put it level with GLM-5.2.

Reflection AI’s Beam is a 501-billion-parameter open-weight model with 23 billion active, announced October 5, 2026. Reflection says the weights ship later in October under Apache 2.0. There is no public API and no price yet.
The verdict, on Reflection’s own numbers: Beam beats the other non-Chinese open models in its table, Thinking Machines’ Inkling and NVIDIA’s Nemotron 3 Ultra, on 13 of the 15 rows each is reported on. It is roughly level with Zhipu’s GLM-5.2, which it beats on 8 of 13 shared rows. It does not beat GLM-5.3 or Qwen 3.8 Max on a single row of its own table. If you want the best open model you can run today, Beam is not it. If you want an Apache-licensed model you can afford to self-host and fine-tune, it is the one to watch when the weights land.
| What you want to know | Answer as of October 8, 2026 |
|---|---|
| Price | None published. No public API |
| Who can use it now | A select group from an early-access waitlist at platform.reflection.ai |
| Weights | “Later this month” (October 2026), with the technical report and model card |
| Licence | Apache 2.0 |
| Size | 501B total parameters, 23B active per token (sparse mixture of experts, 52 layers) |
| Context | Up to 1M tokens after midtraining. Text in, text out (no image input) |
| Independent scores | None yet. Artificial Analysis had no Beam page on October 8, 2026 |
What is Reflection AI’s Beam?
Beam is Reflection AI’s first open-weight model, introduced in Introducing Beam: Reflection’s 501B open-weight model, the lab’s own announcement dated October 5, 2026. Reflection presents it as a step forward for “the Western open-weight frontier,” and calls it “the first model in a series.”
The architecture is a sparse mixture-of-experts transformer. Of its 501 billion parameters, 23 billion fire for any given token, routed through fine-grained experts with interleaved local and global attention across 52 layers. It reads and writes text only. Reflection’s own demos have it reason about game visuals and map grids “in text mode”, which is a workaround, not vision.
The training numbers are large for a first model, and Reflection published more of them than most labs do:
- Pretraining: 23.8 trillion tokens on 6,144 NVIDIA GB300 NVL72 GPUs, in under four weeks, at 92.3% goodput (the share of wall-clock time spent on training steps kept in the final model).
- Reinforcement learning: 10.5K GB300 GPUs for four weeks, more than 100 million rollouts, a pool of nearly one million coding, agentic and STEM environments, and roughly 1.3 billion sandboxes for training and grading.
- Context: rollouts ran up to 256K tokens. Midtraining extended the effective context to 1M.
- Effort control: a reasoning-effort parameter trades shorter answers for longer reasoning on hard tasks.
That RL run is the core of the pitch. Reflection calls it “one of the largest scale RL runs conducted by any open lab to date,” and gives Inkling’s 30 million rollouts as the comparison. Nobody can check that from outside yet, but the rollout count is a concrete number to hold the technical report to.
How does Beam score against GLM-5.2?
Close, and the closeness is the claim. Reflection says Beam “is competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks,” and that it matches GLM-5.2 on reasoning “while using 3–4× less inference compute.” Every score below comes from Reflection’s own benchmark table. None has been reproduced independently.
| Metric | Beam | GLM-5.2 |
|---|---|---|
| AIME 2026 | 97.8 | 99.2 |
| GPQA Diamond | 90.5 | 91.2 |
| Terminal Bench 2.1 | 80.1 | 81 |
| IFBench | 79.7 | 73.3 |
| AA-LCR | 79.3 | 78.3 |
| MCP Atlas | 78.7 | 77.8 |
| SWE-bench Pro v1 | 65.5 | 62.1 |
| LongBench v2 | 65.5 | 64 |
| DeepSWE v1.1 | 44.4 | 44 |
| HLE (no tools) | 36.2 | 40.5 |
| tau3 banking | 38 | 37.1 |
| AutomationBench | 37 | 26.2 |
| CritPt | 16.3 | 20.9 |
Read the chart by category and a pattern shows. The parity claim is softest exactly where Reflection makes it. On reasoning, GLM-5.2 leads every shared row: by 1.4 points on AIME 2026, 0.7 on GPQA Diamond, 4.3 on Humanity’s Last Exam and 4.6 on CritPt. “Comparable” is fair for AIME and GPQA. It is generous for HLE and CritPt. Beam’s real wins are agentic: 10.8 points on AutomationBench and 6.4 on IFBench.
Against the rest of the open field, the same table is less kind:
| Rival in Reflection’s table | Rows reported for both | Beam ahead | Rival ahead |
|---|---|---|---|
| Inkling (Thinking Machines) | 15 | 13 | 2 (IFBench, Omniscience) |
| Nemotron 3 Ultra (NVIDIA) | 15 | 13 | 1, plus a tie on AA-LCR |
| GLM-5.2 (Zhipu) | 13 | 8 | 5 |
| DeepSeek V4.1 Flash | 9 | 2 | 7 |
| Kimi K3 (Moonshot) | 14 | 1 | 13 |
| GLM-5.3 (Zhipu) | 12 | 0 | 12 |
| Qwen 3.8 Max (Alibaba) | 13 | 0 | 13 |
The coding rows show the gap most plainly. On Terminal Bench 2.1, Beam scores 80.1 against 88.2 for GLM-5.3, 88.3 for Kimi K3 and 90.6 for DeepSeek V4.1 Flash. On DeepSWE v1.1 it scores 44.4 against 74.2 for V4.1 Flash. Reflection doesn’t hide any of this. Its own post concedes that “frontier open models like Kimi K3 remain ahead on raw capability.”
What does “3–4× less inference compute” actually measure?
Reflection estimates compute with a stated formula: FLOPs ≈ 2 × active parameters × mean generated tokens per attempt. For a mixture-of-experts model that uses the 23 billion active parameters, not the 501 billion total. So the claim rests on two things: Beam has a small active count, and it reasons in fewer tokens than GLM-5.2 for a similar score.
Reflection names what the formula leaves out: “prompt prefill, context-dependent attention operations, and serving overhead.” It calls the result “an approximate compute comparison rather than measured inference cost.” That matters for agents in particular. A coding agent re-reads a long context on every turn, and prefill and attention are exactly the costs that grow with context. The 3–4× figure is a decode-side estimate. It isn’t a bill.
Some arithmetic on the active count is still useful (this site’s own, using Reflection’s formula, decode only):
- Beam: 2 × 23B = 46 GFLOPs per generated token.
- Mistral Large 4, at 52B active: 104 GFLOPs per token, about 2.3 times Beam’s.
- Inkling, at 41B active: 82 GFLOPs per token.
Per generated token, Beam needs less than half the decode compute of Mistral Large 4 and a little over half of Inkling’s. Whether it is cheapest per finished task depends on how many tokens it spends, and only an independent benchmark that publishes run costs can answer that.
When can you download or use Beam?
Not yet, for most people. Reflection says Beam “is undergoing final red-teaming and evaluations.” Right now it is open only to “a select group of users” from a waitlist. The weights, technical report, model card and “developer artifacts” are promised “later this month,” which means October 2026. No day is given.
The launch plan in the post has two parts:
- Weights under Apache 2.0, along with documentation and “the full stack for running, evaluating, and fine-tuning the model.”
- Distribution partners and open-source integrations, so Beam runs inside existing libraries and agent harnesses.
Apache 2.0 is the detail that matters most to a business user. It is the same licence Inkling shipped under. It allows commercial use, modification and redistribution, with no user-count thresholds and no field-of-use restrictions, so a legal team can approve it without reading a bespoke model licence.
What will Beam cost to run?
Nobody knows the API price yet, so the honest answer splits in two.
Hosted. Reflection has published no rate card. Once distribution partners list Beam, their per-token prices will set the market, the way host prices did for other open models. The AI model tracker will carry the first verified rate.
Self-hosted. Here arithmetic helps. These figures cover weights only, at a stated precision, before KV cache and activations. Reflection hasn’t said what precision the weights ship in.
| Model | Total params | Weights at BF16 (2 bytes/param) | At 4-bit (measured 0.54–0.63 GB per billion) |
|---|---|---|---|
| Reflection Beam | 501B | ~1.0 TB | ~271–316 GB |
| GLM-5.2 | 754B | 1.51 TB (published download) | 466 GB (published 4-bit build) |
| Mistral Large 4 | 1.05T | ~2.1 TB | ~567–662 GB |
The 4-bit range applies the per-billion sizes this site measured from real Q4_K_M downloads in how much RAM you need to run a local LLM. At 4-bit, Beam would fit a 512GB unified-memory workstation with room left for context. Mistral Large 4 would not. That is the practical side of a 501B model: a big model that a single well-equipped box can still load, once quantized builds exist.
Beam vs Mistral Large 4: who is winning the Western open-weight race?
Two non-Chinese labs launched large open models in the same week, and they made different bets.
Mistral Large 4 is the bigger and broader one: 1.05 trillion parameters, 52 billion active, image input, a public preview API since October 6 and weights due by the end of October. Mistral says it significantly outperforms “any open-weight model developed in the US or Europe,” and its launch post scores Large 4 at 61.7% on DeepSWE v1.1. Reflection’s own table puts Beam at 44.4 on the same benchmark. On that row, Mistral’s claim holds.
Beam is the smaller and more permissive bet. It has fewer than half of Large 4’s active parameters, text only, and an Apache 2.0 licence that Reflection has stated outright. Mistral had not named a licence for Large 4 when this was written.
Both launches are framed against China, and the scoreboard explains why. On Reflection’s own table, Kimi K3, GLM-5.3 and Qwen 3.8 Max lead Beam on nearly every row. The October releases close part of the gap. They don’t close it. What the Western labs offer instead is provenance and licence terms that some buyers need more than the last ten points of Terminal Bench. The wider field is ranked in the best open-weight AI models of 2026, and every October launch is logged in new AI models released in October 2026.
Which should you use?
- You need the strongest open coding model today: not Beam. On Reflection’s own numbers, DeepSeek V4.1 Flash, GLM-5.3 and Kimi K3 all lead it on Terminal Bench and DeepSWE.
- You need a non-Chinese open model under a permissive licence: wait for Beam’s weights. Unless the independent scores contradict the launch table, it replaces Inkling as the default pick, since it leads Inkling on 13 of 15 shared rows under the same Apache 2.0 terms. Inkling stays relevant for multimodal input, which Beam doesn’t have.
- You want to self-host on one machine: Beam’s size is the draw. At 4-bit, the arithmetic above puts it inside a 512GB box.
- You need vision or a callable API this week: Mistral Large 4 has both, in preview.
To get on the list now, sign up for early access at platform.reflection.ai. Then re-check when the weights and technical report land, because the first independent scores will confirm or cut the parity claim.
Frequently asked questions
- What is Reflection AI Beam?
- Beam is the first open-weight model from Reflection AI. Announced October 5, 2026, it is a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active per token, built for coding, reasoning and agentic work. It is text-only, supports up to 1M tokens of context after midtraining, and has a reasoning-effort setting.
- Is Reflection Beam open source?
- Reflection says it will release the weights under Apache 2.0 later in October 2026, along with the technical report, model card, documentation and the stack for running, evaluating and fine-tuning the model. Apache 2.0 permits commercial use, modification and redistribution. As of October 8, 2026 the weights had not been released.
- How much does Reflection Beam cost?
- There is no published price. Beam has no public API, and access is through an early-access waitlist for a select group of users. Hosted prices will come from distribution partners after the weights ship. Self-hosting needs about 1.0 TB for the weights at BF16, or roughly 271 to 316 GB at 4-bit using measured quantized sizes.
- Is Beam better than GLM-5.2?
- They are roughly level on Reflection's own benchmarks. Beam leads on 8 of the 13 rows where both are reported, mainly agentic and long-context tests such as AutomationBench (37.0 vs 26.2) and IFBench (79.7 vs 73.3). GLM-5.2 leads on all four shared reasoning rows, including AIME 2026 (99.2 vs 97.8) and Humanity's Last Exam (40.5 vs 36.2), plus Terminal Bench 2.1. All scores are vendor-reported.
- Is Beam better than Kimi K3 or GLM-5.3?
- No. On Reflection's own table Beam trails GLM-5.3 on all 12 shared rows and trails Kimi K3 on 13 of 14, including Terminal Bench 2.1 (80.1 against 88.2 and 88.3). Reflection itself says Kimi K3 remains ahead on raw capability, and pitches Beam on inference efficiency instead.
- When will Reflection release Beam weights?
- Reflection says later in October 2026, without a specific date. The weights will ship with the technical report, model card and developer tooling, and the model will launch with distribution partners and open-source library integrations.
Sources
Reflection AI (2026). Introducing Beam: Reflection’s 501B open-weight model. Reflection AI, lab announcement with vendor-reported benchmark table. https://reflection.ai/blog/introducing-beam Verified 2026-10-08.
Mistral AI (2026). Introducing Mistral Large 4. Mistral AI, lab announcement with vendor-reported benchmarks. https://mistral.ai/news/mistral-large-4/ Verified 2026-10-08.
Artificial Analysis (2026). Models. Artificial Analysis, independent benchmarking service. https://artificialanalysis.ai/models Checked 2026-10-08: no Beam page listed.
Thinking Machines Lab (2026). Inkling model details, as recorded in this site’s Inkling launch post (41B active parameters, Apache 2.0, 30 million-plus RL rollouts). Verified 2026-07-15.