NEWS
Reflection Bets $25 Billion on Beam’s Cheaper Open Weights
Reflection AI unveiled Beam, a 501 billion parameter open-weight model it says rivals GLM-5.2 at lower estimated compute, as a $25 billion factory bet.
Reflection AI on Monday unveiled Beam, a 501 billion parameter open-weight model it says matches Z.ai’s GLM-5.2 on hard reasoning tests at a fraction of the inference compute. The files are still in final red-teaming, with weights due later in October.
Laskin and Ioannis Antonoglou, both former Google DeepMind researchers, founded the Brooklyn company in 2024. The pitch sits on a $25 billion pre-money valuation and more than $7 billion in Nvidia GB300 access through 2029.
A $25 Billion Lab Ships Its First Proof
Beam is the first public frontier model from a lab that has been raising and renting chips faster than it has been publishing weights. PitchBook puts total capital raised at roughly $4.7 billion, with Nvidia, Sequoia Capital, and Lightspeed Venture Partners among the backers. Laskin is chief executive. Antonoglou, the chief technology officer, worked on AlphaGo, AlphaZero, and MuZero, then on Gemini reinforcement learning from human feedback.
That stack of money and hardware is the wager. Beam is the first artifact investors can point to. The company still will not let the public download the weights, and every benchmark in the launch post is its own.
THE CAPITAL AND COMPUTE STACK
- 2024: Laskin and Antonoglou found Reflection in Brooklyn after leaving Google DeepMind.
- March 2025: Raises $130 million at a $545 million valuation, combining a $25 million seed and a $105 million Series A.
- October 2025: Raises $2 billion at an $8 billion valuation, with Nvidia putting in about $800 million.
- March 16, 2026: Signs an MOU with Shinsegae Group in San Francisco to plan a 250 megawatt sovereign AI factory in South Korea.
- April 2026: Closes a round at a $25 billion pre-money valuation.
- July 1, 2026: SpaceX monthly bill of $150 million starts for GB300 access at Colossus 2, a term that reaches about $6.3 billion if it runs through 2029.
- July 2026: Adds more than $1 billion of GB300 capacity from Nebius through 2029, taking combined compute deals above $7 billion.
- October 5, 2026: Posts Beam as a 501 billion parameter open-weight preview, with full weights due later in October.
The SpaceX line is large and soft at the same time. Either side can end the contract on 90 days’ notice after the first three months, so the first $450 million is the only slice that is hard to walk away from. A $25 billion lab that has not yet published a downloadable frontier checkpoint is still paying rent on that clock.
Introducing Beam: a highly efficient agentic open model with 501B total parameters and 23B active.
– Frontier reasoning efficiency
– Advances the Western open frontier on coding & agentic tasks
– Trained end-to-end from scratchFull weights release this month.
Learn more about… pic.twitter.com/1rMABCywUG
— Reflection (@reflection_ai) October 5, 2026
Beam Trails Kimi and DeepSeek on the Scoreboard
Reflection describes Beam as a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active, trained for coding, reasoning, and agent work. It is text-only. Midtraining extends the effective context window to 1 million tokens. Pretraining used 23.8 trillion tokens from the web, public sources, and licensed sets, with a heavy tilt toward code.
Z.ai’s GLM-5.2, the model Reflection keeps naming as the peer, has about 744 billion total parameters and 40 billion active. Qwen 3.8 Max sits in a larger 2 trillion-plus class. On Reflection’s own table, Beam lands next to GLM-5.2 on a few coding tests and then falls behind Kimi K3 and DeepSeek V4.1 Flash. The company says Kimi K3 remains ahead on raw capability, and that Beam’s edge is inference-time efficiency.
BEAM AGAINST OPEN RIVALS ON CODING TESTS
| Test | Beam | GLM-5.2 | Kimi K3 | DeepSeek V4.1 Flash | Inkling |
|---|---|---|---|---|---|
| DeepSWE v1.1 | 44.4 | 44.0 | 68.0 | 74.2 | NR |
| Terminal Bench v2.1 | 80.1 | 81.0 | 88.3 | 90.6 | 63.8 |
| SWE Bench Pro v1 | 65.5 | 62.1 | NR | NR | 54.3 |
| SWE-bench Verified | 80.9 | NR | NR | NR | 77.6 |
| SWE Bench Pro v2-Hard | 77.2 | NR | 88.2 | NR | 56.9 |
NR means the rival did not publish a score for that row. DeepSWE is the clearest gap: Beam’s 44.4 sits 0.4 above GLM-5.2 and 29.8 points behind DeepSeek V4.1 Flash. Terminal Bench v2.1 is a 0.9 point dip versus GLM-5.2 and a 10.5 point gap versus DeepSeek. On SWE Atlas Codebase QnA, a row not shown above, Beam scored 34.6 against 68.0 for Kimi K3.
The clean Western comparison is Inkling, the open model from Mira Murati’s Thinking Machines Lab, released in July. Beam outscores Inkling on the four coding tests where both published numbers, including 80.9 versus 77.6 on SWE-bench Verified. Inkling is multimodal. Beam is not, so that head-to-head is a partial one. None of these figures has been independently verified.
Where frontier open models like Kimi K3 remain ahead on raw capability, Beam’s advantage is efficiency at inference time.
Reflection AI, Introducing Beam
Alex Polozov, a member of technical staff who joined in November 2025, wrote that pretraining began in earnest in 2026 with a small group and that “on many agentic coding capabilities the model kept going up without any sign of a plateau.” That is the internal read. Outside testers cannot check it until the checkpoint is public.
Reflection’s Efficiency Number Omits Prefill and Serving
The launch’s commercial sentence is the compute ratio, not the top score. Reflection says Beam hits scores comparable to GLM-5.2 on advanced reasoning tests while using 3 to 4 times less inference compute, and more than 4 times less than leading Western open models. It estimated generation compute as FLOPs of about 2 times the active parameter count times mean generated tokens per attempt, counting each multiply-add as two operations, and using active parameters for mixture-of-experts models.
The company is explicit about what that sum leaves out. The estimate skips prompt prefill, context-dependent attention, and serving overhead, and it calls the result an approximate compute comparison rather than measured inference cost. Rival eval numbers in the efficiency charts were taken from Artificial Analysis and DataCurve. If you run agents around the clock, the bill still depends on your prompts, your batching, and your hardware.
That caveat is load-bearing for the factory pitch. A sovereign buyer or a trading firm that wants a local model cares about tokens per dollar on its own cluster, not a FLOPs proxy that ignores attention. The 23 billion active slice is the real design bet: most of the 501 billion parameters stay dark on each token, which is how Reflection wants Beam to undercut both closed APIs and heavier open rivals.
Users can also set a reasoning-effort control. Early in reinforcement learning, a length penalty taught the model to solve tasks with fewer tokens; later, completion lengths grew again as agent skills improved. Shorter traces cut the same FLOPs formula. Longer traces buy harder tasks. The knob is how Reflection wants operators to spend compute on purpose rather than by default.
What the Reinforcement Run Used
Reflection ran reinforcement learning for four weeks on 10,500 Nvidia GB300 GPUs and generated more than 100 million rollouts from about 1 million coding, agent, and STEM environments. Training and grading used about 1.3 billion sandboxes, and scores were still rising when the run stopped. The company calls it one of the largest RL runs any open lab has documented.
THE RL AND PRETRAIN RUN
- Pretrain cluster: Beam Base trained end to end in under four weeks on 6,144 Nvidia GB300 NVL72 GPUs, with 92.3 percent goodput late in the run and nine semi-automatic rewinds.
- RL fleet: 10,500 GB300 GPUs for four weeks, more than 100 million rollouts, a 256K maximum context during that stage, and an average of 110,000 concurrent rollouts.
- Sandboxes: About 1.3 billion used for training and grading, with up to 170,000 concurrent, across more than 20 clusters, two clouds, and four regions.
- Stability: 71 inference incidents were handled without killing the job, with a median recovery of eight minutes and lost capacity of 0.02 percent of serving GPU-minutes.
The architecture work sits under that spend. Beam uses interleaved local and global attention and fine-grained routed experts, with near-uniform expert use at the end of pretraining (the busiest expert at 1.04 times average load). The recipe builds on DeepSeek’s auxiliary-loss-free load balancing method, then adds cosine decay of expert-bias updates and sequence-level balancing so routing stays even when the data mix shifts into RL.
Data work was as aggressive as the GPU count. About 95 percent of raw internet tokens were dropped through parsing, deduplication, and filters, while the pipeline kept about 1.8 trillion high-quality tokens that conventional web filters would have missed, including 87 percent of curated web-code tokens. Custom per-language code filters stripped autogenerated junk. A PDF pipeline ran a vision-language OCR model across petabytes of technical documents so the base model had STEM coverage before RL started.
Some of what the RL stage taught leaked past the task mix. Browsing scores rose during a phase that included no browsing tasks. When the model was given web access, it started searching for other large language models, querying them, and calling OCR APIs to read documents. That is useful in an agent. It is also a behavior to log before anyone points Beam at a live network.
One of the launch demos is a concrete specimen of that habit. Plugged into OpenCode and pointed at Unsloth, Beam built a fine-tuning notebook for the latest small Gemma-4 model on a Text2SQL task, an out-of-distribution job, and raised Gemma’s accuracy on a held-out test set by 66.5 percent. In another test, it filled a 180 by 90 land-and-water grid at 95.5 percent coverage. Those are Reflection’s demos, not third-party trials.
Nvidia, SpaceX, and the Factory Pitch
The product behind the checkpoint is not a chatbot. Reflection is selling “AI factories,” local systems that institutions would train on their own data, with Beam and later models as the open core. Jensen Huang, whose company both backs Reflection and sells it the GPUs, has pushed that factory idea for years. Nvidia wins if the open stack spreads, because the plants still need GB300-class silicon.
Hedge funds and trading firms have been named as early private buyers for that kind of system. The closed labs, OpenAI and Anthropic, still lead on raw frontier quality. Chinese open labs still lead on several of the agent tests Reflection printed. The gap Reflection is trying to occupy is a Western checkpoint that a bank, a ministry, or a retailer can run, inspect, and fine-tune without sending prompts to a US API or depending on a Chinese weight file.
That is why the compute leases matter more than the blog’s leaderboard. SpaceX is renting Colossus 2 capacity in Memphis. Nebius is selling more than $1 billion of the same GB300 generation through 2029. Distribution at launch is supposed to run through hyperscalers and neoclouds, with hooks into open-source libraries, so a factory team can pull Beam into tools it already runs. The model is the sample. The plant is the contract.
Closed-model bills keep rising, and several governments now treat foreign hosted models as a control problem. An open Western workhorse is the political product those buyers can sign. It is also a product that still has to beat Kimi and DeepSeek on the jobs those buyers actually run, once someone other than Reflection holds the files.
South Korea Is the First Sovereign Test
The only named factory customer is still on paper. On March 16, Shinsegae Group and Reflection signed an MOU at the National AI Center in San Francisco to plan a 250 megawatt sovereign AI factory in South Korea, billed as the first project under the US AI Export Program. Commerce Secretary Howard Lutnick attended. Shinsegae would secure land and construction. Reflection would handle design and operations. The build is pegged at least 10 trillion won, about $6.8 billion, with a joint venture due in 2026 and Nvidia supplying the GPUs.
Together with Shinsegae, we will create AI infrastructure that Korea can autonomously evolve.
Misha Laskin, CEO, Reflection AI, at the San Francisco signing
Shinsegae’s own use case is retail: agents that pick products, pay, and ship, plus inventory and logistics tools across E-Mart. That is a narrow, high-volume job, closer to Beam’s coding-and-agent training mix than to a general chatbot bake-off. If the joint venture is built, Korea becomes the first place the factory slide has to survive contact with land, power, and a local partner’s data. An MOU is not a live cluster, and the 250 megawatt plant is still a plan.
Weights Arrive Later in October Under Apache 2.0
Monday’s post is a preview. Reflection is still finishing red-teaming and evaluations. A separate safety and alignment model was trained from the same pretrained checkpoint, then merged through multi-teacher on-policy distillation. Safety numbers are promised in the technical report, and the company says it will open-source the internal safety tests. Until then, there is no public refusal or jailbreak card to compare with GLM or Llama releases.
Early use is gated. Teams can join an early access waitlist for Beam while the checkpoint stays behind the wall. The October drop is the moment the bet becomes testable.
WHAT REFLECTION SAYS SHIPS IN OCTOBER
- The weights: Full parameters under Apache 2.0 license terms, which allow commercial use, modification, and redistribution.
- The paperwork: A technical report, model card, and developer artifacts for running, evaluating, and fine-tuning.
- The numerics: Quantized FP8 and NVFP4 builds aimed at cheaper deployment on Nvidia hardware.
- The pipes: Distribution through hyperscalers and neoclouds, plus integrations with open-source libraries and harnesses.
Apache 2.0 is the legal piece of the factory sale. A ministry that wants to fine-tune on classified data, or a fund that wants to keep prompts off someone else’s GPU, needs a license that does not yank the right to run the model in-house. The quantized builds are how Reflection wants that in-house copy to fit on a smaller fleet than a 40 billion active rival.
The lab is already training the next model, which Laskin has described as much more capable than Beam. That is the tell. Beam is a workhorse sample meant to get Western open weights into the conversation, then into private clusters, while the $150 million monthly meter at Colossus 2 keeps running. Independent tests, and any honest price per task, start only when the files are actually posted.
Until those files are public, the scores and the 3 to 4 times compute claim remain Reflection’s own figures. The waitlist is open, the red-team pass is unfinished, and the next training run is already underway.
-
NEWS1 month agoGeneration Lab’s Secret Youth Shot Has a Copycat Problem
-
AUTO1 month agoTesla Raises Dual Motor Prices as Texas Builds Cybercabs
-
LIFESTYLE2 months agoThe Nantucket Friendship Basket Boom Meets a Maker Shortage
-
NEWS1 month agoOkta Stock Jumps 20% on McKinnon’s Identity Bet
-
BUSINESS1 month agoBristol Myers Quits Cellares as Autoimmune Doses Proceed
-
LIFESTYLE1 month agoLabor Day Mattress Sales Repeat a Familiar Holiday Discount
-
NEWS1 month agoGoogle Ends EU Spam Demotions but Keeps the Ranking Split
-
BUSINESS2 months agoPoland Closes Visa-Free Work for Three Fast-Growing Nationalities
