frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Beam: Reflection's 501B open-weight model

https://reflection.ai/blog/introducing-beam
91•Philpax•59m ago

Comments

htrp•57m ago
> Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.

> Beam’s capabilities come from major investments in both pretraining and reinforcement learning (RL). We pretrained the model on 23.8 trillion diverse, curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming available similar-sized open base models. In parallel, we developed the algorithms, training environments, and infrastructure needed to sustain high-compute RL at exceptional scale. Our high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks of training.

Early access, no weights no tech details, just a sign up here for info

zelphirkalt•36m ago
And also a "proprietary data set" hahaha... Probably just means they don't want to show it, and it is data, that either they shouldn't have, or that there is nothing special about their training data and it is just meant to sound like there is some secret ingredient, while there is none.
wronglebowski•34m ago
I'm all for more open models, but talk is cheap and this is a rather pointless announcement without anything backing it up. Publish your weights and HF repo or shut up IMO.
wg0•47m ago
Suppose I inherited a data center spanning several hundred acres full of GPUs and free electricity.

Where do I get the data?

I mean, this many models. They have to start somewhere.

Hamuko•45m ago
Get data from Claude. That's what the Chinese (allegedly) do.
kbwal7•6m ago
Note that this sort of distillation is NOT for pre-training data (which is tens of trillions of tokens). I think the allegations against Chinese companies by Anthropic is more so that they distill SFT data (which is good for post-training, but you still need a strong base model)
altcognito•41m ago
If you ask a model, they will generally tell you where to get data. Modern frontier models have the large advantage of having tens if not hundreds of millions of users providing use cases to train against to improve their responses.
petu•32m ago
I guess public datasets on HuggingFace and some shadow libraries content is enough to start.

e.g. fineweb dataset is 50TB https://huggingface.co/datasets/HuggingFaceFW/fineweb

lucrbvi•
Ariarule•45m ago
Always glad to see more open-weight models, but this caption on the 2nd demo image had me do a double-take: "Land or Water Generalization Experiment: We recreated the viral X puzzle by asking Beam to create a fixed 180×90 grid for longitudes -179° to 179° and latitudes -89° to 89°, with 16,200 points. This puzzle is a few days old, so could not appear in the training data, thus testing the model’s generalization. Beam gets 95.5% coverage right, putting us between Opus 5 (92.5%) and Fable 5 (97.8%), which shows how well it generalizes to novel new tasks."

Oof, no, this "puzzle is a few days old" is incorrect even if it's a social media trend just recently. Asking a model to generate a world map in this way is _at least_ from August 2025 as it appeared on LessWrong at that time: https://www.lesswrong.com/posts/xwdRzJxyqFqgXTWbH/how-does-a...

extr•27m ago
Yeah I remember when the original post about this came out. Def not recent. Though I think their point survives in that they didn't exactly RL on this.
sharktheone•41m ago
Am I the only one who thought of the BEAM VM after the first word of thee title?
pstuart•29m ago
nopes
drubs•34m ago
I remember being in the room with pretraining day 1 to help monitor the training job launch. Watching this model train from day 1 has been an amazing experience!
aeetes•28m ago
the performance chart puts the better open source models behind the fold making it seem like it outperforms them... but it doesn't! all for open source models but this announcement is misleading
brumbelow•20m ago
Yes. All the link made me realize is that I should checkout Deepseek 4.1 flash
keeganpoppen•27m ago
very curious to see more about what kinds of hardware you can run this on and the perf. characteristics… on the face of it, it seems like optimizing for inference speed might(?) be good for running on smaller hardware, but i suppose it could be the other way around and it is actually much resource-hungrier for the number of parameters, etc. …
NorwegianDude•27m ago
Bigger and still worse than existing free Chinese models that are smaller? Open weight models are nice, but at this point it seems western models are very far behind Chinese ones, despite Chinese companies publishing a lot of their findings. I hope we get more open models and more providers, as being stuck with a model from China or US with no competition is risky.

Google does do a great job with Gemma models. It's one of the few language models actually good at language. OpenAI's top closed models can't even write norwegian correctly.

onlyrealcuzzo•26m ago
This appears to be larger than DeepSeek v4.1 Flash, more expensive to run, and worse on every measured metric.

Am I missing something?

martini333•15m ago
Beam goes brrrr
dotancohen•7m ago
We're still at the stage where every new entrant is welcome in my opinion. Doesn't need to be record-breaking upon initial release.
jstummbillig•6m ago
Apparently there is more to making good models than copying everything on the internet.
zopper•16m ago
Open model that is not yet open or widely accessible via API. Primarily comparing to non-SOTA models like Inkling and GLM 5.2. Included comparison to GLM 5.3 and DeepSeek V4.1 Flash in the table, but not in the charts (I assume they would make them look bad). Also no results from AA Index or Arena.
28m ago
There are a lot of open-research on pre-training, post-training and RL data mixtures and sourcing.

I recommend checking papers from Datalogy, Nvidia Nemotron, Ai2 (Ollmo, Tulu, ...) and the recent model from Aleph Alpha if you want to learn more.

konfusinomicon•7m ago
forget the data....sell it and go live your life!

Beam: Reflection's 501B open-weight model

https://reflection.ai/blog/introducing-beam
94•Philpax•59m ago•25 comments

Why Plain Text Is Still One of the Best Technologies We Have

https://deadparrotbbs.com/why-plain-text-is-still-one-of-the-best-technologies-we-have/
63•speckx•1h ago•18 comments

Web Search API

https://developers.cloudflare.com/changelog/post/2026-10-02-introducing-web-search-api/
393•tosh•9h ago•188 comments

The future of independence is interdependence

https://onlys.ky/independence-is-interdependence/
114•eustoria•4h ago•71 comments

Making a GTK application in Haskell, part 1

https://floreal.tech/blog/2026/making-a-gtk-app-in-haskell-part-1/
109•Vosporos•5h ago•21 comments

OpenAI "rogue" agent activities found on Wikimedia projects

https://diff.wikimedia.org/2026/10/05/openai-rogue-agent-activities-found-on-wikimedia-projects/
198•brokensegue•2h ago•133 comments

Competitive Programmer's Handbook (2018) [pdf]

https://cses.fi/book/book.pdf
21•vinhnx•2d ago•1 comments

Linux containers in 500 lines of code

https://blog.lizzie.io/linux-containers-in-500-loc.html
62•mkornaukhov•6h ago•13 comments

Denmark Data Breach Exposes 8.8M People's Personal Data

https://www.cpr.dk/cpr-nyt/nyhedsarkiv/2026/okt/omfattende-uautoriseret-adgang-til-borgeres-cpr-o...
429•clan•12h ago•307 comments

Show HN: Nightwatch – a Mac menu-bar app that tells you when tonight is clear

https://github.com/rsutcliffe/nightwatch
44•delphidolphin•1d ago•5 comments

Anthropic reported diary entry to police, woman faces felony charge

https://www.techspot.com/news/114091-florida-woman-used-claude-diary-anthropic-reported-shoot.html
324•emptybits•14h ago•245 comments

Norway Eyes Partial Ban of Smart Glasses

https://www.barrons.com/news/norway-eyes-partial-ban-of-smart-glasses-e65dc239
47•pseudolus•2h ago•15 comments

Huawei and Qualcomm announce broad patent license agreement

https://www.huawei.com/en/news/2026/10/qualcomm-broad-patent-agreement
159•0xedb•12h ago•97 comments

One person is now a quorum at the SEC

https://www.ft.com/content/3120782c-1ea0-4fdc-9462-0a4b4658f70f
36•mooreds•1h ago•9 comments

The designer lamps in my house

https://arslan.io/2026/10/05/the-lamps-in-my-house/
61•farslan•6h ago•34 comments

AI Companies Are Parasites

https://www.coryd.dev/posts/2026/ai-companies-are-parasites
40•cdrnsf•46m ago•17 comments

The technology to eradicate mosquito-borne disease exists

https://worksinprogress.co/issue/mosquitoes-are-a-choice/
199•benbreen•1d ago•151 comments

Differences Between `Foldl` and `Foldr`

https://blog.haskell.org/foldl-and-foldr/
116•signa11•4d ago•25 comments

Greenvolt begins building 600 MW/2.4 GWh BESS in Poland

https://www.ess-news.com/2026/09/25/greenvolt-begins-building-600-mw-2-4-gwh-bess-in-poland/
25•msalsas•1h ago•15 comments

Beating the Compiler

https://www.mattkeeter.com/blog/2024-07-12-interpreter/
50•andsoitis•4d ago•32 comments

Mold Linker Version 3.0.0 Release – Rewritten in Rust

https://github.com/rui314/mold/releases/tag/v3.0.0
206•roflcopter69•8h ago•119 comments

2026 Nobel Prize in Physiology or Medicine: Deisseroth, Hegemann, Nagel

https://www.nobelprize.org/prizes/medicine/2026/summary/
83•lode•10h ago•39 comments

Pixel 11 doesn't yet meet the GrapheneOS security standards and may be skipped

https://discuss.grapheneos.org/d/41564-pixel-11-doesnt-yet-meet-the-grapheneos-security-standards...
364•finnlab•7h ago•202 comments

US closely monitoring case of lab worker who possibly died of plague in Siberia

https://www.theguardian.com/world/2026/oct/05/russia-lab-worker-possibly-dies-of-plague-siberia-q...
174•tosh•3h ago•156 comments

Martian chaos terrain

https://en.wikipedia.org/wiki/Martian_chaos_terrain
75•tiagod•3d ago•19 comments

Hot Flashing Guide Rev. 2.0 (2004)

https://archive.techarp.com/showarticle504a.html?pgno=0
6•userbinator•2d ago•0 comments

In the wake of Tippett Studios’ closure, a digital archive appears online

https://filmstories.co.uk/news/tippett-studios-in-the-wake-of-its-closure-a-digital-archive-of-an...
231•rdmuser•23h ago•30 comments

Mystery Function

https://codeset.ai/function
35•andre15silva•3d ago•10 comments

A browser-native classic Visual Basic VB6 IDE

https://wieslawsoltes.github.io/VB6/
382•wiso•1d ago•126 comments

We ported the original Doom to SQL

https://cedardb.com/blog/sqldoom/
290•Vaslo•1d ago•44 comments