frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: Train 300M/32-Layer Model in 1.5GB RAM on Base M1 Mac

1•vlad_kalinkin•53m ago
Hello again! Since my last post about Ullis, the project has gone through several changes as I was searching for the right architecture. As it turned out, KAN and Hyena were quite resource-heavy, and I couldn't get anything viable out of them. Then I tried RWKV, specifically the latest RWKV-8 Heron version with 1-bit ROSA activation. This ultimately proved to be the most viable architecture of all. As an example: on my 2020 MacBook Pro M1 with 8GB RAM and a 68GB/s memory bandwidth, I managed to start training a model with ~272 million parameters, a 2048 context length, and 32 layers—all within a 1.5GB memory footprint. On my specific hardware, this is still a highly taxing task. Due to the low bandwidth, the total latency per step is around half a minute, and the throughput drops to about 70 tokens per second. To be clear right away: I had to keep ROSA SAM on the CPU because it proved to be more efficient for these tasks than the GPU. Here are the logs as proof:

cargo run --release -- train \ --data data/ullis_dataset.jsonl \ --run runs/hard \ --config train_config.json \ --steps 20000 \ --learning-rate 0.005 \ --checkpoint-every 500 \ --bpe-train-mib 150

ullis: token stream compacted to u16 (156 MiB) ullis: compiling Metal shaders ullis: Metal ready in 0.2s ullis: Metal train: LN/QKV/CMix/head on GPU; ROSA SAM on CPU (~603979776 bytes idx+y+out/step) ullis: starting loop after 379.0s of setup ullis: clipped SGD lr=0.005 rosa_grad=StopGradBits (QKV frozen; window-mean CE on FP16, token-sum STE on BinaryConnect; |w0|=0.01) step 1/20000 loss=7.9256 ema=7.9256 p10=5.218 p50=8.193 p90=10.197 unigram=5.958 unique=597 n=1389 flips=0/0/0 (head/cmix/o) embed_grms=6.40e-5 scale_grms=8.09e-3 cmix_vrms=0.013 |w|=0.010 dw=4.37e-6 bias_rms=2.033 resid=3.21e-8 rss=1643MiB 27881ms 73 tok/s ullis: step 1 phases embed=8ms fwd_ln=370ms fwd_rosa=7220ms fwd_cmix=5897ms head=1070ms bwd_ln=625ms bwd_cmix=9754ms bwd_rosa=2678ms embed_sgd=57ms step 2/20000 loss=7.0537 ema=7.8384 p10=1.670 p50=7.518 p90=9.969 unigram=5.564 unique=590 n=1520 flips=0/0/0 (head/cmix/o) embed_grms=6.46e-5 scale_grms=1.06e-2 cmix_vrms=0.013 |w|=0.010 dw=4.68e-6 bias_rms=2.033 resid=5.67e-8 rss=1628MiB 27997ms 73 tok/s ullis: step 2 phases embed=8ms fwd_ln=314ms fwd_rosa=5717ms fwd_cmix=4738ms head=1231ms bwd_ln=786ms bwd_cmix=11755ms bwd_rosa=3180ms embed_sgd=60ms

With scaled-down settings, training runs with pretty satisfactory performance. Here is an example: since a multi-stage training pipeline with pre-training is not implemented yet, I trained it on a large Claude Opus distillation dataset. After 5500 steps, here is the result:

cargo run --release -- generate \ --checkpoint runs/ullis_b1_32m/checkpoint.safetensors \ --prompt "2+2" \ --temperature 0.4 \ --top-p 0.8

    Finished `release` profile [optimized] target(s) in 0.60s
     Running `target/release/ullis generate --checkpoint runs/ullis_b1_32m/checkpoint.safetensors --prompt 2+2 --temperature 0.4 --top-p 0.8`
- *No a single = = = = = = i < (r: $=3_x(d: $d = \frac{[b_t + 2>[b_t + 1 + 3-c)$

- The correct answer is a list of the number of the first n > 0

As you can see, the model is capable of learning, but since I cannot afford to keep my main work laptop running training tasks 24/7, I had to stop at this relatively modest result. There is, of course, a possibility that some training algorithms might have implementation bugs, but this is the current outcome. I hope you find this project interesting. Honestly, pulling off a project like this entirely on your own is quite tough, especially considering I'm developing it alongside an LLM assistant (which constantly tries to break things). I'm not deeply experienced in this field yet, so I'd be incredibly glad to find testers, people with domain expertise, or anyone willing to share any kind of feedback. Thank you!

Link: https://github.com/Vladislav-Kalinkin/ullis

Physically Immutable Optical Archive Libraries

https://savartus.com/solutions/enterprise-laser-storage/
2•thunderbong•4m ago•0 comments

One Nix flake to rule them all

https://fzakaria.com/2026/08/28/one-flake-to-rule-them-all
2•ingve•9m ago•0 comments

British Public Oppose Secret Surveillance Powers and Want Strong Protections

https://cdt.org/press/british-public-oppose-secret-surveillance-powers-and-want-strong-protection...
2•DeepLogin•10m ago•0 comments

When it comes to China, America has a plan

https://mondediplo.com/2026/06/11us
1•Anon84•17m ago•0 comments

Show HN: 1endpoint – Cheaper access to AI models

https://1endpoint.dev
1•DustinPham12•18m ago•1 comments

How Culture Shapes the Stories We Tell About Our Emotions (2024)

https://behavioralscientist.org/culture-shapes-stories-of-emotion/
1•the-mitr•19m ago•0 comments

Some Scientists Have 'Magic Hands' in the Lab. This A.I. Is Learning Why.

https://www.nytimes.com/2026/08/27/science/scientists-experiments-replication-ai.html
1•mitchbob•20m ago•1 comments

What It Took to Dismantle the Most Powerful Company in the World

https://www.nytimes.com/2026/08/28/opinion/ai-power-lobbying-military-britain-east-india-company....
1•Anon84•23m ago•0 comments

Atlantropa

https://en.wikipedia.org/wiki/Atlantropa
1•yitchelle•24m ago•0 comments

Iceland votes 'No' to reopening EU accession talks

https://www.euronews.com/my-europe/2026/08/30/iceland-votes-no-to-further-eu-accession-talks-in-t...
2•sehw•26m ago•0 comments

Rehearse a specific English work conversation before you have it

https://www.unfreezy.com
1•czaar1•26m ago•0 comments

Groom of the Stool

https://en.wikipedia.org/wiki/Groom_of_the_Stool
2•soupspaces•27m ago•0 comments

IguanaTex

https://github.com/Jonathan-LeRoux/IguanaTex
1•Eridanus2•28m ago•0 comments

Innovation Doesn't Happen on Schedule

https://www.thesignalist.io/s/innovation-doesnt-happen-on-schedule/
1•kodesko•29m ago•0 comments

The largest electric plane takes flight today

https://www.popsci.com/technology/worlds-largest-electric-plane/
1•thinkingemote•29m ago•0 comments

Scientists Create the Littlest Big Bang to Study the Universe's Origins

https://www.wired.com/story/scientists-create-littlest-big-bang-to-study-universe-origins/
1•joozio•30m ago•0 comments

When Led Zeppelin bombed in front of 1.9B people [video]

https://www.youtube.com/watch?v=PPJJdQfVxyg
1•jurf•30m ago•0 comments

Show HN: The Thousand – 1k founders, 1 question, one book

https://thethousand.co
1•jabed•30m ago•2 comments

Google Maps – Lake America

https://www.google.com/maps/place/Lake+America/@43.7036865,-79.2490163,8z/data=!3m1!4b1!4m6!3m5!1...
1•pseudolus•31m ago•0 comments

Columbia House Is Shutting Down

https://pitchfork.com/story/columbia-house-is-shutting-down/
1•pseudolus•39m ago•0 comments

He Did Not Conquer: Franklin's Failure to Annex Canada

https://reviewcanada.ca/magazine/2025/12/how-vain-an-attempt-review-he-did-not-conquer/
2•Bluestein•40m ago•0 comments

Show HN: Get your free AI search visibility scan

https://www.visiscan.app
1•not_wowinter14•40m ago•0 comments

Tortoise and the Birds – an African folktale recreated with Lego animation [video]

https://www.youtube.com/watch?v=EVHlqLa7-u0
1•CodePapi_•44m ago•1 comments

Show HN: Nuzzle – adorable live wallpapers of your pets

https://usenuzzle.com
1•13001r•44m ago•0 comments

Against Dumb Vibe Coding, Not Vibe Coding

https://funnygeek.com/essays/vibe-coding-speech/
1•ifoo•46m ago•1 comments

Show HN: Drop a SQL schema, get an interactive ER diagram

https://mcdview.dev/
2•PatuDev•46m ago•0 comments

Mitchell Hashimoto: Basic Superlogical Demo

https://twitter.com/mitchellh/status/2093451043661316217
2•mdrachuk•48m ago•0 comments

The U.S. Munitions Crisis

https://www.cfr.org/articles/the-u-s-munitions-crisis
2•Betelbuddy•49m ago•1 comments

Show HN: Time tracking for solo consultants that ends in a PDF invoice

https://hourtobill.com/demo
1•Stackedboost•50m ago•0 comments

Tja – TUI for learning German verbs in your terminal, one-armed-bandit style

https://github.com/progapandist/tja
2•progapandist•50m ago•2 comments