frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: Running PrismML's Bonsai inside DRAM by breaking DDR4 timing rules

16•pcdeni•4d ago
The excitement surrounding PrismML’s 1-bit/ternary Bonsai models has the industry closely watching how smartphone giants, particularly Apple, will implement LLMs on edge devices.

Moving AI on-device is a brilliant and necessary strategy. It ensures absolute user privacy in alignment with EU regulations, fundamentally shifts the economics away from costly cloud inference, and paves the way for a significant hardware upgrade supercycle as users seek true AI-capable silicon.

To create a smart on-device "Semantic Router," models need to reach the 27B+ parameter scale. Achieving this on a phone requires extreme quantization, such as PrismML’s ternary weights.

However, a critical hardware reality often overlooked by the software world is that fitting the weights in RAM is not equivalent to moving them. Running a 27B ternary model on standard LPDDR encounters a significant memory bandwidth limitation. Transferring gigabytes of data across the SoC bus for each token generation can lead to thermal throttling of the NPU and excessive battery drain.

This raises an important question: why are we still transferring data to the compute? Why not execute AI inference natively within the memory?

Frustrated with academic PIM simulations that overlook bare-metal physics, I developed CaSA, an architecture that performs ternary LLM inference directly inside COTS DRAM through charge-sharing, completely bypassing the memory bus.

Software quantization is a great initial step, and CaSA provides the physical hardware substrate needed to complete the bridge: https://github.com/pcdeni/CaSA

Comments

SwellJoe•40m ago
Jebus, that is some sloppy prose. Can people not even be bothered to write the summary themselves, anymore? AI doesn't want anything, so they can never have a point of view, so their prose rambles incoherently across all the various prompts they've seen in a project. This project sounds like the ramblings of a crazy person. Even though the fact that DRAM can do any computation is interesting, nobody should have to read this mess.

"why are we still transferring data to the compute? Why not execute AI inference natively within the memory?"

You already answered that question: 47.5 seconds per token from a tiny 2B 1-bit model model.

deivid•35m ago
Interesting project, but the slop readme made me quit reading halfway
butvacuum•34m ago
Very interesting. I don't see it mentioned so I'll ask:

Would being able to alter voltage levels on the fly (eg, cells x y and z get 1.25 while abc get 1.2) expand the ability here?

ilaksh•18m ago
You are saying this is 13 times faster? More proof please. How do we set it up? I really want this to be a real thing.

Show HN: FeyNoBg – Automatic background removal model and training library

https://usefeyn.com/blog/feynobg/
27•snyy•1h ago•8 comments

Show HN: Running PrismML's Bonsai inside DRAM by breaking DDR4 timing rules

17•pcdeni•4d ago•4 comments

Show HN: Infrawrench – a tool to manage cloud and svcs with workflows and chat

https://infrawrench.com
8•astrid__•2h ago•0 comments

Show HN: Darkslide – A keyboard-centric photo editor

2•roman-volkov•1h ago•0 comments

Show HN: Physically accurate black hole you can put in your room

https://blackhole.plav.in
451•aplavin•4d ago•164 comments

Show HN: Pure-PHP server outperforms NginX for static files + 10x PHP throughput

https://github.com/Qbix/webserver
4•EGreg•3h ago•0 comments

Show HN: 1,250 SwiftUI components, and an MCP that writes them into your app

https://nibware.dev/
4•leonickson•3h ago•4 comments

Show HN: multiaes – hardware-accelerated, constant-time AES, two-file drop-in

https://github.com/ttarvis/multiaes
3•lemaudit•3h ago•0 comments

Show HN: Watch 14-Byte AI "brains" attempt to solve a 2D maze (Its hard)

https://con-dog.github.io/MINIMIO-PUBLIC-FRONTEND/
20•purple-leafy•11h ago•4 comments

Show HN: Pilot Protocol – a network where AI agents find tools and each other

6•teocalin37•4h ago•4 comments

Show HN: Tilde Pay – Give your AI agent a bank account to pay for things

https://my.tildepay.ai/
2•solsol94•4h ago•1 comments

Show HN: Dozenal – A Game of Spatial Arithmetic

https://dozenal.game
13•sarreph•7h ago•8 comments

Show HN: Watch random code typed out on an MS-DOS IDE

https://hackerman.specr.net/
14•vunderba•14h ago•6 comments

Show HN: Reverse Minesweeper

https://sunflowersgame.com/
245•pompomsheep•1d ago•85 comments

Show HN: Descript wanted $24/mo, I built an open-source alternative in a weekend

https://github.com/wassgha/rescript
35•wassimgr•12h ago•27 comments

Show HN: I mapped every US golf course

https://golfcoursebrowser.com/
214•rickmf•1d ago•161 comments

Show HN: CheapSecurity – Lightweight, Self-Hosted CCTV for Linux SBCs

https://github.com/gmrandazzo/CheapSecurity
135•zeldone•1d ago•34 comments

Show HN: Case study: A coding agent refactors a 750k LOC app, no code review

5•bonjourjoel•6h ago•0 comments

Show HN: PumpProof – Scores a hyped stock's dilution risk from its SEC filings

https://pumpproof.com
5•jekjek•6h ago•0 comments

Show HN: Tool to turn a repo into map

https://codemap.gitbiased.com
4•skyfantom•6h ago•1 comments

Show HN: Turn SSH X11 Forward to WebSocket – remote apps flow to you workstation

3•kbradero•7h ago•1 comments

Show HN: Horus-runtime – High Performance Computing workflow manager

https://github.com/temple-compute/horus-runtime
2•chdominguez•7h ago•0 comments

Show HN: JavaScript/JSON and CSS minifier lib in minimal C89

https://fossil.wanderinghorse.net/r/cssminc/doc/ckout/README.md
3•sgbeal•8h ago•0 comments

Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

479•adam_rida•3d ago•223 comments

Show HN: Markwright – a 2MB native screen annotator with no telemetry

https://www.aloneguid.uk/projects/mw/
2•aloneguid•8h ago•1 comments

Show HN: Paper – A Journaling CLI for Developers

https://paper.rewrlution.com
9•dmsehuang•19h ago•11 comments

Show HN: Agent Console – A Local Dashboard for Codex and Claude Code

https://github.com/buhuipao/agent-console
2•buhuipao•8h ago•0 comments

Show HN: Gitwig – Mouse-drivable Git TUI and multi-repo dashboard in Rust

https://gitwig.dev/
2•tareqmy•9h ago•0 comments

Show HN: We built an MCP server for document generation

https://docuqueue.com/#mcp
3•dvcoolarun•9h ago•0 comments

Show HN: adCasa OS – AI marketing workspace built with Bayesian attribution

https://adcasa.io/
2•adcasa•10h ago•0 comments