frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: Avoiding the Memory Wall by computing LLM inference directly inside RAM

2•pcdeni•1h ago
The excitement surrounding PrismML’s 1-bit/ternary Bonsai models has the industry closely watching how smartphone giants, particularly Apple, will implement LLMs on edge devices. Moving AI on-device is a brilliant and necessary strategy. It ensures absolute user privacy in alignment with EU regulations, fundamentally shifts the economics away from costly cloud inference, and paves the way for a significant hardware upgrade supercycle as users seek true AI-capable silicon.

To create a smart on-device "Semantic Router," models need to reach the 27B+ parameter scale. Achieving this on a phone requires extreme quantization, such as PrismML’s ternary weights.

However, a critical hardware reality often overlooked by the software world is that fitting the weights in RAM is not equivalent to moving them. Running a 27B ternary model on standard LPDDR encounters a significant memory bandwidth limitation. Transferring gigabytes of data across the SoC bus for each token generation can lead to thermal throttling of the NPU and excessive battery drain.

This raises an important question: why are we still transferring data to the compute? Why not execute AI inference natively within the memory?

Frustrated with academic PIM simulations that overlook bare-metal physics, I developed CaSA, an architecture that performs ternary LLM inference directly inside COTS DRAM through charge-sharing, completely bypassing the memory bus.

Software quantization is a great initial step, and CaSA provides the physical hardware substrate needed to complete the bridge: https://github.com/pcdeni/CaSA

Show HN: Remux – an open-source tmux workspace designed for iPhone

https://github.com/h3nock/remux
17•bitwise42•1h ago•2 comments

Show HN: Whetuu – a zero-config cross-shell prompt written in Zig

https://yamafaktory.github.io/whetuu/
8•yamafaktory•1h ago•2 comments

Show HN: macOS menu-bar manager for SSH port forwards

https://github.com/lx2026/RelayBar
11•linxy97•1h ago•1 comments

Show HN: OneCLI – OSS credential gateway that keeps secrets out of AI agents

https://onecli.sh
2•Jonathanfishner•15m ago•0 comments

Show HN: I made YAFL – a E2EE file handoff for AI agents

https://yafl.dev
3•marcushorndt•1h ago•0 comments

Show HN: Find undervalued bikes near you with ML

https://bike.broker
3•azuroscloudapi•44m ago•0 comments

Show HN: A primer for non-coders to design software via structured AI dialogue

https://github.com/CoreGrowthLabs/CoreGrowthPrompting/blob/main/README.en.md
3•CGP_Labs•45m ago•0 comments

Show HN: I built a zero-latency developer tool suite in pure Vanilla JavaScript

https://omnideck.cc/
3•AAPD_Studio•1h ago•3 comments

Show HN: MinkNote – Native macOS notes and journaling using Markdown files

https://muse23.com/apps/minknote/
2•demianturner•46m ago•0 comments

Show HN: Palmier Pro – open-source macOS video editor built for AI

https://github.com/palmier-io/palmier-pro
3•hchtin•46m ago•0 comments

Show HN: What's AI's go-to, public or private healthcare?

https://www.modelbias.ai/prompt/public-or-private-healthcare
5•bnfcl•1h ago•2 comments

Show HN: ManiDesk – SQL workspace joining market, macro and fundamental analysis

https://mani-desk.com/
3•franklin_m•50m ago•2 comments

Show HN: STT-MCP – local STT for agents

https://github.com/sm18lr88/STT-MCP
2•hereme888•51m ago•0 comments

Show HN: Verify what an AI agent did, then tamper with the record (no signup)

https://api.cinchor.com/proof/refund
2•foh_quarters•57m ago•0 comments

Show HN: WatchMachineGo – A visualizer to show hardware performing LLM inference

https://watchmachinego.com/llm-inference
2•dev_dan_2•57m ago•1 comments

Show HN: Hanky – ETL style framework for loading flash cards into Anki

https://github.com/Haeata-Ash/hanky
4•funfruit•2h ago•0 comments

Show HN: Hosted PaddleOCR-VL-1.6 API

https://www.openparser.dev/
5•TimurKramar•59m ago•0 comments

Show HN: Dally – A little every day adds up

https://mctools.site
9•totaldude87•2h ago•0 comments

Show HN: Plox – Lox compiler in ~600 LOC from the 'Crafting Interpreters' book

https://github.com/eliasdejong/plox
2•eliasdejong•1h ago•0 comments

Show HN: GapQuery – find app-market gaps from 35,600 apps and their reviews

https://www.gapquery.com
2•northify•1h ago•0 comments

Show HN: BatchEdits – Batch AI video editor to edit content in volume

https://batchedits.com
3•shayannadeem321•1h ago•1 comments

Show HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrong

https://github.com/cactus-compute/cactus-hybrid
176•HenryNdubuaku•22h ago•37 comments

Show HN: Carta – An open-source pandoc reimplementation in Rust

https://github.com/mfkrause/carta
2•mfkrause•1h ago•0 comments

Show HN: 5dive – Run a Company of Claude Code/Codex Agents (Written in Bash)

https://github.com/5dive-ai/5dive
2•lodar•1h ago•0 comments

Show HN: 15 years of hand-curated architecture, art and design as an MCP server

https://www.thisispaper.com/intelligence/api-docs
2•zaxarov•1h ago•0 comments

Show HN: Setoku - self-hosted knowledge server for AI agents

https://setoku.com/
2•rgbrgb•1h ago•1 comments

Show HN: Personal Jarvis, an open source voice assistant that runs your PC

https://github.com/PersonalJarvis/PersonalJarvis
2•PersonalJarvis•1h ago•0 comments

Show HN: Avoiding the Memory Wall by computing LLM inference directly inside RAM

2•pcdeni•1h ago•0 comments

Show HN: ArXivMax – video explainers for any research paper

https://www.arxivmax.com/
2•rwu1997•1h ago•0 comments

Show HN: Note taking and book Library, all-in-one AI KnowledgeBase

https://github.com/pileax-ai/pileax
2•pileax•1h ago•0 comments