frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

https://github.com/magnitudedev/magnitude
41•anerli•56m ago•18 comments

Commit Description as a Thinking Tool

https://yedhu.me/posts/commit-description-as-a-thinking-tool/
44•yedhukrishnan•1h ago•7 comments

A brief history of the Bloomberg terminal

https://spectrum.ieee.org/bloomberg-terminal
120•rbanffy•4h ago•42 comments

You Said No MCP

https://earendil.com/posts/you-said-no-mcp/
498•yarapavan•8h ago•278 comments

Burning Man Death Rates – A Short Lesson in Statistics

https://ihavenapkinthoughts.substack.com/p/burning-man-death-rates-a-short-lesson
44•viraj_shah•1d ago•22 comments

I Could've Accessed 17T Microsoft Records

https://blog.faav.net/how-i-couldve-accessed-17-trillion-microsoft-records
154•luispa•1d ago•72 comments

SDF vs. MSDF vs. Slug: GPU Text Rendering

https://alphapixeldev.com/sdf-vs-msdf-vs-slug-vs-rive-gpu-text-rendering/
88•ibobev•4h ago•35 comments

Moist-Electric Wallpaper for Indoor Energy Harvesting and Humidity Management

https://advanced.onlinelibrary.wiley.com/doi/10.1002/aenm.71603
33•croes•2h ago•24 comments

Bild AI (YC W25) Is Hiring a Founding Product Engineer

https://www.ycombinator.com/companies/bild-ai/jobs/dAbC3Gd-founding-product-engineer
1•rooppal•1h ago

What TLA+ can and can't check

https://buttondown.com/hillelwayne/archive/what-tla-can-and-cant-check/
49•b-man•4h ago•9 comments

SDF Public Access Unix System ... est. 1987

https://sdf.org/
37•kmstout•3h ago•3 comments

5x faster Edge Functions: V8 isolates to Firecracker MicroVMs

https://www.netlify.com/blog/edge-functions-firecracker-microvms/
5•jbott•16m ago•0 comments

The last time my family was replaced by technology

https://manuel.darcemont.fr/posts/the-last-time-my-family-was-replaced-by-technology/
68•megalomanu•5h ago•147 comments

Reverse-engineering a $35 backup camera display (AMT630A)

https://github.com/mogrinz/AMT630A
40•mogrinz•1d ago•11 comments

Livenerf: Has Opus 5.5 been nerfed yet?

https://github.com/ninjahawk/livenerf
836•bryan0•19h ago•352 comments

Show HN: JBR-001 – An open-source 3D printable desktop robot

https://projecthub.arduino.cc/syntheticaidata/jbr-001-a-desktop-companion-robot-powered-by-arduin...
108•gvuksic•1d ago•24 comments

Energy Timelines Photovoltaic

https://www.eia.gov/kids/history-of-energy/timelines/photovoltaic.php
4•thelastgallon•1d ago•1 comments

LinkedIn Larpmaxxing

https://hereticpleb.vercel.app/blog/linkedin-larpmaxxing/
63•BurnerBurner•14h ago•70 comments

Solving Factorio Quality

https://exyr.org/2026/solving-factorio-quality/
228•laurenth•1d ago•84 comments

Mathematical Origami

https://mathigon.org/origami
92•signa11•1d ago•21 comments

Getting out of the way: my robotics crash course

https://thisismypersonalblog.com/posts/2026-09-25-getting-out-of-the-way/
44•systemerror•2d ago•9 comments

Dots: Always-on agents

https://openai.com/index/introducing-dots/
724•alvis•1d ago•611 comments

Vermont replacing power plants with home batteries

https://www.bbc.com/future/article/20260928-a-virtual-power-plant-hidden-in-vermont-homes-is-keep...
345•devonnull•1d ago•255 comments

NASA asked several former SR-71A staffers to help secret restart

https://aviationweek.com/defense/aircraft-propulsion/nasa-asked-several-former-sr-71a-staffers-he...
285•ilamont•1d ago•329 comments

Show HN: Ledge.sh – Runnable Markdown Notes

https://ledge.sh
33•dancablam•18h ago•29 comments

Show HN: Dental Scope – Interactive 3D dental anatomy

https://dental-scope.com/
83•Zeruxe•1d ago•49 comments

Show HN: A working 3D model of an Enigma machine

https://enigma.design
80•primitivesuave•1d ago•31 comments

Show HN: Strata – an expressive semantic layer that can say no to your LLM

https://strata.do/try/
4•ajoski9•3h ago•0 comments

Show HN: Real-time Solar System with 526k asteroids and all tracked satellites

https://space.bl2.net/
359•wanick•23h ago•95 comments

U.S. postal inspectors shut down website selling counterfeit postage labels

https://postalemployeenetwork.com/news/2026/09/26/u-s-postal-inspectors-shut-down-website-selling...
270•ilamont•23h ago•161 comments
Open in hackernews

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

https://github.com/magnitudedev/magnitude
41•anerli•56m ago
Hey HN, Anders and Tom here. We're building Magnitude, an inference engine for agents that optimizes itself to run as fast as possible on your hardware. It works on Mac, Linux, and Windows on any hardware and is up to 2x faster than llama.cpp.

We're both software engineers and previously built an open source browser agent to 4k+ GH stars and 100k+ downloads. We increasingly wanted to run it on local models, but found that no inference engine worked for our use case.

Inference engines today all make a performance tradeoff. They are either:

- Built for batched inference on datacenter hardware at the cost of single-session performance (vLLM, SGLang) - Designed for broad compatibility instead of optimizing for specific hardware (llama.cpp, Ollama) - Specialized for specific hardware or models but lacking engine completeness (oMLX, ds4)

Plus none of them are designed for running agents locally. Sessions are long, several often run at once, and you still want to use your computer for other things.

Magnitude is built for maximum performance on your hardware and running local agents:

- On-device compilation and tuning: Kernels are written with flexible parameters that are tuned on your actual device before the model runs. This gives you broad hardware compatibility with the same performance ceiling as hardware-specific kernels. - Focus on best architectures: We write our tunable, highly efficient kernels for the most popular open-weights families. This allows us to achieve and surpass the performance of hardware or model specialized engines, without forcing ourselves to over-generalize at the cost of performance. - Dynamic memory allocation: Magnitude reserves only enough memory up front to hold model weights. As your agent sessions grow, the memory heap dynamically increases, and frees itself when agents stop. Your hardware can still be used for other stuff while agents run. - Hybrid paged attention: We borrow the best ideas from engines like SGLang to allow concurrent sessions to share prefix caches, but optimize placement for memory-adjacency so single-session performance doesn't suffer.

Magnitude is fully open source (Apache 2.0). We built it in Rust, including a custom GPU kernel runtime and autotuner. We take inspiration from the best innovations in inference from academics (e.g. FlashAttention, FlashInfer, TurboQuant) as well as other engines (e.g. SGLang radix attention) to reach the performance ceiling.

Benchmarked against llama.cpp with Qwen 3.6 35B A3B (4 bit), 64k context, no speculative decoding:

Metal (Mac M4 Pro 48 GB) - 92% faster decode (30 tok/s → 57 tok/s) - 9% faster prefill (466 tok/s → 507 tok/s) - 28% less per-agent memory usage

CUDA (DGX Spark) - 19% faster decode (49 tok/s → 58 tok/s) - 23% faster prefill (2,033 tok/s → 2,507 tok/s) - 27% less per-agent memory usage

Magnitude ships as a desktop app that you can easily connect with whatever agents you already use (Pi, OpenCode, Hermes, Codex, and more). It automatically runs models on demand when these agents actually need them, and shuts them down after inactivity. Here's what it looks like: https://www.youtube.com/watch?v=0qE8BWEZu7o

We're excited to push Magnitude further to let you run bigger models on the same hardware while continuing to improve performance. Our plans include:

- Expert streaming: store experts on RAM or disk and load them just-in-time. This lets you run models bigger than what otherwise would fit on your GPU. - Kernel compiler: our current kernels tune a few parameters to fit your hardware. We can take this further with a fully custom compiler that automatically chooses how to fuse kernels and which implementations to use, to make it fit to your hardware even better. - Multi-device utilization: Make the best possible use of all hardware on a system (CPU, GPUs, RAM, disk) by detecting these and automatically solving for the best model layout.

We'd love for more people to try it out and give us feedback. Feel free to comment here, we'll be around all day!

Comments

p-e-w•49m ago
What is the business model?
anerli•27m ago
We envision a future where workloads are hybrid. Average consumer hardware will be able to handle a lot with local models, but you’ll still want to use cloud models for harder tasks. Magnitude will make it seamless to switch between the two, even for the same tasks (without breaking your prefix cache). We’ll charge per token for our inference cloud, using the same efficiencies we unlock for local inference to pass the savings on to you.
amirhesham•43m ago
Oh this is so cool. Curious about the business model, too.
nateb2022•42m ago
Any source on the benchmarks/methodology besides the image? There's a ton of variance possible in llama.cpp's performance depending on how it was configured. I'd also like to see benchmarks against MLX.
anerli•15m ago
The benchmark we cited here is a simple prose-repetition task. We put the content of Moby Dick up to 64k context in the request, and then ask it to repeat the last section.

For llama.cpp, we try to make the comparison as fair as possible by using similar settings. No speculative decoding, default prefill batch sizes, flash attention on.

We tried also quantizing the KV cache to 8-bit keys and 4-bit values like we do in Magnitude, but this bombed decode speed for llama.cpp in our testing. Since it seems llama.cpp did not optimize that path, we used 16-bit KV instead.

The source for the benchmark is available here also: https://github.com/magnitudedev/magnitude/tree/main/inferenc...

anerli•13m ago
Compared to MLX - we've done some rough benchmarking and we are outperforming any of the MLX-based engines we've compared to so far. Going to do more in depth benchmarking and release it soon.
sgtwompwomp•42m ago
This is dope, is this kind of like Wafer.ai but for local models? As in a coding agent optimizes the kernels so the local model runs continuously better? Cause that is compelling if so. If it’s more simple that’s cool too
kenzic•39m ago
How long does tuning take (on an M3 MacBook Pro for example)?
anerli•24m ago
Tuning is a one-time process that takes around ~1 minute whenever you download a new model. This is generally enough time to tune all the kernels' parameters to the point where tuning any longer asymptotes. Time can vary a little based on the hardware though.
kenzic•14m ago
Wow, that's impressive.
yolandac•35m ago
does it allow us to run larger models that weren't possible before?
kmike84•27m ago
This seems to be a good idea. However, beating llama.cpp on speed is a low bar :)

I found it to be a good baseline, but at least on Mac there was always something way faster, and/or with better memory requirements - like you said, ds4, omlx, mtplx, etc. It seems if you use local LLMs for real, there is very little reason not to use one of the more optimized engines.

3 main failure modes I observed in the engines:

* Not using best available spec decoding

* Using too much VRAM for KV cache (e.g. KV cache used to take almost nothing in ds4, but huge amount of VRAM on unsloth/llama.cpp for deepseek models)

* Degraded performance at large context sizes - benchmarks at 4K or 32K are awesome, but at realistic 100-200K it's slower than some stupid baseline

anerli•8m ago
Yeah these are all things that we directly tackle!

Spec decoding: Models in our catalog come assigned with an assigned drafter model for speculative decoding based on the best known method and model available for that target model (support DFlash, DSpark, and DFlash2).

Using too much memory for KV cache: We use a TurboQuant-inspired quantization of KV cache to 8-bit keys and 4-bit values. This drops KV memory usage by over half and also speeds up decode. Based on long context quality benchmarking we've done it does not seem to negatively impact retrieval or coherence over long context.

Large context sizes: our KV quantization helps a lot for this, and we focus our optimizations on specifically longer-context requests since that's what most agent inference actually looks like.

lxe•21m ago
On my local inference box I have a perpetual codex thread open in my llama.cpp checkout that I periodically ask to take a look at currently pending llama.cpp PRs, do some research on latest MTP, Dflash and other prediction or attention optimizations, do research on the latest model quants and finetunes, take a look at localLlama Reddit threads and just do essentially a sweep of the frontier.

Then it rebuilds latest llama.cpp, grabs the PRs it finds relevant to test against, and then it performs a benchmark and finalizes the upgrade and verifies what model, variant, or even a separate finetune that we should be running.

Occasionally, it performs its own optimizations and commits, which then gets superseded by pull requests and merged code that essentially validates the model's own optimization directionality.

msdz•18m ago
Congratulations on the launch, it looks like an impressive product and tool!

Q: From my (very, very limited!) understanding, I’m under the impression that part of the “inference engine inertia” is model- or at least architecture-specific code for most, if not each new open-weight model coming out. Assuming I got that right, do you plan on supporting everything vLLM/llama.cpp can do, such that Magnitude becomes a drop-in replacement for as many (economically/pareto-viable) models as possible, or do you want to focus on the best possible support for only a select few models/classes of models?

anerli•2m ago
Yeah, generally being able to focus on specific architectures lets you optimize better for those. However models of the same family (for example Qwen 3.5/3.6/ some 3.8 models) share the same architecture, so you only need to optimize once and new models can use the same kernels. There's also shared algorithms and kernels that can be optimized once and used across different families, so it's a bit nuanced.

We plan to support any model architecture that we believe is somewhere along or close to the pareto frontier. There's some model families that are outdated or more niche that we don't necessarily want to put our focus into.

herf•7m ago
I have two NVIDIA GPUs (16GB+16GB) here, and it detects them each twice (says I have 4 GPUs). But then, it says most models are too big (anything >8GB?) and seems to run only on one GPU (5070ti).

Unfortunately even with my 5070ti, llama.cpp seems to be about 20-30% faster at decode, running as:

set CUDA_VISIBLE_DEVICES=0 build\bin\Release\llama-server -hf google/gemma-4-12B-it-qat-q4_0-gguf -ngl 99 --no-mmproj-offload -mg 0 -c 262144 -fa on --host 0.0.0.0

teabee89•5m ago
How does this compare to ZML's llmd https://zml.ai/llmd/ ?