frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: InstinctFlash – Run 5B world-action models in real time on Jetson Thor

https://github.com/General-Instinct/InstinctFlash
10•guanming0717•3h ago
Hey HN, Guanming here, cofounder of General Instinct. We just released InstinctFlash, a high-performance serving framework for robotics models. It’s licensed under AGPL-3.0.

On Jetson Thor, we see speedups about 1.2x to 7.9x from runtime optimizations alone and up to 33.78x for LingBot-VA when we combine those runtime optimizations with a distilled few-step diffusion scheduler, going from the original 25 visual / 50 action steps to 2 / 4 steps. Across 50 Robotwin2.0 tasks, we evaluated 1,153 episodes per configuration, LingBot-VA with InstinctFlash at 2 visual / 4 action steps achieved a 90.5% success rate, compared with 92.1% for the baseline at 25 visual / 50 action steps.

Here’s an optimized 5B world action model, running in real time on a Jetson Thor: https://youtu.be/nku65iyL5Fw

InstinctFlash currently supports 8 VLA / world-action model families, including pi0.5 and NVIDIA Cosmos Policy, across RTX 4090 / 5090 and Jetson Thor.

Just give it your fine-tuned checkpoint and InstinctFlash handles the rest, exposing the accelerated model through a Python runtime or an OpenPI-compatible WebSocket server.

We started working on this because we kept running into the same problem while deploying robot policies, the models were getting much better, but inference was often way too slow for the control loop we actually wanted.

For pi0.5, mixed-precision GEMMs and CUDA graphs speed up computation and reduce launch overhead. For Cosmos, caching avoids redundant computation across diffusion steps. World-action models’ diffusion denoising step depends on the previous one which motivated our work on few-step distillation.

Right now, InstinctFlash contains 6 aspects of optimization.

- Graph: CUDA graph capture, memory planning and separating prefill from repeated execution.

- Cache: Reusing KV and conditioning state across diffusion steps and prediction calls.

- Attention: Specialized attention paths for different model architectures.

- Kernels: Fused operations and kernels tailored to specific backends and tensor layouts.

- Precision: FP8 and mixed-precision execution.

- Model: Few-step distillation for diffusion and action generation.

Teams at Samsung, Siemens, and other robotics startups have used InstinctFlash for model acceleration on VLAs, WAMs, and diffusion-based world models. Now we are opening up access to you.

Feel free to try it here: https://github.com/General-Instinct/InstinctFlash

More implementation details and benchmarks: https://general-instinct.com/blog/instinctflash-edge-inferen...

Would love to hear your feedback!

Show HN: Drop – A rootless Linux sandbox with gVisor support

https://droprun.sh/
119•mixedbit•4h ago•39 comments

Show HN: AI·rete·RAG – a Rete rule engine decides, RAG explains why

https://ai-rete-rag.com/
19•ZaharaHussain•2h ago•0 comments

Show HN: InstinctFlash – Run 5B world-action models in real time on Jetson Thor

https://github.com/General-Instinct/InstinctFlash
10•guanming0717•3h ago•0 comments

Show HN: graf (1000x faster graphify in Rust)

https://github.com/ctxrs/graf
3•ripped_britches•1h ago•0 comments

Show HN: OpenMCP – An open, code-first fork of the Model Context Protocol

https://github.com/enclawed/omcp
2•enclawed•3h ago•0 comments

Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM

https://github.com/volotat/mini-AGI/
269•volotat•1d ago•70 comments

Show HN: Venya lets AI agents use secrets without seeing them

https://github.com/tabith-llc/venya
2•tabith•4h ago•0 comments

Show HN: Foremerge – Catch intent conflicts between parallel coding agents

https://github.com/naw103/foremerge
45•naw103•1d ago•15 comments

Show HN: Matcha – fresh job postings pulled directly from company career pages

https://matchajobs.co/
2•jpercy•6h ago•1 comments

Show HN: A competition for small neural networks that play strategy games

https://tinybrains.dev
108•codetiger•2d ago•40 comments

Show HN: Lossless-memory – a personal AI memory that never summarizes

https://github.com/aru-labs/lossless-memory
64•aru-labs•1d ago•29 comments

Show HN: Radius – A Meetup.com Alternative

https://radius.to/
162•radius89•2d ago•77 comments

Show HN: Watch all the AI agents on your machine

https://github.com/markwylde/all-your-agents
5•turblety•7h ago•1 comments

Show HN: An atlas of system designs with interactive architecture diagrams

https://atlas-sysdes.vercel.app/
3•mertkahyaoglu•7h ago•2 comments

Show HN: Vellum, the best diagram editor you'll ever use

https://vellum.blueprintr.io
20•JoshJamesAnthon•1d ago•15 comments

Show HN: jevals – replacing LLM judges with typed Jev decisions

https://github.com/openlayer-ai/jevals
43•gbayomi•1d ago•6 comments

Show HN: Sigabrt.dev – cronjob monitor with an SSH TUI

https://sigabrt.dev
81•4815162342•3d ago•35 comments

Show HN: CUA-S1 – A System One Model for Computer Use

https://github.com/trycua/cua
90•frabonacci•3d ago•10 comments

Show HN: Gdocs-me-up: a high-fidelity Google Docs exporter

https://github.com/behdad/gdocs-me-up
20•behdad•1d ago•8 comments

Show HN: Share your AI Setup, Learn from others

https://mysetup.ai/
249•steveybrown•5d ago•138 comments

Show HN: Jev Powered Obsidian Search

https://github.com/Emlembow/jev-graph-search
4•MikeLembo•18h ago•1 comments

Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash

https://cactuscompute.com/needle
236•HenryNdubuaku•4d ago•92 comments

Show HN: Lightspeed, build real time apps the Laravel way (+asteroids demo)

17•jv22222•1d ago•5 comments

Show HN: Snapdrop: Instantly share files between devices. No setup, no signup

https://snapdrop.me
111•Capira•4d ago•51 comments

Show HN: WebGCM – a global climate model running in the browser on WebGPU

https://gcm.echorelay.net/
5•jlhawn•21h ago•0 comments

Show HN: Blackgit – use Git to only download those cared files

https://github.com/zhuzhonghua/blackgit
3•zhonghua•22h ago•2 comments

Show HN: Scry, programmable internet search w/ congestion pricing

https://scry.io/
60•Xyra•4d ago•26 comments

Show HN: Differential – a TUI diff reviewer with semantic grouping

https://github.com/thepartly/differential
3•gogoout•22h ago•1 comments

Show HN: A GUI for non programmers for Epic's version control system Lore

https://www.anchorpoint.app/lore
6•m_niedoba•1d ago•1 comments

Show HN: Ambits – agentic grep/rg tool will history tracking

https://github.com/joshLong145/ambits
5•joshLong145•1d ago•0 comments