frontpage.

Show HN: Morph Reflexes – Multi-head classifiers for agent traces

3•bhaktatejas922•2h ago

The most common failures for production agents are behavioral: looping, reasoning leakage, user frustration, and more. Using a frontier model like GPT or Sonnet to judge every turn is too expensive and slow to run at scale.

What Reflexes are: semantic signals from agent traces, served fast and cheap over API. Built on custom kernels and a custom inference engine forked from vLLM.

Under the hood, it is a small LLM architected around multi-head inference. Small models need to be trained for specific tasks, but running 50 separate small models on the same input for 50 tasks makes no sense.

How it works: We use a modern LLM with hybrid attention and remove the decode step. We built an inference engine that lets prefill compute be 99% reused from reflex to reflex, similar in spirit to older 2019-era BERT/HYDRA and older multiple-head techniques. we built the inference engine to reuse the KV/cache across inputs and compute across all reflexes. One shared backbone reads the trace once, then many heads classify different signals. Our inference engine reuses the same KV/cache and compute across all reflexes, giving us sub-30ms inference with less than 0.1% overhead for each additional reflex.

We took the same high-level idea and did the hard work to make it work with a modern architecture and attention. On it, we can run inference in under 30ms and serve the full request in under 90ms. If you run 4 reflexes or 100, the extra overhead is less than 2ms.

Why does optimizing this matter?

If you’re even a medium-sized startup, you’re dealing with tens of thousands of agent runs and millions of turns. If you want to track things like user frustration rates over time, frontier LLM-as-judge does not scale.

I built a similar stack at Tesla. When ML engineers needed to sample data across petabytes for signals like `is_camera_obfuscated=true`, along with 200 other things, you need to 1) spin them up quickly 2) run at scale efficiently

What it is not: A dashboard. 99% of dashboards go unused. 100% API first and made for devs who want to use this to trigger their own stuff.

vibetrain a custom reflex in our dashboard, and/or then let it self improve in production: https://www.morphllm.com/dashboard/reflex

Docs: https://docs.morphllm.com/sdk/components/reflexes/index

I’d love feedback from people running agents in prod: what sorts of things do you wish you could track over time across 100% of turns but cant right now?

TLDR: semantic signals from agent traces, super fast, cheap via API

Show HN: My 13-year-old built an ant colony tracker

Show HN: Free Online GIS Viewer and Format Converter

Show HN: TakoVM – open-source sandboxing for your agent's code

Show HN: Open-source restreaming and live studio

Show HN: Morph Reflexes – Multi-head classifiers for agent traces

Show HN: Openleetcode – LeetCode runner where tests live in the repo

Show HN: Kage, verification and freshness for Google's OKF agent memory

Show HN: Jensen – a Deus Ex: Human Revolution theme for 30 developer apps

Show HN: Clusy – Cursor for data science notebooks in cloud

Show HN: Shot-scraper video tool for recording YAML-defined webapp feature demos

Show HN: I made a heatmap of 3400 VCs who are open to cold emails

Show HN: Makes local LLMs faster and more reliable by optimizing for your device

Show HN: I built an AI agent to yell at me about my ADHD

Show HN: fenic – LLMs as dataframe operators, query meaning and structure

Show HN: Openleetcode – local LeetCode runner with open test suites

Show HN: Don't ask if devs cheat with AI, test if they're good with it

Show HN: Classic Minesweeper

Show HN: OM Core – multidimensional models without spreadsheet cell formulas

Show HN: Curvytron 2, I rewrote my browser party game, 10 years later

Show HN: Shoaku – Your Coding Navigator

Show HN: Second opinion – A skill to query different models

Show HN: PDFMergely – In-browser PDF tools that never upload your files

Show HN: Agentic Orchestrator, a TUI for long-running coding agents

Show HN: DRM-Free Books

Show HN: TraceAIO – open-source LLM visibility tracker

Show HN: Zanagrams

Show HN: NodePad – AI agent on a canvas instead of a linear chat

Show HN: Running Gemma-4 26B at 124 tokens/SEC on a CPU, no GPU

Show HN: Privacy policy generator for AI apps (LLM disclosure, EU AI Act)

Show HN: Bash4LLM+ – A lightweight, dependency-free Bash wrapper for LLM APIs

Show HN: Morph Reflexes – Multi-head classifiers for agent traces

Show HN: My 13-year-old built an ant colony tracker

Show HN: Free Online GIS Viewer and Format Converter

Show HN: TakoVM – open-source sandboxing for your agent's code

Show HN: Open-source restreaming and live studio

Show HN: Morph Reflexes – Multi-head classifiers for agent traces

Show HN: Openleetcode – LeetCode runner where tests live in the repo

Show HN: Kage, verification and freshness for Google's OKF agent memory

Show HN: Jensen – a Deus Ex: Human Revolution theme for 30 developer apps

Show HN: Clusy – Cursor for data science notebooks in cloud

Show HN: Shot-scraper video tool for recording YAML-defined webapp feature demos

Show HN: I made a heatmap of 3400 VCs who are open to cold emails

Show HN: Makes local LLMs faster and more reliable by optimizing for your device

Show HN: I built an AI agent to yell at me about my ADHD

Show HN: fenic – LLMs as dataframe operators, query meaning and structure

Show HN: Openleetcode – local LeetCode runner with open test suites

Show HN: Don't ask if devs cheat with AI, test if they're good with it

Show HN: Classic Minesweeper

Show HN: OM Core – multidimensional models without spreadsheet cell formulas

Show HN: Curvytron 2, I rewrote my browser party game, 10 years later

Show HN: Shoaku – Your Coding Navigator

Show HN: Second opinion – A skill to query different models

Show HN: PDFMergely – In-browser PDF tools that never upload your files

Show HN: Agentic Orchestrator, a TUI for long-running coding agents

Show HN: DRM-Free Books

Show HN: TraceAIO – open-source LLM visibility tracker

Show HN: Zanagrams

Show HN: NodePad – AI agent on a canvas instead of a linear chat

Show HN: Running Gemma-4 26B at 124 tokens/SEC on a CPU, no GPU

Show HN: Privacy policy generator for AI apps (LLM disclosure, EU AI Act)

Show HN: Bash4LLM+ – A lightweight, dependency-free Bash wrapper for LLM APIs