Ask HN: How are you controlling costs and enforcing limits for LLM calls?

3•8dazo•2d ago

I’ve been running into an issue with LLM/agent systems where unexpected loops or repeated calls can quickly drive up costs.

Most tools I’ve seen focus on observability (logs, traces, dashboards), but not actual enforcement at runtime.

Curious how people here are handling this in production:

- Are you enforcing hard limits (budget, rate, etc.) or just monitoring?

- Do you handle this at the app level or via some middleware/proxy?

- Have you built something in-house for this?

Feels like an unsolved problem, especially with agents.

Would love to hear how others are dealing with it.

Comments

jackycufe•2d ago

Certainly. I use LiteLLM to get more cache and save more money

brandonharwood•1d ago

It’s a bit of a chicken & egg thing and depends a ton on how LLM is applied within an app. I always start at the core design of the integration and focus hard on the problem it solves. Why are you using an LLM in the first place? What is/are the function/s it needs to perform in the context of the user interaction? These are the kind of questions that help you understand the constraints you need to implement. So for example; a project I’m working on is a diagramming tool, and I’m implementing an AI layer on top of it so users can refine/edit/generate diagrams. The tool creates maps structured into a JSON schema, but these can get really long, sometime s thousands of lines depending on the complexity of the diagram. Obviously feeding an entire diagram or having the AI generate an entire diagram is expensive here, so the fix was building a deterministic translation layer that compressed the diagram into a compact semantic model for the LLM, stripping visual noise (x/y coordinates), deduplicating relationships, resolving references etc.

With this and keeping the interact, we cut token usage by ~75% across the app. On the output side, the LLM only produces changes needed, not the full diagram. Layout, validation, and rendering are computed client-side for free so costs only scale with what the user asks for. With good UX as well, we can pay attention to what users ask for, and create “quick actions” that use the LLM within closed loop subsystems. Since we assign a credit system for AI tool usage, we’re better able to accurately assign credit costs to quick actions because each action has a defined scope.

TLDR: make the LLM do less, then put hard limits around the smaller set of things it’s allowed to do

Show HN: I built a Harvey-style tabular review app, then open sourced the code

Microsoft's executive shake-up continues as developer division chief resigns

Debugy: Runtime Logs for Coding Agents

We measured copyrighted-text memorization in 81 open-weight language models

PKG47: AI-Controlled Package Registry

Afterchain – Deterministic inheritance protocol for digital assets

Creating the Futurescape for the Fifth Element

Do links hurt news publishers on Twitter? Our analysis suggests yes

Nigel Farage wants to build a British ICE. Starmer may have handed him the tools

Fast, cheap AI-assisted decompilation of binary code is here

Engineers Are Great for Marketing

Largest Dutch pension fund cuts ties with controversial tech firm Palantir

Cisco: Cybersecurity Remains Top Challenge as Industrial AI Adoption Expands

FalconFly 3dfx Archive

Influence Campaign on TikTok Uses AI Videos to Boost Hungary's Orbán

Reallocating $100/Month Claude Code Spend to Zed and OpenRouter

Škoda's Duobell bicycle bell outsmarts ANC headphones

Content Giant Slashed Telemetry Cost 79%, Saved $1.2M

A study linked various SAT test scores to favorite bands

We Have Become Obsessed with Attachment. And It Is Causing Harm

Some Better Defaults for Emacs

PBXN-110

Ask HN: What is the future of Devs, after launch of Anthropic's Glasswing?

No fine-tuning, no RAG – boosting Claude Code's bioinformatics up to 92%

Opera 130 stable arrives with Chromium 146 and Twitch support

cppreference.com has been under maintenance for a year

Veteran artist behind Mass Effect, Halo, & Overwatch 2 weighs in on Nvidia DLSS5

I was copy-pasting to Claude from WhatsApp – so I fixed that

From bytecode to bytes: automated magic packet generation

Show HN: Giving My First Pitch at 1M Cups Using a Custom Mobile App