frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: I built a tool showing how AI providers (should) throttle their models

https://throttle.staffinganalytics.io/?src=hn
4•eliotho•1h ago
OP here: this project was born out of the frustration/paranoia that AI providers are throttling their models when their server load is too high. So, I set out to model and study the problem mathematically to understand what was happening, what I found was quite surprising.

The idea seems natural: as the data center demand increases momentarily through the day, throttling their models (either using quantized versions, reducing the context window or lowering the tier of the model to a smaller one) seems appealing as the replacement model in principle uses less electricity. The problem is that this can cause the opposite effect: as users are trying to solve a question, if the degraded AI model gives a bad answer, the user is likely to keep re-asking. On the AI provider side this looks paradoxical: throttling to a lower model creates in fact more demand for their data center.

This problem is even worse for agentic workflows, as these are more likely to create a re-ask storm, and maybe explains the outages and anecdotal experiences of users that feel the models are degraded.

The model: I used mainly queueing theory arguments solving the optimal scheduling serving for an AI fleet with heterogeneous users solving a finite horizon Dynamic Programming optimization problem.

Insights: The industry standard practice of throttling once the number of users in system exceeds a given threshold is in fact what’s causing the problem, the optimal rule implies separating users that won’t feel degradation as much with users that are very sensitive to it (agents and power users vs users doing simple tasks).

Limitations: The visualization and paper examples are a toy example to illustrate the problem, only the providers have enough data to properly calibrate these instances. In the paper there are some interesting calibrated instances.

Technical Details: The visualization is around 100 lines of flask plus js frontend (LLM assisted with ground truth based on the original numerical example of the paper).

Paper with proofs/theory: https://arxiv.org/abs/2608.23986

The curious symmetry between quantum error correction and the brain

https://arxivblog.substack.com/p/the-curious-symmetry-between-quantum
1•dmonay•34s ago•0 comments

In the product end game, every change carries significant risk, episode 2

https://devblogs.microsoft.com/oldnewthing/20260826-00/?p=112649
1•ibobev•40s ago•0 comments

You can run JavaScript/Python in textlog in a tweet now

https://textlog.cc/tag/exec
1•stagas•1m ago•0 comments

UniFi Travel Router Long-Range

https://store.ui.com/us/en/products/utr-lr
1•HiroProtagonist•1m ago•0 comments

AI Agent Guardrails Don't Work. Safer Platforms Do

https://www.syntasso.io/post/ai-agent-guardrails-don-t-work-safer-platforms-do
1•danielbryantuk•1m ago•0 comments

Open source mobile optimized control pane for coding agent

https://omg.dev/
1•bennykokmusic•2m ago•1 comments

Autistici/Inventati case sets a new counterterrorism precedent, Irdi says

https://decode39.com/16319/autistici-inventati-case-sets-a-new-counterterrorism-precedent-irdi-says/
2•iamnothere•2m ago•0 comments

Get your Windows license refund

https://en.refund4freedom.org/
1•smartmic•3m ago•0 comments

Two Alleged 'TeamPCP' Hackers Arrested in Australia

https://krebsonsecurity.com/2026/08/two-alleged-teampcp-hackers-arrested-in-australia/
1•Errorcod3•3m ago•0 comments

Apple's rare Bay Area layoffs hit Vision Pro and Siri teams amid AI shakeup

https://www.latimes.com/business/story/2026-08-27/apple-lays-off-147-employees
1•1vuio0pswjnm7•3m ago•0 comments

Foiling a Protohackers email spam bot

https://incoherency.co.uk/blog/stories/foiling-protohackers-spam-bot.html
1•surprisetalk•3m ago•0 comments

A Sense of Where You Are (1965)

https://www.newyorker.com/magazine/1965/01/23/a-sense-of-where-you-are
1•jasonshen•3m ago•0 comments

Basic 10Liner Contest 2026

https://gkanold.wixsite.com/homeputerium/rules-2026
1•ethanpil•4m ago•0 comments

Now Hiring: Senior Open Source Maintainer

https://nesbitt.io/2026/08/28/now-hiring-senior-open-source-maintainer.html
1•pimterry•5m ago•0 comments

Cross-Agent Memory

https://lanes.sh/use-cases/shared-memory-across-agents
1•s-xyz•7m ago•0 comments

The Finn – an agent that lives in my router and complains about it

https://github.com/YuriKovalov22/the-finn
1•YuriiKovalov•7m ago•0 comments

Grafana – GeoMap Carto Base layer now requires API key

https://github.com/grafana/grafana/issues/131638
1•yabones•8m ago•0 comments

Google Engineer Accused of Polymarket Insider Trading Says He Was Just Gambling

https://www.wired.com/story/google-engineer-accused-of-polymarket-insider-trading-says-he-was-jus...
2•1vuio0pswjnm7•8m ago•0 comments

Do Rural Residents Have the Power to Say No to Data Centers?

https://barnraisingmedia.com/2026-election-do-rural-residents-have-the-power-to-say-no/
1•iamnothere•8m ago•0 comments

Walmart will accept Apple Pay

https://finance.yahoo.com/personal-finance/banking/article/walmart-will-finally-accept-apple-pay-...
1•mgh2•8m ago•0 comments

Show HN: SkyRoads (1993) – the DOS classic, ported natively to macOS and Linux

https://pedrocatalao.github.io/skyroads-sdl/
1•pedrocatalao•8m ago•0 comments

Do not look inside – you will

1•klopik•8m ago•0 comments

Over 8,300 Gitea servers vulnerable to code execution attacks

https://www.bleepingcomputer.com/news/security/over-8-300-gitea-servers-vulnerable-to-code-execut...
2•speckx•9m ago•0 comments

Ask HN: Why does HN over-advertise useless shit?

3•anslopic•12m ago•1 comments

Htmx 4.0.0 has been released

https://four.htmx.org/announcements/2026-08-28-htmx-4.0.0-is-released
2•rmsaksida•16m ago•0 comments

KHMS – a file-based long-term memory an LLM agent installs into itself

https://github.com/kostey/khms-memory
2•ksxcz•17m ago•0 comments

1931 China Floods

https://en.wikipedia.org/wiki/1931_China_floods
2•sieste•17m ago•0 comments

We need to talk about migrations with AI

https://blog.pragmaticengineer.com/the-pulse-we-need-to-talk-about-migrations-with-ai/
4•_1•17m ago•0 comments

The Loss of Changelogs

https://amxmln.com/blog/2026/the-loss-of-changelogs/
2•speckx•19m ago•0 comments

Short Time Outs

https://www.jefftk.com/p/short-time-outs
2•torutofu•19m ago•0 comments