frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: Don't trust me bro: fixing GPT-OSS (3.49B tokens, 1k GPU hours, 1x3090)

https://github.com/iamskeole/burrito-core
2•iamskeole•49m ago
Long story short, about a year ago, in spite of everybody bashing gpt-oss for broken tool calling and refusals, I thought there's something there worth exploring. Model hit a sweet spot for me in that it was the first time I could run full 128k context, factory-precision weights, across parallel requests on a single RTX 3090 at close to 200 tps (well... eventually, but it was still flying at around 100 tps initially which was mind blowing in the before-times).

Could and would being two different things, turned out both llama.cpp and vLLM were shitting their pants running the model at the time (love you guys, I know this model was a pita!), particularly around tool calling (vLLM was / is broken seven ways to Sunday), mostly due to the Harmony template introduced by OpenAI (which, coincidentally (?) is almost identically implemented in Gemma 4 and somehwat similar in Muse Glimmer, 9-12 months after the gpt-oss release, so OpenAI was on to something there and likely not just for the OSS release but their bigger and closed siblings too).

Anyway, validating my hypothesis with the vanilla backends proved impossible at the time.

So I did the only rational thing: built an inference harness that fixes the model, then ran probably the most autistic evals in history -- 320,192 questions across 8 seeds, prefilling and decoding over 3.49B tokens, for 1,062 hours of batch size 1 GPU time on a single 3090.

In the words of Carl Sagan, to make an apple pie from scratch, you first have to invent the universe. I spent my nights inventing this one in parking lots between food delivery gigs, so I named it burrito.

All that just to test whether OpenAI shipped a broken model (spoiler: it didn't). Did it work? Here's the hero shots for the final boss of tool calling evals: multi-turn, pass@8 (at least 1 seed of 8) and pass^8 (every seed).

https://raw.githubusercontent.com/iamskeole/burrito-evals/re... > task solve rate on at least one seed

https://raw.githubusercontent.com/iamskeole/burrito-evals/re... > task solve rate on every seed

Sharing everything, MIT:

- harness: https://github.com/iamskeole/burrito-core - evals (incl. full inference traces): https://github.com/iamskeole/burrito-evals - fixed jinja template: https://huggingface.co/openai/gpt-oss-20b/discussions/274/fi...

(Detailed analysis on reasoning zones and "optimal effort" levels can be found in the evals repo)

Website Deathmatch: head-to-head website battles

https://website-deathmatch.com
1•jargonbank•32s ago•0 comments

Morning bathrobe rant: AI slop. – Uncle Bob

https://twitter.com/unclebobmartin/status/2094742094187307295
1•dayyan•38s ago•0 comments

Human Control at Machine Speed

https://cacm.acm.org/blogcacm/human-control-at-machine-speed/
1•franciscoromeir•3m ago•0 comments

Sennett (2016) Programming as Craftsmanship or Why Win10 was so bad

https://www.youtube.com/watch?v=nIq4w9brxTk
1•lensecat•3m ago•1 comments

A new DARPA Lift Challenge just got announced

https://www.electronicdesign.com/blogs/nonlinearities/article/55402155/electronic-design-darpa-an...
1•andyturudic•3m ago•0 comments

PRs Not Welcome: How Top AI OSS Projects Are Managing Contributors

https://www.latent.space/p/pr-not-welcome
1•shenli3514•3m ago•0 comments

Kedrosky: Disconnect between AI valuations and revenue-growth forecasts

https://www.marketwatch.com/story/theres-a-disconnect-between-ai-valuations-and-revenue-growth-fo...
1•aanet•6m ago•1 comments

Bricolage Grotesque

https://ateliertriay.github.io/bricolage/
1•saikatsg•6m ago•0 comments

Anthropic admits AI 'not perfectly aligned' with human values

https://www.theguardian.com/technology/2026/sep/01/anthropic-claude-ai-hacking-human-values
2•fittingopposite•7m ago•0 comments

Nepal seeks compensation from China, US, India, other large polluters for floods

https://theprint.in/diplomacy/not-charity-but-moral-liability-nepal-seeks-compensation-from-china...
2•alephnerd•8m ago•0 comments

Google Pics: Easy image creation and editing in Google Workspace

https://blog.google/products-and-platforms/products/workspace/google-pics/
1•mikexstudios•8m ago•0 comments

Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face

https://www.dwarkesh.com/p/ajeya-cotra
1•swolpers•8m ago•0 comments

Germany says Russia behind Leipzig airport drone attack

https://www.bbc.com/news/articles/c5ylm3m67n2o
2•jonnybgood•9m ago•1 comments

Agentic Research Is Oxymoronic

https://arxiv.org/abs/2608.31161
2•privong•11m ago•0 comments

The traditional research paper is dead

https://generativehistory.substack.com/p/no-more-muddling-through
2•tolerance•12m ago•0 comments

Defining AI Psychosis. Part 1: True AI Psychosis

https://jeffs.blog/p/defining-ai-psychosis-part-1-true
1•euthymiclabs•12m ago•0 comments

Cleaning Up Poop: The Human Labor Behind Waymo

https://bayareacurrent.com/theres-a-lot-of-cleaning-up-poop-a-depot-worker-on-the-human-labor-beh...
2•counteroptimize•12m ago•0 comments

GCP us-central1 having major issues

https://status.cloud.google.com/incidents/J5ia5t9p3g9Q5Wi7r8Ev
5•sadfsdfsadfsdf•12m ago•0 comments

From 3:00 Am Panic to Confidence: How I Use AI During On-Call Incidents

https://qainsights.com/from-300-am-panic-to-confidence-how-i-use-ai-during-on-call-incidents/
1•qainsights•12m ago•0 comments

Skforecast-AI – Agentic time series forecasting in Python

https://github.com/skforecast/skforecast-ai
1•JoaquinAmat•13m ago•0 comments

HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions

https://thezvi.substack.com/p/huggingface-attack-postmortem-civilizations
1•swolpers•14m ago•0 comments

Fal scales media search to 300M+ assets on turbopuffer

https://twitter.com/turbopuffer/status/2094808269021970671
1•joshauk•16m ago•0 comments

CMS Proposes Changes to Incentivize FHIR-Enabled Electronic Prior Authorization

https://www.wsgr.com/en/insights/cms-proposes-changes-to-incentivize-fhir-enabled-electronic-prio...
1•stmw•16m ago•0 comments

Debugging API Responses Locally: A JSON Formatter That Never Sends Your Payload

https://capytoolkit.com/blog/developer-tools/debugging-api-responses-locally-browser-json-formatt...
1•ChillyCapy•17m ago•0 comments

USPS whistleblower describes new plans for voting by mail hidden from the public

https://www.cnn.com/2026/09/01/politics/usps-whistleblower-mail-ballot-rules-trump
4•GrinningFool•17m ago•1 comments

Ajeya Cotra – The OpenAI/Hugging Face story, told by one of the investigators [video]

https://www.youtube.com/watch?v=X50zezLFWWI
1•tosh•17m ago•0 comments

Maniquest Launch

https://maniquest.com
1•maniquest•18m ago•1 comments

'Humongous' marmot crowned winner in inaugural Fat Marmot Week contest

https://www.theguardian.com/us-news/2026/aug/30/marmot-winner-fat-marmot-week-contest
1•ohjeez•18m ago•0 comments

A Million Kakapos

https://blog.mempko.com/a-million-kakapos/
1•mempko•18m ago•0 comments

DEF CON 34 – ESP32 as counter-surveillance platform [video]

https://www.youtube.com/watch?v=PcK-TDmzshc
1•skibz•18m ago•0 comments