frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Fable 5 – Median thinking declined in August

https://twitter.com/Lon/status/2101793422487204027
30•espeed•46m ago

Comments

alexjplant•22m ago
I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).

I wonder what their official explanation for this behavior is.

QwenGlazer9000•11m ago
Last time they were called out, it was a regression in Claude code itself.

At least that's their explanation. Either way, it wasn't a good look for "vibecoding" but it got brushed over.

CamperBob2•19m ago
How do you measure thinking tokens? They don't send those back to the client.
ivanbakel•13m ago
They tell you how many tokens are used, however, right? Otherwise you couldn't see your own token consumption.
CamperBob2•12m ago
Good point. I suppose watching the number go up is useful information in itself.

I have been using CC with DeepSeek 4.1 Flash lately, and it's nice to see how the sausage is being made (even if it's partly illusory, as CoT always is.)

r2-129•14m ago
Obviously. The standard pattern is that model X is basically AGI and wins all benchmarks, followed the next day by Y and Z, which both win all benchmarks, too.

Then weeks later people find out that they have been duped and complain that the models have been quantized or employ worse inference.

Buy decent coffee instead of your $200 subscription and sidestep all the scams.

CamperBob2•6m ago
You forgot a stage or two:

1: "Our model will bring about the end of all things. Flee, flee for your lives"

2: "Our model is basically AGI"

3: "Our model will be available in limited release next week"

4: "Everybody who subscribes at the $200 level gets access now"

5: "Everybody who subscribes at the $20 level gets access now"

6, at least at Google: "Our model will be shoved down your throat every time you do a search, whether you want it or not"

ltbarcly3•6m ago
Well I happen to enjoy coffee and $200 AI plans. What if Blue Bottle started watering down it's coffee? Is your answer to stop drinking coffee and make myself tea instead?

Information that vendors are watering down or otherwise being misleading in what they are delivering in their product is important to share even if you don't use that product yourself.

mlmonkey•9m ago
Anecdotally, I have found the same. I spend a lot of time with these frontier models, brainstorming, etc. and the drop in performance from, say, week 1 to week 8 is often massive. Whereas in the beginning, it seemed like a capable research assistant, by the end of week 8 or so it starts acting like a puppy dog eager to make its 'master' happy for a few treats.
cloudking•7m ago
How do create repeatable tests in a non-deterministic system? Every time you send the same prompt you get a different answer.
ssivark•5m ago
[delayed]

I don't care about majors, minors, or honors programs. Do you?

https://blog.computationalcomplexity.org/2026/09/i-dont-care-about-majors-minors-or.html
1•mathattack•27s ago•0 comments

Ask HN: How close are we to AI bubble bursting?

1•bg117•36s ago•0 comments

The Questionable Legality of ICE's Palantir Elite System

https://www.techpolicy.press/from-records-to-raids-the-questionable-legality-of-ices-palantir-eli...
1•cdrnsf•46s ago•0 comments

The Influence of AI on Human Decisions in DHS Surveillance

https://www.techpolicy.press/the-influence-of-ai-on-human-decisions-in-dhs-surveillance/
1•cdrnsf•1m ago•0 comments

What Jev Means for the Future of Evals

https://armank.com/thoughts/3
1•armank-dev•3m ago•0 comments

Hands-On with Googlebooks

https://www.tomshardware.com/laptops/hands-on-with-googlebooks-five-models-the-new-googlebook-os-...
1•Shuddown•4m ago•0 comments

North Korea Stopped Nuclear Testing in 2017, Triggered 1k Earthquakes Since

https://www.sciencealert.com/north-korea-stopped-nuclear-testing-in-2017-but-triggered-nearly-140...
1•gmays•4m ago•0 comments

The Sun Doesn't Shine on Me (2006)

https://fullhoffman.com/2006/03/20/the-sun-doesnt-shine-on-me/
1•simonebrunozzi•9m ago•0 comments

BBC removes 11 Mitchell and Webb comedy sketches from iPlayer

https://www.bbc.co.uk/news/articles/c6jdvpr30k8jo
1•benj111•9m ago•0 comments

You're a Meat Proxy

https://twitter.com/i/status/2101779116056088732
1•Michelangelo11•10m ago•0 comments

Antigravity Spaces – Switch Antigravity IDE Projects Easily

https://github.com/ApollosWave/antigravity-spaces
1•apolloswave•11m ago•0 comments

Iron willpower is an illusion – composure is the real engine of self-control

https://drdeborst.substack.com/p/iron-willpower-is-an-illusion-composure
2•dr-j•11m ago•0 comments

Show HN: Jobs at Recently Funded Startups

https://www.vcbacked.co/jobs
1•veritas9•12m ago•0 comments

Chess Design Timeline

https://chess-timeline.vercel.app/
2•msotomorras•13m ago•1 comments

Database Normalization

https://en.wikipedia.org/wiki/Database_normalization
1•gregsadetsky•15m ago•0 comments

We Asked a Urologist Whether Icing Your Testicles Boosts Testosterone

https://geiselmed.dartmouth.edu/news/2026/we-asked-a-urologist-whether-icing-your-testicles-reall...
1•harry_nutsachs•17m ago•0 comments

Guess the Country from Demographic Data

https://www.demoguessr.com/
1•BjarkeErichsen•19m ago•1 comments

Show HN: TickerWhale – nightly 0-100 scores and fair values for 1,300 stocks

https://tickerwhale.com
1•nacer222•21m ago•0 comments

Googlebook – Meet the Lineup and pre-order

https://googlebook.google/shop/
1•bretpiatt•23m ago•1 comments

Grok 4.7 Intelligence, Performance and Price Analysis

https://artificialanalysis.ai/models/grok-4-7
4•theanonymousone•23m ago•0 comments

Show HN: Agent Chaperone – Screen AI agent tool calls and results with Jev

https://github.com/agent-chaperone/agent-chaperone
2•sepehrsafari•25m ago•0 comments

Show HN: flyOS – A fruit fly connectome simulated in real time on an iPhone

https://www.becomethefly.com
2•coryetzkorn•25m ago•0 comments

Saudi Arabia's Ceer launches flagship electric vehicles

https://www.agbi.com/manufacturing/2026/09/saudi-arabias-ceer-launches-flagship-electric-vehicles/
4•abdullahalharir•27m ago•0 comments

Modulate ML Team Announces New Public Entity Transcription Benchmark

https://www.modulate.ai/blog/modulate-ml-team-announces-new-public-entity-transcription-benchmark
1•voicevector5•28m ago•0 comments

WTF Is Up with Napster's AI Pivot?

https://tedium.co/2026/09/21/napster-ai-pivot/
2•speckx•28m ago•0 comments

Alcor: Simulate cpuid and sgdt/sidt results per-process

https://github.com/er-azh/alcor
2•Tiberium•29m ago•0 comments

Google hit with €403M fine by Irish data watchdog over GDPR violations

https://www.bbc.co.uk/news/articles/ck1e52v16ngxo
2•philbo•29m ago•0 comments

Raspberry Pi founder Eben Upton: 'I'm an Omni-geek

https://www.ft.com/content/2285c111-b103-4b04-96c7-fe26a3c04c3e
1•imichael•30m ago•0 comments

Grok 4.7 is here with Electrical engineering benchmark which beats fable 5.1 max

https://twitter.com/hive_echo/status/2102072241420898708
1•echohive42•31m ago•0 comments

Delta A21N at Kahului on Sep 19th 2026, fuel fumes on board

https://avherald.com/h?article=541628f4
2•r2sk5t•34m ago•0 comments