frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Ask HN: What default model do you use and why?

16•stikit•48m ago
I use claude for most of what I do, and Fable is largely overkill for me and frequently burns through my Max plan's session credits in minutes (!) when just doing an initial mobile app planning with 4 agents. After I waited out the timeout period 6 hours later, and I picked up again, the cache had timed out so it burned through 2% of the session in less than a minute. Opus 4.8 is now my goto and I will be avoiding 5 until I see a reason to switch back. 4.8 is 'good enough' for what I need and has been a great value. It mostly gets things right. Most of what I do is web and mobile , largely cloud backend.

Comments

mariocesar•29m ago
I have an 'ask' alias in my shell that just uses Haiku. I use it daily for pretty much everything

For more long work, I now use Fable to create a PLAN.md. I tell it to make a plan that will be executed by other models, and most of the time it ends up choosing Opus or Sonnet.

I didn't start doing this recently. Before that, I would just use the top model for everything. Splitting the work across different models depending on the task has helped a lot. They run faster, and I usually get much better results

sourcecodeplz•9m ago
i can not believe you really use Haiku?
bellowsgulch•28m ago
mimo-v2.5-free, mimo-v2.5, deepseek-flash in that order, honestly don’t even bother using qwen3.6-35b-a3b now unless i need uncensored tasks finished, mostly reverse engineering

most engineering tasks don’t require frontier llms

when they get stuck, then i consider moving up to more capable models

purchasing a claude plan seems widely unnecessary to me

the tasks they do better than the average engineer cut both ways: unless you have an existing portfolio of well written and designed work done pre-llms, it looks like you’re producing slop that pretends to be well designed

poor typography choices despite using the mode,

poor layout choices despite using popular CSS frameworks

etc

bad engineers will always be bad engineers

tools don’t make up for it

edit: a follow up to this— everyone is using eyebrows in their layouts and have no fucking clue why it was done to begin with

everyone has a status pill floating above their front page hero display text and its not fucking status related

so gross

jraedisch•27m ago
opusplan, for the same reasons.
alstonite•26m ago
Astra low for pretty much everything gets me through the work week with a bit to spare on the 5x max plan.
ewindisch•14m ago
How are you managing your tokenomics so well?

Astra low on the $200/mo "20x Pro" plan gets me through a single day.

vallerie•26m ago
I've found good success with the new Gemini models on Antigravity. Granted I use my models either:

- like a fancy auto complete (here are some stub methods, they should do X, fill them in)

- using fairly detailed plans and test harnesses, so blowing up the world is hard

The 3.X Flash family have been fairly capable models, and the selling point for me is just raw speed. Gemini is noticeably faster than the competition, about 3-4x, and I just get work done faster with it.

That said I'm keeping an eye on Open Weights. DS4 Flash was good until price hikes, and finding a provider that serves at high speed and without quantisation at the prior price is tricky.

traverseda•22m ago
Claude, not necessarily because it's better. I find deepseek flash to be of similar quality. But because it's so heavily subsidized.
hmokiguess•18m ago
I think for me stuff peaked around Opus 4.7, I was leaning heavily on the model with paired supervision from reviewing the output manually every step of the way. Ever since that things got a little more complicated and in an unsustainable pace for me, I am trying to remove myself from the equation and build verifiable and reliable tests with quick feedback loops that let frontier models run autonomously but in all honesty not seeing it scale well, I need to take a step back and reassess if the trade off was worthwhile. Frontier models are being incredible at making me feel like they passed my tests only to eventually reveal some tech debt that forces me to take large pivots. It seems the speed is sexy but the results are questionable, models may have hit a limit in my workflow and I think harness engineering is more important than anything. Would love to hear feedback on this take and if others have experienced similar things and what they did to overcome this. (Context would be solo founder bootstrapping greenfield work with full autonomy and sometimes more room for rapid iteration)
herpdyderp•17m ago
- Default: Codex with Terra Max (because it's crazy cheap)

- Preferred: Claude Code with Opus 5 Medium

expedited123•15m ago
Qwen 3.8 flash-next or Gemma 26B because I care about responsible and sustainable usage of LLM's.
philbo•14m ago
Glm-5.3-flash for me. I like faster models because I stay involved all the way through. I don't delegate full control to the agent because it's harder to understand the end result that way.
o_m•13m ago
I used to main Claude, but I can't stand how it writes. I feel like I'm wasting too much time trying to decipher the output. Adding writing rules does not seem to work. Now I use GPT 5.6 Terra high fast mode, with Luna for everything else. I might consider using Sol for planning. I can't stand using Sol or smarter models for coding, because they will eventually try to rewrite everything in the codebase.

I also don't want to use the Claude Code and Codex agent harnesses. The good thing with Codex subscription is that it can be used in other harnesses, unlike Claude. As far as I know, only Anthropic has this restriction.

ewindisch•13m ago
I'm using a lot of gpt-5.6-Luna and glm-5.3-flash. Astra is really fantastic but it's too expensive. I average about 50-90B/tok/mo.
sourcecodeplz•10m ago
i like Luna but only until 130k context, over that it becomes dumb.
sourcecodeplz•12m ago
Muse Spark 1.3 Contribs

unbeatable price/intel ratio per M tokens:

$0.10 (input)

$0.20 (output)

$0.002 (cached-input)

jinnko•10m ago
I was using various open weights models until glm-5.3-flash came out recently. It's incredibly capable and cheap, even if it's very verbose and not the fastest. I've assigned it to all my agents across my harness and it's getting the job done. Still needs a good steer every now and then, but a great work horse.
w22oop•9m ago
I use gemma 4 12B and Qwen 3.8 9B
the__alchemist•8m ago
Sol High-Xhigh, and Opus 5.

Granted, yesterday I threw a few tasks to Astra which the former 2 botches; it produced clean, correct solutions quickly, so pending further eval, this may take over.

IMO unless it's a mechanical tasks, it's worth it to use carefully -crafted queries on the more expensive models, than iterate through messier solutions on the cheaper ones.

codazoda•7m ago
I use Sonnit for my personal work and Opus for my professional work.

For my personal stuff, I'm on a small $20 plan, so I need to use tokens conservatively. I was very rarely exceeding limits until I built a Dark Software Factory. It's not as efficient at token use.

At work I have a $100 plan that I rarely exceed so I use Opus. I have access to Fable too, and I used it a lot while it was new, but I don't find it improves most of my work by too much.

I also use Sol as a secondary for my personal work. I pay for it because I like ChatGPT for various work-loads. I like to let sol write plans for me, it does a better job at certain tasks, and it saves me some Claude tokens.

Show HN:TurboBench:The Data Compression Benchmark. The Compression Lie Detector

https://github.com/powturbo/TurboBench
1•powturbo•1m ago•1 comments

Orcrist: A Coding Agent using LLM state machines

https://github.com/simone20a/Orcrist
1•simone20a•2m ago•0 comments

Show HN: MCP that gives Codex/Claude your SEO and AI visibility data

https://bloomiro.com/mcp
2•johnbuildss•2m ago•0 comments

Lessons I learned as a software engineer

https://www.thetrueengineer.com/p/10-lessons-from-the-journey-what
1•adletbalzhanov•4m ago•0 comments

A Misalignment of AI and Cancer Biology

https://twitter.com/mbeisen/status/2098582265190351334
1•harporoeder•5m ago•0 comments

Ask HN: Can we do a karma reset?

4•HNKarmaReset•5m ago•0 comments

Cities like Raleigh finish construction before e.g. San Francisco approves them

https://twitter.com/esoltas/status/2098162675989565925
1•JumpCrisscross•7m ago•0 comments

The Unravelling of the American Age Is Accelerating

https://phillipspobrien.substack.com/p/the-unravelling-of-the-american-age
1•JumpCrisscross•8m ago•0 comments

Show HN: OpenVurp – An open-source alternative to Grok Bot

https://github.com/openvurp/openvurp
1•vforno•8m ago•0 comments

Why Spotify Is Not Using Bayesian A/B Testing

https://engineering.atspotify.com/2026/9/why-spotify-is-not-using-bayesian-a-b-testing
1•sanj•8m ago•0 comments

Show HN: Redis Lua scripts as typed Python functions

https://github.com/IgnaceMaes/redis-lua-py
1•Nebulic•10m ago•0 comments

LG responds to TV spying allegations

https://www.theverge.com/tech/994333/lg-responds-to-tv-spying-allegations
3•OuterVale•13m ago•0 comments

GrapheneOS: Nearly all of our recent posts on Hacker News are being flagged

https://twitter.com/GrapheneOS/status/2098792171058889006
1•timedude•15m ago•1 comments

Rapidly scaling online storage to serve over 1B ChatGPT users

https://openai.com/index/scaling-storage-one-billion-users-part-one/
1•tamnd•15m ago•0 comments

YoooClaw C·ONE: a Chinese AI card that captures context instead of processing it

https://sinosphereapp.substack.com/p/yoooclaw-cone-the-ai-gadget-that
1•Nyloc1•16m ago•0 comments

Anthropic CEO calls for immediate slowdown in AI development

https://www.axios.com/2026/09/12/anthropic-ai-amodei-pacing
4•doctoboggan•17m ago•1 comments

Why Balaclava?

https://www.espacecagoules.com/en/blogs/tout-sur-la-cagoule/pourquoi-balaclava
1•jruohonen•18m ago•1 comments

Is This a New WhatsApp Hack?

3•MarcellusDrum•18m ago•0 comments

Reddit Now Blocking Firefox for Android with uBlock Origin

3•chucksmash•18m ago•1 comments

Accelerating MiniMax-H3 768p Video Generation on a Nvidia DGX Spark in 1 Minute

https://nvlabs.github.io/Sana/Sol-Engine/Sol-H3-Spark/
2•Lwrless•19m ago•0 comments

I spent $4,000 on a robot dog from China

https://www.understandingai.org/p/i-spent-4000-on-a-robot-dog-from
2•Brajeshwar•20m ago•0 comments

Exploring denser on-chip AI memory with two-transistor gain cells

https://arxiv.org/abs/2602.21278
1•casey2•20m ago•0 comments

Is it time for a Luddite Renaissance?

https://www.npr.org/2026/09/08/nx-s1-5955618/is-it-time-for-a-luddite-renaissance
3•mooreds•22m ago•0 comments

LLMs, Board Layout and Human Engineering

https://gregtidanian.substack.com/p/llms-board-layout-and-human-engineering
1•silustech•23m ago•0 comments

Caroline Ellison Joins Manifund

https://manifund.substack.com/p/caroline-ellison-has-joined-manifund
2•DarkContinent•23m ago•0 comments

What to Expect When You Are Expecting Recursive Self Improvement

https://mark-riedl.medium.com/what-to-expect-when-you-are-expecting-recursive-self-improvement-24...
1•mooreds•23m ago•0 comments

A.I. Was Supposed to Give Us New Killer Apps. What Happened?

https://www.nytimes.com/2026/09/12/opinion/ai-software-coding-apps.html
5•_tk_•24m ago•0 comments

Robert Friedland on the Monumental Shortage of Copper [audio]

https://www.bloomberg.com/news/audio/2026-09-12/robert-friedland-on-the-world-s-monumental-shorta...
1•mooreds•24m ago•0 comments

Scraping barnacles off a cruise ship (the game)

https://barnacleaner.com
1•bryanhan•26m ago•1 comments

Prompts Aren't Real

https://evaluation.club
1•mcfunley•27m ago•0 comments