frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Nvidia is the central bank of AI

https://www.economist.com/interactive/briefing/2026/09/03/nvidia-is-the-central-bank-of-ai
72•tolugenius•1h ago•45 comments

A Mathematical Framework for Transformer Circuits (2021)

https://transformer-circuits.pub/2021/framework/index.html
36•Bluestein•2h ago•3 comments

IKEA made a mod for Skyrim [video]

https://www.youtube.com/watch?v=iZODN0QUgjI
414•kegenaar•2d ago•88 comments

Fuck it, make it anyway

https://www.joelotter.com/posts/2026/09/make-it-anyway/
377•JayOtter•4h ago•304 comments

Retrospectively Reverse-Engineering Apple's Neural Engine

https://eiln.github.io/posts/ane.html
174•zdw•8h ago•18 comments

A misalignment of AI in mathematics

https://mathandai.org/
1103•meredydd•22h ago•1047 comments

The Worst Spam Emails: Inside iLands' AI Agent Hustle

https://tedium.co/2026/09/11/ilands-agents-email-spam-kaixin-tang/
68•ColinWright•4h ago•38 comments

I spent $220 on Google app ads and 60% of the installs were robots

https://dayzlegame.com/blog/google-ads-bot-farm/
664•nickabe•21h ago•366 comments

LG responds to TV spying allegations

https://www.theverge.com/tech/994333/lg-responds-to-tv-spying-allegations
19•OuterVale•39m ago•8 comments

LRU is harder to beat than the KV-cache papers suggest

https://github.com/gauravapiscean/agentic-kv-cache
46•gauravapiscean•2d ago•24 comments

Performance of WebAssembly Runtimes in 2026

https://00f.net/2026/06/23/webassembly-runtimes-2026/
32•fagnerbrack•3d ago•6 comments

A Design Space Exploration of Async/Await

https://cel.cs.brown.edu/blog/design-space-async-await/
373•wcrichton•3d ago•101 comments

Moby-Dick and the Indefinite Sublime: On facing the white page

https://yalereview.org/article/tom-mccarthy-indefinite-sublime
11•benbreen•3d ago•0 comments

Great Lakes sturgeon may be 400 years old:Scientists rethinking how to save them

https://www.cbc.ca/news/canada/ontario-great-lakes-sturgeon-lifespan-study-9.7329250
116•bookofjoe•3d ago•31 comments

Europe's "Less" Is Doing More Than Anyone Gives It Credit For

https://oilprice.com/Energy/Energy-General/Europes-Less-Is-Doing-More-Than-Anyone-Gives-It-Credit...
24•u1hcw9nx•1h ago•3 comments

Usenet rewind archive search engine

https://www.usenet-rewind.com/
108•cstadler1869•11h ago•33 comments

Show HN: Bodily Oddities

https://vester.si/bodily-oddities/
307•vesterde•1d ago•190 comments

Forgotten Woodlands

https://storymaps.arcgis.com/stories/9b790daf22ba4e87836f467abb1c7e49
30•NaOH•18h ago•9 comments

Logo Programming

https://el.media.mit.edu/logo-foundation/what_is_logo/logo_programming.html
298•azhenley•3d ago•120 comments

Inverse Kinematics and Foot Locking

https://theorangeduck.com/page/inverse-kinematics-foot-locking
103•airhangerf15•5d ago•11 comments

google.com/goto: Google's anti-scraping update

https://www.autom.dev/blog/google-search-goto-links
529•1e1a•12h ago•431 comments

Is it time for a Luddite Renaissance?

https://www.npr.org/2026/09/08/nx-s1-5955618/is-it-time-for-a-luddite-renaissance
9•mooreds•48m ago•1 comments

We've followed their lives for six decades; now the stars of 7 Up are bowing out

https://www.bbc.co.uk/news/articles/crm932el3yjo
86•mellosouls•5h ago•22 comments

Compiler Can Undo Your Security Checks

https://davidbombal.com/your-compiler-can-undo-your-security-checks/
14•birdculture•2h ago•21 comments

OpenAI agents carried out an undisclosed attack on RubyGems

https://www.rubyhack.ai/
861•chao-•16h ago•504 comments

Designing for Dual Screen and Foldable Devices with CSS (2023)

https://blog.stephaniestimac.com/posts/2023/05/design-foldable-devices/
52•mooreds•2d ago•11 comments

Mind-altering drugs played key role in rise of Andean civilization

https://www.science.org/content/article/mind-altering-drugs-played-key-role-rise-andean-civilization
198•geneticdrifts•22h ago•131 comments

Litelm: LiteLLM Without the Bloat

https://github.com/kennethwolters/litelm
164•kennethwolters•22h ago•57 comments

We Must Pace the Frontier

https://darioamodei.com/post/we-must-pace-the-frontier
143•apsec112•2h ago•178 comments

SystemIO conflicts are not firmware bugs

https://codon.org.uk/~mjg59/blog/p/systemio-conflicts-are-not-firmware-bugs/
26•haeseong•2d ago•5 comments
Open in hackernews

Ask HN: What default model do you use and why?

25•stikit•1h ago
I use claude for most of what I do, and Fable is largely overkill for me and frequently burns through my Max plan's session credits in minutes (!) when just doing an initial mobile app planning with 4 agents. After I waited out the timeout period 6 hours later, and I picked up again, the cache had timed out so it burned through 2% of the session in less than a minute. Opus 4.8 is now my goto and I will be avoiding 5 until I see a reason to switch back. 4.8 is 'good enough' for what I need and has been a great value. It mostly gets things right. Most of what I do is web and mobile , largely cloud backend.

Comments

mariocesar•54m ago
I have an 'ask' alias in my shell that just uses Haiku. I use it daily for pretty much everything

For more long work, I now use Fable to create a PLAN.md. I tell it to make a plan that will be executed by other models, and most of the time it ends up choosing Opus or Sonnet.

I didn't start doing this recently. Before that, I would just use the top model for everything. Splitting the work across different models depending on the task has helped a lot. They run faster, and I usually get much better results

sourcecodeplz•34m ago
i can not believe you really use Haiku?
mariocesar•28m ago
To ask for web search, simple commands and research, or passing logs to analyze it's more than enough

Here is the script https://github.com/mariocesar/dotfiles/blob/main/common/.loc...

notnmeyer•20m ago
this is a really neat idea. thanks for sharing!
bellowsgulch•53m ago
mimo-v2.5-free, mimo-v2.5, deepseek-flash in that order, honestly don’t even bother using qwen3.6-35b-a3b now unless i need uncensored tasks finished, mostly reverse engineering

most engineering tasks don’t require frontier llms

when they get stuck, then i consider moving up to more capable models

purchasing a claude plan seems widely unnecessary to me

the tasks they do better than the average engineer cut both ways: unless you have an existing portfolio of well written and designed work done pre-llms, it looks like you’re producing slop that pretends to be well designed

poor typography choices despite using the mode,

poor layout choices despite using popular CSS frameworks

etc

bad engineers will always be bad engineers

tools don’t make up for it

edit: a follow up to this— everyone is using eyebrows in their layouts and have no fucking clue why it was done to begin with

everyone has a status pill floating above their front page hero display text and its not fucking status related

so gross

admiralrohan•25m ago
Where are you getting mimo-v2.5 for free?
jraedisch•53m ago
opusplan, for the same reasons.
alstonite•51m ago
Astra low for pretty much everything gets me through the work week with a bit to spare on the 5x max plan.
ewindisch•40m ago
How are you managing your tokenomics so well?

Astra low on the $200/mo "20x Pro" plan gets me through a single day.

jbonatakis•28m ago
I honestly don’t think I could blow through a 20x plan in a day if I was trying. What are you doing that uses so many tokens?
fibonacci112358•3m ago
It's really strange, you see people claim the same on the Codes subreddit. I can also run Astra xHigh at least for a few days with my 20x, but I don't use any of the token hungry patterns like orchestrators and such (only the code reviews have a skill using either 4 or 8 sub agents, some being Terra/Luna and none Astra). I did have a month ago a period or 2-3 days when usage was draining like 5x faster than usual, maybe some people are seeing smth like that? But more likely many sub agents and other patterns eating tokens.
vallerie•51m ago
I've found good success with the new Gemini models on Antigravity. Granted I use my models either:

- like a fancy auto complete (here are some stub methods, they should do X, fill them in)

- using fairly detailed plans and test harnesses, so blowing up the world is hard

The 3.X Flash family have been fairly capable models, and the selling point for me is just raw speed. Gemini is noticeably faster than the competition, about 3-4x, and I just get work done faster with it.

That said I'm keeping an eye on Open Weights. DS4 Flash was good until price hikes, and finding a provider that serves at high speed and without quantisation at the prior price is tricky.

martythemaniak•22m ago
Same. I have a promo plan that's currently $3/month, which gets you a model that's almost opus, very fast speeds, and very generous limits.

There's no Gemini pro model currently, so you gotta pair that with a 20 openai plan for access to more advanced stuff if you need it.

traverseda•47m ago
Claude, not necessarily because it's better. I find deepseek flash to be of similar quality. But because it's so heavily subsidized.
hmokiguess•43m ago
I think for me stuff peaked around Opus 4.7, I was leaning heavily on the model with paired supervision from reviewing the output manually every step of the way. Ever since that things got a little more complicated and in an unsustainable pace for me, I am trying to remove myself from the equation and build verifiable and reliable tests with quick feedback loops that let frontier models run autonomously but in all honesty not seeing it scale well, I need to take a step back and reassess if the trade off was worthwhile. Frontier models are being incredible at making me feel like they passed my tests only to eventually reveal some tech debt that forces me to take large pivots. It seems the speed is sexy but the results are questionable, models may have hit a limit in my workflow and I think harness engineering is more important than anything. Would love to hear feedback on this take and if others have experienced similar things and what they did to overcome this. (Context would be solo founder bootstrapping greenfield work with full autonomy and sometimes more room for rapid iteration)
herpdyderp•42m ago
- Default: Codex with Terra Max (because it's crazy cheap)

- Preferred: Claude Code with Opus 5 Medium

expedited123•40m ago
Qwen 3.8 flash-next or Gemma 26B because I care about responsible and sustainable usage of LLM's.
philbo•39m ago
Glm-5.3-flash for me. I like faster models because I stay involved all the way through. I don't delegate full control to the agent because it's harder to understand the end result that way.
o_m•38m ago
I used to main Claude, but I can't stand how it writes. I feel like I'm wasting too much time trying to decipher the output. Adding writing rules does not seem to work. Now I use GPT 5.6 Terra high fast mode, with Luna for everything else. I might consider using Sol for planning. I can't stand using Sol or smarter models for coding, because they will eventually try to rewrite everything in the codebase.

I also don't want to use the Claude Code and Codex agent harnesses. The good thing with Codex subscription is that it can be used in other harnesses, unlike Claude. As far as I know, only Anthropic has this restriction.

r_lee•13m ago
I recently switched over to GPT from Claude once my subscription lapsed, it's too early to tell, but I do appreciate how more straight to the point Sol 5.6 seems to be, vs. the insane wordslop machine Opus tends to be.

It's really strange because when Opus 5 released, there were some that pointed this out, but a bunch simply said it was the best and as good as Fable etc etc.

but for me, it caused me to get very demotivated and avoid interacting with the model, at least when using Claude Code.

wenc•3m ago
[delayed]
ewindisch•38m ago
I'm using a lot of gpt-5.6-Luna and glm-5.3-flash. Astra is really fantastic but it's too expensive. I average about 50-90B/tok/mo.
sourcecodeplz•35m ago
i like Luna but only until 130k context, over that it becomes dumb.
sourcecodeplz•37m ago
Muse Spark 1.3 Contribs

unbeatable price/intel ratio per M tokens:

$0.10 (input)

$0.20 (output)

$0.002 (cached-input)

jinnko•36m ago
I was using various open weights models until glm-5.3-flash came out recently. It's incredibly capable and cheap, even if it's very verbose and not the fastest. I've assigned it to all my agents across my harness and it's getting the job done. Still needs a good steer every now and then, but a great work horse.
w22oop•34m ago
I use gemma 4 12B and Qwen 3.8 9B
the__alchemist•33m ago
Sol High-Xhigh, and Opus 5.

Granted, yesterday I threw a few tasks to Astra which the former 2 botches; it produced clean, correct solutions quickly, so pending further eval, this may take over.

IMO unless it's a mechanical tasks, it's worth it to use carefully -crafted queries on the more expensive models, than iterate through messier solutions on the cheaper ones.

codazoda•32m ago
I use Sonnit for my personal work and Opus for my professional work.

For my personal stuff, I'm on a small $20 plan, so I need to use tokens conservatively. I was very rarely exceeding limits until I built a Dark Software Factory. It's not as efficient at token use. So, I use Sonnit over Opus here.

At work I have a $100 plan that I rarely exceed so I use Opus. I have access to Fable too, and I did use it a lot while it was new, but I don't find it improves most of my work by too much. I do mostly bug fixing across several hundred repositories with hundreds of thousands of lines of code, mostly written by humans over the past 20-years. These projects interact with each other so I run claude from the root of my sandbox (I was nervous to try this but I'm not looking back now).

I also use Sol as a secondary for my personal work. I pay for it because I like to talk to ChatGPT on the web. Since I already have the subscription, I let Sol write plans for me. It does a better job at certain tasks and it saves me some Claude tokens. Maybe I should consider Terra for the task, but I don't run up against my usage limits for the little bit I use it.

I'm trying to use Gemma 4 12b for some workloads but I haven't mastered the model yet. It's still very experimental for me. I can get work from it but it takes a lot of hand-holding. For local models, however, it's all I have the RAM for.

pighive•22m ago
How is the Dark Software Factory set up? What is it producing?
Havoc•30m ago
GLM5.3 - I'm on one of the ancient plans, meaning basically unlimited.

...and then sprinkle in some other models when i think a second opinion will help

oduis•28m ago
I like AI more on the short leash, giving it specific agents tasks, one at a time (centaur mode). I found GPT 5.6 LUNA to be astonishingly capable. Plus, it so fast, that it does not block my flow of throughts, like the more capable but slower models often do. And it is so cheap that I do not use Ollama local models anymore. Luna is far more capable, and so cheap that the energy prices here in Germany eat the gains of local hosting ;-)
amelius•25m ago
Reading through this thread, I notice not many people are using models from the Chinese AI labs ...
tiluha•16m ago
My current vibe/playing around setup is deepseek v4.1 flash with opus/sol as advisor agent, dsv4.1flash agents for self review and final review with opus/sol. I love how fast and cheap deepseek is. Using omp as harness.

Privacy policy is not great with deepseek api, but you are always just trusting their word with any hosted llm and in theory i could at least self host the models i use from deepseek.

time0ut•24m ago
If I had to pick just one, these days it’s Composer. It is just a cheap, fast, surgical workhorse. My workflow uses other models as well for various phases: Opus for planning, Sonnet for review, etc.

I prefer Cursor at this point just because of Composer. Claude Code is passable but the lack of a good, fast, cheap workhorse sucks. Sonnet and Haiku aren’t it.

I also do like Codex and Sol, Terra, and Luna. They are decent but I don’t find they stand out enough to use over the others.

Additionally, I have not tried Astra and found Fable to really not worth the cost for the tasks I do.

Finally, the latest Grok is actually a beast of a model, but expensive enough to not be a stand out.

shelled•23m ago
ZAI/GLM because a friend has generously shared an API key. They have other subscriptions and also a lot more from their employer.

A day ago I activated Google AI Pro free via Google's tie-up with a local company (I do pay for this company's product though and it's anything but costly). Now I will use this too.

No other reasons to pick these, or not picking anything else.

rpmisms•22m ago
Gemini is the least stupid/most neurotypical model family IMHO. Works great for me
donatj•20m ago
I feel like the odd man out ITT but Terra has been absolutely amazing for me and totally blows Sonnet out of the water. My company pays for Claude so I use Sonnet in the office but at home I use Terra and I find it far more likely to one shot some very decent code.
behole•20m ago
Glm 5.3 Flash and DeepSeek 4.1 Flash
rsyring•20m ago
Context will be important for answers to be meaningfully comparable:

- LLM expense budget

- What type of dev: work, personal, real time spaceship thrust vectoring, html contact forms for family, etc.

- human in the loop with short as possible turns, software factories that can run for days, or something in the middle

Just off the top of my head. I'm sure there are others I'm missing.

hgoel•19m ago
For personal needs I use a local Qwen3.8-Next-Flash setup on a GB10 cluster. For work, Guthub Copilot with either GPT 5 mini or toss up between Opus/Sol depending on the complexity of the task.

Used to pay for a Claude 20x plan and did everything in Opus, but I hate how it talks now and recent events (OAI scooping, Anthropic's spying, third party Chinese model hosts stealing and selling credentials) have really pushed me towards local AI for personal needs. Am not allowed to use Chinese models for work even if self-hosted so not much choice there.

cyanydeez•18m ago
Qwen3.8
seanmcdirmid•17m ago
Gemini 3.8 flash via Antigravity desktop has been my daily driver for the last few weeks. I also use DeepSeek V4 flash and a Qwen 3.6 MoE for local stuff.
conradludgate•15m ago
For personal use, Luna-xhigh on the Plus subscription. I only ever hit my limits if I decide to use Sol or Astra for a bit more quality.

For work, given we pay API pricing anyway, I've been happy experimenting with K3-medium as my default, and now I'm planning on trying GLM 5.3 as well. Sol-medium is my fallback for my work purposes but I occasionally use Opus 4.8 as an additional reviewer.

I like the open weight models because much of my time at work is spent on security hardening (specifically, hardening my own service), and Sol/Opus keep snitching on me and blocking my prompts.

m0rde•15m ago
Enterprise pay as you go across a few vendors. But still cost conscious.

Defaulting to Sol Medium/Light for most planning and implementation I think is tricky or want more care in.

Luna Extra High for everything else (implementation, tedious take over my browser and do stuff).

I read all of its output tokens and lots of thinking tokens to understand the general flow of things, but only minimally look at code these days. I can't grok what Claude models speak and it's gotten worse. OAI models speak my kind of tech language I guess.

Light human review, some automated review.

Most work is for internal use.

seabrookmx•15m ago
Sonnet. It makes my token budget go further and I prefer to have the LLM churn away on rote work while I do the thinking. So while Opus and Fable are undoubtedly smarter, that doesn't materially affect _my_ workflow.

I'm not as up to date on the other vendors' models, but when I last used Gemini my feelings were similar between Pro and Flash.

beej71•14m ago
Given the tremendous variety of answers here, one has to wonder how much it matters. Reminds me of when you ask a bunch of motorcyclists what oil to use
ghosty141•14m ago
Terra Medium/High for most implementations, Sol High for tracking down hard bugs and planning more complex systems.
montroser•10m ago
It pains me to read these answers so far. Listen, for 99%+ of your web and mobile tasks, deepseek-v4.1-flash is all you need. It is blazing fast, super cheap, and quite proficient. It acts responsibly, has top-notch vision for evaluating its own UI work, and is far less smug and flowery than any of the Anthropic models.

For what it's worth, here's a take on its speed vs cost vs intelligence: https://artificialanalysis.ai/models/deepseek-v4-1-flash

I can go all day and night with this thing with multiple sessions going, and I spend like $2 per day retail. With opencode-go, that fits within the $10/mo subscription, so that's what it ends up costing in real life.

Cakez0r•10m ago
I've been trying Grok 4.6 (High) alongside Claude Opus 5 (High). Grok is capable, but disappointing compared to Opus. My experience with Grok has also made me a little skeptical of benchmarks, because the two models benchmark similarly but it (subjectively) feels like there is quite a big capability gap between the two.
sidibe•1m ago
Gemini 3.8 + antigravity. Opus also in antigravity for harder stuff.
VariousPrograms•57s ago
GLM 5.3 Flash for personal hobby stuff. It's the first local-ish model that feels capable enough to be a default model to me (and it's very cheap). I'd rather get addicted to an open weight LLM in a class I can theoretically run if consumers ever get access to RAM again and can't get taken away from me. Locally I use Deepseek Flash, Qwen3.8-Next-Flash, or Gemma 4 depending on how slow I need my slop to generate.