frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Ask HN: How come everyone is an LLM expert?

3•delis-thumbs-7e•10h ago
It is a bit strange how new model is released and hour after there is commentators declaring it complete trash and embarrassment to the AI industry, or the best thing since sliced bread. Surely they have not had the opportunity to test the ins and outs of the model yet? Or do people just blindly trust benchmarks as if they were not pretty easy to manipulate, as research has shown quite a few times now? Or is it just all vibe?

So how do you measure how one model is better than another?

Comments

tolugenius•10h ago
There is probably far, far more people trusting benchmarks and "I remade x thing in 1 prompt with y model" claims than you'd imagine, just ignore all of it. You know your workflow and what better should be and could be, measure on what works for you. You should note (and I may be wrong, not active in these part) a lot of those demos are very toy, recreating a known game, known app, known workflow, etc. very interesting but again a very toy example that should be taken with that in mind.
bigyabai•10h ago
Benchmarks can still be useful, even for benchmaxxing labs. For example, the recent Beam model benchmarked much worse than DS4.1 Flash and GLM 5.3, both of which perform extremely well outside of benchmarking. Regardless of whether or not Beam was benchmaxxed, it's performance suggested that it wasn't capable of solving problems that other models in it's weight class could do easily.

The best-case scenario is that Reflection was being honest about their model's disappointing performance. The worst-case scenario is that they benchmaxxed, and it still managed to underperform compared to it's peers.

spottedmarley•9h ago
I built my own benchmarking arena that tests local models on all of things the I need a model to do well. I don't look at any of the existing benchmark data that is out there. When a new model drops, I run it through my arena and see how it compares to previous models. If I talk about a model being good I am referencing my own accumulated knowledge on how a model performs for me on tasks that I care about. I generally will never be heard talking negatively about a model (except maybe a frontier/hosted model, they all suck in their own ways) because if a model sucks it just gets deleted and I move on to other things. I suppose I'd consider myself somewhat of an 'expert' when it comes to analyzing local model performance, but I don't really listen too much to what anyone else says about them, or which benchmarks tell them which things about a model. Just test them on the things that are important to you.

Ask HN: Are there AI models for generating sounds based on a text and reference?

22•onemiketwelve•1d ago•12 comments

Ask HN: What do you think about Fractional CFOs

3•mtmosestn•1h ago•3 comments

Ask HN: Why is Ask HN only showing me 14 posts?

40•Gooblebrai•5h ago•37 comments

Ask HN: What Is Your Personal Backup Methods

2•Crontab•1h ago•0 comments

Ask HN: What do you run on a $5 VPS that's worth keeping online 24/7?

11•mariocesar•3h ago•4 comments

Ask HN: What games do you religiously play on your phone browser?

6•saimiam•3h ago•13 comments

Ask HN: Do you still own a printer?

9•nunorbatista•8h ago•13 comments

Ask HN: How to deal with AI "true believer" leadership at work

5•microflash•9h ago•3 comments

Ask HN: How do you describe what your relationship brings you?

2•maxignol•7h ago•3 comments

Ask HN: What would you like to see in a new AI technology release?

2•ikishade•8h ago•2 comments

Ask HN: Vibecoding Follies

4•vegnus•10h ago•2 comments

Ask HN: How come everyone is an LLM expert?

3•delis-thumbs-7e•10h ago•3 comments

ChatGPT self-distances when admitting fault

3•chrisjj•4h ago•1 comments

Tell HN: GitHub refuses to remove cracked copies of my software after a month

54•IvanK_net•8h ago•52 comments

Ask HN: Show me your agentic SDE semantics

2•danielovichdk•12h ago•0 comments

Ask HN: Is observability broken for you?

2•tmach32•14h ago•0 comments

Ask HN: My Apple ID is locked for a week, Apple is a single point of failure

11•akg_67•15h ago•3 comments

Ask HN: Agent access to chat data should require participant consent?

2•maxwellito•17h ago•0 comments

Ask HN: Alternatives to NeurIPS, ICLR and ICML

3•john-titor•19h ago•0 comments

Ask HN: What brings you back to personal AI agents like Instinct and Muse?

3•sdrth•21h ago•1 comments

Orcah Studio: A local-first video agent that can search your videos

8•iliashad•1d ago•5 comments

Ask HN: Why does Astra compact context so frequently vs. Fable?

2•yesitcan•1d ago•0 comments

Ask HN: What do you think about AI generated slides for conferences?

4•Hixon10•1d ago•1 comments

Ask HN: How are AI budgets changing in your company?

5•bobby-cb•1d ago•2 comments

Tell HN: 2026 is the year of the Linux desktop, agent-adjusted

3•dvrp•1d ago•2 comments

Tell HN: Uceprotect is extorting website owners

101•goldenmember•1d ago•59 comments

Our digital privacy is being aggressively disbanded this week

12•Steaglsz•1d ago•1 comments

DOS Game Stunts Port to Linux,Windows,Browser,etc.

4•LowLevelMahn•2d ago•1 comments

Who is cleaning up all the garbage LLMs generate?

6•kbrannigan•2d ago•7 comments

Ask HN: Anyone accepted into OpenAI "Codex for open source"?

4•awb•1d ago•0 comments