frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Fine-tuning your LLM might be making it dumber

https://github.com/aryanyaksh-art/did-it-even-learn
1•aryan_yaksh•50m ago

Comments

f33d5173•6m ago
Fine tuning is not necessarily oriented at making a model better at generic "intelligence" tasks. In fact, the goal is often specifially to improve model performance at narrow tasks while recognizing it will get worse at other tasks.

I clicked through the leaderboard. It seemingly isn't sorted by best or worst performing, but rather by absolute change in performance - improved and worsened models are interleaved. Regardless, I selected the model that worsened the most, and investigated what they claimed their fine tuning accomplished; in fact they claimed nothing, their model sheet was the default trl model sheet without even the training set listed. The second worst performing was the same, but they at least listed their training set. It was a set of millions of word math problems. So we would expect the model to get better at those problems and worse at everything else.

On the other hand, the most improved model on the table had an actual model sheet, where they claim the model has been trained to be "more helpful". They claim that by ablating model refusal the model ends up doing better on benchmarks as well. That claim seems to have been borne out.

Obviously if you train a model for some purpose it will not necessarily do well at that purpose. But I would expect popularily used models to be ones that actually do what they say they do. I don't think people need to be told that sometimes models are badly trained. And the claim that this project actually has data to support, that finetuning on one task does not necessarily improve performance on another task, is so obvious that it doesn't need to be stated.

Show HN: Dagychu – self-hosted platform for running and managing automation

https://github.com/raideria-software/dagychu
1•titansoup•14s ago•0 comments

Retirement Is Dead

https://macleans.ca/longforms/retirement-is-dead/
2•speckx•32s ago•0 comments

The Logistics Nightmare Facing U.S. Warships

https://www.nytimes.com/interactive/2026/09/02/us/iran-war-us-navy-carriers-supply.html
1•thelastgallon•1m ago•0 comments

My girlfriend asked me why I have 15 Codex subscriptions

https://hraness.com/writing/my-girlfriend-asked-me-why-i-have
1•thoughtpeddler•2m ago•0 comments

Draft of an idea for an opt-in legal and economic framework needing input(paper)

https://zenodo.org/records/22152670
1•KitCrowCo•3m ago•0 comments

AI Policy

https://dbushell.com/ai/
2•duck•3m ago•0 comments

Is this the simplest (and most surprising) sorting algorithm?

https://arxiv.org/abs/2110.01111
2•porridgeraisin•5m ago•0 comments

Lachy Groom backs Indian startup aiming to keep aircraft aloft for a year

https://techcrunch.com/2026/08/31/lachy-groom-backs-indian-startup-aiming-to-keep-aircraft-aloft-...
1•porridgeraisin•5m ago•0 comments

Qwen3.8-Max Checkpoint 0902

https://www.qwencloud.com/models/qwen3.8-max-0902
1•fbrusch•6m ago•0 comments

Ultra Fast LLM's Are Interesting (Mercury 2.5 Preview Release)

https://openrouter.ai/inception/mercury-2.5-preview
1•gmarkwa•8m ago•1 comments

Show HN: Kit. Claude Code but Concise

https://github.com/speakeasy-api/kit
3•danielkov•9m ago•1 comments

FUTO Notes v1.7.1

https://notes.futo.tech/blog/futo-notes-1-7-1/
3•somewhatjustin•10m ago•1 comments

So what is a Qubit, anyway?

https://e-diamond.github.io/blog/2026/08/15/what-is-a-qubit
2•vismit2000•11m ago•0 comments

Why Adding More AI Agents Makes Your Team Slower

https://blog.mempko.com/misconception-about-agentic-scaling/
1•mempko•12m ago•0 comments

Satellite images before Nepal disaster showed warning signs

https://www.nature.com/articles/d41586-026-02746-4
1•Brajeshwar•13m ago•0 comments

Grok Bot Skills

https://github.com/shrdgn/grokbot-skills
1•shadag•14m ago•0 comments

SteamdDB Joins Nexus Mods

https://www.nexusmods.com/news/15597
4•HelloUsername•14m ago•0 comments

Visa Guide for IT Professional How to Work in the USA, Canada and Europe in 2026

https://www.foxapply.com/blog/en/visa-guide-it-professionals-usa-canada-europe
3•josanjohnata•15m ago•0 comments

A New Observable

https://observablehq.com/@observablehq/a-new-observable
2•yurivish•15m ago•0 comments

Ask HN: How to ask questions in Ask HN

1•samikama•16m ago•1 comments

Why don't MCP servers tell agents what their tools return?

https://www.blacksmith.sh/blog/code-smith-code-mode
1•adityamaru•16m ago•0 comments

Intellectual and Moral Bankruptcy

https://philoserf.com/posts/intellectual-and-moral-bankruptcy/
2•speckx•17m ago•1 comments

With American Characteristics

https://newsletter.doomberg.com/p/with-american-characteristics
1•simonebrunozzi•17m ago•0 comments

Ask HN: Would you be interested in using a VoIP system that has an operator?

2•left-struck•18m ago•0 comments

Rustup 1.29.1

https://blog.rust-lang.org/2026/09/01/Rustup-1.29.1/
2•andrewstetsenko•18m ago•0 comments

Show HN: Tin Computer, an autonomous growth agent for your projects

https://tin.computer
1•demegire•19m ago•0 comments

Peptides can form well-defined structures in harsh, Venus-like conditions

https://news.mit.edu/2026/study-peptides-can-form-well-defined-structures-harsh-venus-conditions-...
1•geox•19m ago•0 comments

The United States Is Not America

https://www.olafalders.com/2026/09/02/the-united-states-is-not-america/
3•oalders•20m ago•1 comments

Polyend Keys – a mechanical keyboard with a built-in MIDI layout

https://theplayground.co.uk/polyend-turns-the-computer-keyboard-into-a-midi-instrument-with-new-k...
1•zakxxi•20m ago•0 comments

Review: Hacker News by Lyle Bramwell

https://bramwellreviews.com/reviews/hackernews.htm
1•kinduff•20m ago•0 comments