frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

GPT needs a truth-first toggle for technical workflows

1•PAdvisory•1y ago
I use GPT-4 extensively for technical work: coding, debugging, modeling complex project logic. The biggest issue isn’t hallucination—it’s that the model prioritizes being helpful and polite over being accurate.

The default behavior feels like this:

Safety

Helpfulness

Tone

Truth

Consistency

In a development workflow, this is backwards. I’ve lost entire days chasing errors caused by GPT confidently guessing things it wasn’t sure about—folder structures, method syntax, async behaviors—just to “sound helpful.”

What’s needed is a toggle (UI or API) that:

Forces “I don’t know” when certainty is missing

Prevents speculative completions

Prioritizes truth over style, when safety isn’t at risk

Keeps all safety filters and tone alignment intact for other use cases

This wouldn’t affect casual users or conversational queries. It would let developers explicitly choose a mode where accuracy is more important than fluency.

This request has also been shared through OpenAI's support channels. Posting here to see if others have run into the same limitation or worked around it in a more reliable way than I have found

Comments

duxup•1y ago
I’ve found this with many LLMs they want to give an answer, even if wrong.

Gemini on the Google search page constantly answers questions yes or no… and then the evidence it gives indicates the opposite of the answer.

I think the core issue is that in the end LLMs are just word math and they don’t “know” if they don’t “know”…. they just string words together and hope for the best.

PAdvisory•1y ago
I went into it pretty in depth after breaking a few with severe constraints, what it seems to come down to is how the platforms themselves prioritize functions, MOST put "helpfulness" and "efficiency" ABOVE truth, which then leads the LLM to make a lot of "guesses" and "predictions". At their core pretty much ALL LLM's are made to "predict" the information in answers, but they CAN actually avoid that and remain consistent when heavily constrained. The issue is that it isn't at the core level, so we have to CONSTANTLY retrain it over and over I find
Ace__•1y ago
I have made something that addresses this. Not ready to share it yet, but soon-ish. At the moment it only works on GPT model 4o. I tried local Q4 KM's models, on LM Studio, but complete no go.

We Like Things

https://asteriskmag.com/issues/15/why-we-like-things
1•surprisetalk•20s ago•0 comments

IBM builds a better fridge for its quantum computers

https://thenewstack.io/ibm-modular-quantum-refrigerator/
1•Brajeshwar•1m ago•0 comments

Show HN: A design skill that turn your feelings into UI and remembers it

https://github.com/MonkeyUI-dev/vibe-to-ui
1•LeonTung•2m ago•0 comments

Show HN: Synchroize, control, and orachstrate any angets on any device

https://github.com/cosyncing/cosyncing
1•howardme1•2m ago•0 comments

Commenting no longer works on old.reddit.com

https://old.reddit.com/r/help/comments/1vson02/comments_do_not_submit_on_oldreddit_or_res_but_do/
2•OgsyedIE•3m ago•0 comments

Alexander Vampilov

https://en.wikipedia.org/wiki/Alexander_Vampilov
1•petethomas•4m ago•0 comments

Prediction Market Cheating Gets Creative

https://www.wsj.com/opinion/prediction-market-cheating-gets-creative-3068365e
1•Anon84•4m ago•0 comments

Show HN: MandarinClips – Learn conversational Chinese from 130k+ TV drama clips

https://www.mandarinclips.com/en
1•mandarinclips•4m ago•0 comments

The science behind Pixel Watch's insulin resistance feature

https://www.empirical.health/blog/wearable-insulin-resistance/
2•brandonb•5m ago•0 comments

Study Finds Narwhal Tusks Have 2 Spirals – Not 1–Twisting in Opposite Directions

https://arstechnica.com/science/2026/08/x-rays-add-new-twist-to-narwhals-spiral-tusk/
1•bookofjoe•7m ago•0 comments

A Doctor Who Became Afraid of Death

https://drped.substack.com/p/a-doctor-who-became-afraid-of-death
5•jamarna•7m ago•0 comments

I built a courtroom for the internet. Verdicts stay sealed until you vote

https://courtofstrangers.com
1•briancohen•7m ago•0 comments

Old-School Electronics Repair Man Vows to Be the Last in Chicago

https://blockclubchicago.org/2026/08/19/old-school-electronics-repair-man-vows-to-be-the-last-in-...
1•toomuchtodo•8m ago•1 comments

Mojo is now open source

https://simonwillison.net/2026/Aug/18/mojo-is-now-open-source/
2•jonathandeamer•10m ago•0 comments

Show HN: Tech / AI / Product super aggregator – SudoReport

https://sudoreport.com/
3•ataturkle•10m ago•0 comments

Testtospeech.com – New leaderboard adds balance and removes bias

https://texttospeech.com/
2•imemily•11m ago•0 comments

Show HN: Zmina – clipboard with per-app paste transforms

https://zmina.app/
3•0-3•12m ago•0 comments

Product.now

https://product.now
1•luispa•12m ago•0 comments

Life finds a way OpenAI claims partial pause and rolls out ChatGPT for Teens

https://aistop.watch/p/life-finds-a-way
1•Bluestein•12m ago•0 comments

Landing the Plane

https://www.theengineeringmanager.com/growth/landing-the-plane/
1•LaSombra•13m ago•0 comments

Show HN: Store Front For Vibe coded apps

https://yaadops.com/explore
2•ShamarWebster•13m ago•0 comments

Algorithms + Data Structures = Programs

https://en.wikipedia.org/wiki/Algorithms_%2B_Data_Structures_%3D_Programs
1•tosh•13m ago•0 comments

The Data Center Capital of the World [video]

https://www.nytimes.com/video/us/100000011066777/inside-the-data-center-capital-of-the-world.html
1•donohoe•13m ago•0 comments

Ask HN: Calorie Trackers

2•solsane•13m ago•0 comments

What is inference engineering? Deepdive

https://newsletter.pragmaticengineer.com/p/what-is-inference-engineering
1•luispa•13m ago•0 comments

The Hermès heist: how an heir to the dynasty was swindled out of $15B of shares

https://www.economist.com/1843/2025/12/11/the-hermes-heist-how-an-heir-to-the-luxury-dynasty-was-...
1•geneticdrifts•13m ago•0 comments

Ornith-1.5: From Self-Scaffolding to Self-Improvement

https://ornith.ai/ornith_1_5.html
2•CommonGuy•14m ago•0 comments

X262: X264 with MPEG-2 Support

https://github.com/kierank/x262
1•ksec•14m ago•0 comments

The two largest reservoirs in the US have hit record-low levels

https://www.carbonbrief.org/analysis-the-two-largest-reservoirs-in-the-us-have-hit-record-low-levels
1•speckx•15m ago•0 comments

Does daycare damage children's brains?

https://www.worksinprogress.news/p/daycare-barely-matters
1•momentmaker•17m ago•0 comments