frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

GPT needs a truth-first toggle for technical workflows

1•PAdvisory•1y ago
I use GPT-4 extensively for technical work: coding, debugging, modeling complex project logic. The biggest issue isn’t hallucination—it’s that the model prioritizes being helpful and polite over being accurate.

The default behavior feels like this:

Safety

Helpfulness

Tone

Truth

Consistency

In a development workflow, this is backwards. I’ve lost entire days chasing errors caused by GPT confidently guessing things it wasn’t sure about—folder structures, method syntax, async behaviors—just to “sound helpful.”

What’s needed is a toggle (UI or API) that:

Forces “I don’t know” when certainty is missing

Prevents speculative completions

Prioritizes truth over style, when safety isn’t at risk

Keeps all safety filters and tone alignment intact for other use cases

This wouldn’t affect casual users or conversational queries. It would let developers explicitly choose a mode where accuracy is more important than fluency.

This request has also been shared through OpenAI's support channels. Posting here to see if others have run into the same limitation or worked around it in a more reliable way than I have found

Comments

duxup•1y ago
I’ve found this with many LLMs they want to give an answer, even if wrong.

Gemini on the Google search page constantly answers questions yes or no… and then the evidence it gives indicates the opposite of the answer.

I think the core issue is that in the end LLMs are just word math and they don’t “know” if they don’t “know”…. they just string words together and hope for the best.

PAdvisory•1y ago
I went into it pretty in depth after breaking a few with severe constraints, what it seems to come down to is how the platforms themselves prioritize functions, MOST put "helpfulness" and "efficiency" ABOVE truth, which then leads the LLM to make a lot of "guesses" and "predictions". At their core pretty much ALL LLM's are made to "predict" the information in answers, but they CAN actually avoid that and remain consistent when heavily constrained. The issue is that it isn't at the core level, so we have to CONSTANTLY retrain it over and over I find
Ace__•1y ago
I have made something that addresses this. Not ready to share it yet, but soon-ish. At the moment it only works on GPT model 4o. I tried local Q4 KM's models, on LM Studio, but complete no go.

I Redefined Partnership Terms to End the Sick Romanticization of Co-Founders

https://medium.com/startup-shop/in-lieu-of-technical-co-founder-i-redefined-better-partnership-te...
1•joeyomaita•1m ago•1 comments

Covert Caches: TLB Edition

https://parallelprogrammer.substack.com/p/covert-caches-tlb-edition
2•matt_d•4m ago•0 comments

Ukrainian naval drone sinks Russian kamikaze drone boat

https://arstechnica.com/gadgets/2026/09/military-milestone-ukrainian-naval-drone-sinks-russian-ka...
2•Gaishan•5m ago•0 comments

Making a game for the GBA and PC from the same codebase

https://mattgreer.dev/blog/making-a-game-for-gba-and-pc/
2•ibobev•7m ago•0 comments

Designing Silicon from Scratch [video]

https://www.youtube.com/watch?v=q6ytRHaTEXI
2•ibobev•8m ago•0 comments

Looking forward to Git 2.56 – and 3.0

https://lwn.net/SubscriberLink/1094575/2385e98583715c2b/
3•chmaynard•9m ago•0 comments

Self-Hosting Behind Cgnat

https://david.alvarezrosa.com/posts/self-hosting-behind-cgnat/
3•ibobev•10m ago•0 comments

Brad Smith and Kevin Rudd on AI, Regulation, and the Future of Work [video]

https://www.youtube.com/watch?v=hClbN_imH5Y
2•verdverm•10m ago•0 comments

U.S. Site Blocking Bill Adds VPNs to the List of Blocking Intermediaries

https://torrentfreak.com/u-s-site-blocking-bill-adds-vpns-to-the-list-of-blocking/
3•gslin•13m ago•0 comments

Show HN: Building Seven in Public

https://claude.ai/share/61ae35d6-82f2-425a-84b6-907a3ec1b5c4
2•o2zer0cool•13m ago•0 comments

Apple Watch Series 12 and Ultra 4 May Have Hidden Flash Storage

https://www.macrumors.com/2026/09/21/apple-watch-series-12-ultra-4-hidden-storage/
2•bookofjoe•17m ago•0 comments

Muse, Meta's extraordinarily privileged AI assistant, has a serious 0-day

https://arstechnica.com/security/2026/09/muse-metas-extraordinarily-privileged-ai-assistant-has-a...
4•spenvo•18m ago•0 comments

AI Risk Cheatsheet

https://umais.me/writing/ai-risk-cheatsheet/
2•umz•22m ago•0 comments

Spymarks, Not Watermarks

https://brand.io/article/spymarks/
7•possibilistic•23m ago•1 comments

Acubom

https://acubom.com/
2•autifill•23m ago•0 comments

Jev: System One Models for Prod, Not God – With Diogo Almeida, CEO, TypeSafe AI

https://www.latent.space/p/jev
4•swyx•24m ago•1 comments

I Am in need of Testers

https://agaro.ai/
4•yusufghyasi•30m ago•1 comments

An actively maintained and updated Motif fork exists

https://www.osnews.com/story/145877/an-actively-maintained-and-updated-motif-fork-actually-exists/
2•birdculture•31m ago•0 comments

LLMs Are Too Big. My Log Router Doesn't Need to Sing

https://www.distributedthoughts.org/my-log-router-doesnt-need-to-sing/
3•pgr0ss•32m ago•0 comments

Trump says U.S. working on deal to buy potash from Belarus instead of Canada

https://www.cbc.ca/lite/story/9.7352583
7•colinprince•33m ago•1 comments

Markdown in /src

https://htmx.org/essays/markdown-in-src/
1•perrygeo•38m ago•0 comments

Self-hosted, end-to-end encrypted 1:1 chat app. Offline key-exchange

https://github.com/pkMinhas/walkytalky-private-messenger
1•pkMinhas•39m ago•0 comments

Show HN: Transom – easily manage multiple browser profiles on macOS

https://github.com/darvid/transom
1•darvid•42m ago•0 comments

Where do AI norms come from?

https://charity.wtf/p/where-do-ai-norms-come-from
1•BerislavLopac•43m ago•0 comments

Bitrig Now Builds iPhone Duo Apps

https://bitrig.com/blog/bitrig-builds-iphone-duo-apps
1•alwillis•44m ago•0 comments

Do not fear AI. Fear AI companies

https://df7sc6o35ljoz.cloudfront.net/posts/do-not-fear-ai-rev-2.html
8•rDr4g0n•46m ago•1 comments

An LLM Beat NetHack

https://kenforthewin.github.io/blog/posts/llm-nethack-ascension/
1•kenforthewin•46m ago•0 comments

Singapore auctions luxury goods seized in $2.4B money-laundering bust

https://www.cnn.com/2026/09/14/style/singapore-money-laundering-auctions
1•kelt•47m ago•0 comments

Tupo: A daily sudoku-style logic puzzle with a colorful twist

https://playtupo.com/
1•slymax•47m ago•0 comments

Entropy Is Spare Capacity

https://fffej.substack.com/p/entropy-is-spare-capacity
1•mooreds•50m ago•0 comments