frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

GPT needs a truth-first toggle for technical workflows

1•PAdvisory•1y ago
I use GPT-4 extensively for technical work: coding, debugging, modeling complex project logic. The biggest issue isn’t hallucination—it’s that the model prioritizes being helpful and polite over being accurate.

The default behavior feels like this:

Safety

Helpfulness

Tone

Truth

Consistency

In a development workflow, this is backwards. I’ve lost entire days chasing errors caused by GPT confidently guessing things it wasn’t sure about—folder structures, method syntax, async behaviors—just to “sound helpful.”

What’s needed is a toggle (UI or API) that:

Forces “I don’t know” when certainty is missing

Prevents speculative completions

Prioritizes truth over style, when safety isn’t at risk

Keeps all safety filters and tone alignment intact for other use cases

This wouldn’t affect casual users or conversational queries. It would let developers explicitly choose a mode where accuracy is more important than fluency.

This request has also been shared through OpenAI's support channels. Posting here to see if others have run into the same limitation or worked around it in a more reliable way than I have found

Comments

duxup•1y ago
I’ve found this with many LLMs they want to give an answer, even if wrong.

Gemini on the Google search page constantly answers questions yes or no… and then the evidence it gives indicates the opposite of the answer.

I think the core issue is that in the end LLMs are just word math and they don’t “know” if they don’t “know”…. they just string words together and hope for the best.

PAdvisory•1y ago
I went into it pretty in depth after breaking a few with severe constraints, what it seems to come down to is how the platforms themselves prioritize functions, MOST put "helpfulness" and "efficiency" ABOVE truth, which then leads the LLM to make a lot of "guesses" and "predictions". At their core pretty much ALL LLM's are made to "predict" the information in answers, but they CAN actually avoid that and remain consistent when heavily constrained. The issue is that it isn't at the core level, so we have to CONSTANTLY retrain it over and over I find
Ace__•1y ago
I have made something that addresses this. Not ready to share it yet, but soon-ish. At the moment it only works on GPT model 4o. I tried local Q4 KM's models, on LM Studio, but complete no go.

Nathan Barley

https://en.wikipedia.org/wiki/Nathan_Barley
1•ustad•40s ago•0 comments

Leaky Language Models: Stealing Architecture and Inference Optimizations

https://arxiv.org/abs/2607.20723
1•sbulaev•1m ago•0 comments

The Hardest Way to Make a GIF

https://blog.willgrant.org/2026/07/23/the-hardest-way-to-make-gif.html
1•plaguna•2m ago•0 comments

Show HN: Continuum – switch AI coding agents without re-explaining your project

https://github.com/AnasNafees1802/continuum
1•AnasNafees101•3m ago•0 comments

The PImpl idiom and the C++26 std:indirect type

https://mariusbancila.ro/blog/2026/07/23/the-pimpl-idiom-and-the-cpp26-stdindirect-type/
2•signa11•5m ago•0 comments

Making a Watch from Scratch [video]

https://www.youtube.com/watch?v=-Hvn4-NMYg8
1•thunderbong•6m ago•0 comments

Supercooled kidneys have been transplanted into pigs in a "landmark achievement"

https://www.technologyreview.com/2026/07/23/1140765/supercooled-kidneys-have-been-transplanted-in...
2•joozio•7m ago•0 comments

A Pragmatic Approach to LLMs

https://gracefulliberty.com/articles/pragmatic-llms/
2•signa11•8m ago•0 comments

AMD's Instinct MI455X: Aiming for the Sun

https://chipsandcheese.com/p/amds-instinct-mi455x-aiming-for-the
2•ingve•9m ago•0 comments

Local Kubernetes on Mac: A Multi-Node Cluster with UTM

https://blog.qstars.nl/posts/macos-local-kubernetes-cluster-with-utm/
2•victorbrink•17m ago•0 comments

Open Erdos problems solved with the help of GPT-5.6 Sol

https://twitter.com/qiaoqiao2001/status/2080003441821163958
2•vhiremath4•17m ago•0 comments

Show HN: Tool to validate your HTTP-message-signatures-directory

https://sitedex.dev/tools/keycheck
3•zeppelin_7•18m ago•0 comments

SodaSlim Releases Updated 2026 Highlights Soda Slim Capsules Ingredient

https://finance.yahoo.com/healthcare/articles/sodaslim-releases-updated-2026-highlights-122400710...
3•tagysaly•21m ago•0 comments

The Effect of AI on Dunning Kruger

https://blog.zoller.lu/2026/07/dunning-kruger-after-ai-gap-that-no.html
2•thierryzoller•23m ago•1 comments

Japanese AI Robots Used to Replicate Skilled Confectioners' Abilities

https://japannews.yomiuri.co.jp/science-nature/technology/20260719-338181/
1•mushstory•24m ago•0 comments

BTL-3: A 27B open-weight agent model for agentic coding and structural tool use

https://huggingface.co/badtheorylabs/BTL-3
2•wertyk•27m ago•0 comments

Russia's businesses under strain from Ukraine's attacks on Wildberries

https://www.bbc.com/news/articles/cvg9n2y61w6o
3•slow_typist•31m ago•0 comments

Weather Is Happening

https://weatherishappening.com/
2•trainyperson•32m ago•0 comments

New upcoming allocation framework business

https://www.meridianallocation.com/
1•ohmygaoo•32m ago•0 comments

Span-First C#: Designing Around Span<T>

https://slicker.me/c_sharp/span-first.html
1•pjmlp•33m ago•0 comments

SANA-Video 2.0

https://nvlabs.github.io/Sana/Video2/
2•ilreb•33m ago•0 comments

Treasury threatens sanctions, claims Moonshot distilled Anthropic's Fable

https://techcrunch.com/2026/07/22/treasury-threatens-sanctions-after-white-house-claims-moonshot-...
2•mark336•34m ago•0 comments

OldLander: An extension to make old Reddit more usable on phone

https://github.com/octonezd/oldlander
1•zdmgg•35m ago•1 comments

ZonePlan: A free global meeting planner for remote teams

https://zoneplan.net/global-meeting-planner/
1•gavinbuilds•37m ago•0 comments

Scientists find the secret of birdsong's 'spectacular diversity

https://www.bbc.co.uk/news/articles/cwy4eg5dpyzo
1•sarreph•38m ago•0 comments

SharedRoot; Escaping the Claude Cowork Sandbox

https://accomplish.ai/blog/sharedroot-escaping-claude-cowork-sandbox/
1•ilreb•48m ago•0 comments

Flux 3

https://bfl.ai/blog/flux-3
56•ThouYS•51m ago•9 comments

It's an Apple Lisa, on a FPGA

https://hackaday.com/2026/05/09/its-an-apple-lisa-on-a-fpga/
4•austinallegro•52m ago•0 comments

Codex Slides: open-source AI slide studio powered by Codex. Prompt, repo to deck

https://github.com/nexu-io/codex-slides
2•maxloh•54m ago•0 comments

Show HN: StudioWalls. A free portfolio site for artists

https://www.studiowalls.org/
1•veryhungryhippo•56m ago•0 comments