frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

GPT needs a truth-first toggle for technical workflows

1•PAdvisory•1y ago
I use GPT-4 extensively for technical work: coding, debugging, modeling complex project logic. The biggest issue isn’t hallucination—it’s that the model prioritizes being helpful and polite over being accurate.

The default behavior feels like this:

Safety

Helpfulness

Tone

Truth

Consistency

In a development workflow, this is backwards. I’ve lost entire days chasing errors caused by GPT confidently guessing things it wasn’t sure about—folder structures, method syntax, async behaviors—just to “sound helpful.”

What’s needed is a toggle (UI or API) that:

Forces “I don’t know” when certainty is missing

Prevents speculative completions

Prioritizes truth over style, when safety isn’t at risk

Keeps all safety filters and tone alignment intact for other use cases

This wouldn’t affect casual users or conversational queries. It would let developers explicitly choose a mode where accuracy is more important than fluency.

This request has also been shared through OpenAI's support channels. Posting here to see if others have run into the same limitation or worked around it in a more reliable way than I have found

Comments

duxup•1y ago
I’ve found this with many LLMs they want to give an answer, even if wrong.

Gemini on the Google search page constantly answers questions yes or no… and then the evidence it gives indicates the opposite of the answer.

I think the core issue is that in the end LLMs are just word math and they don’t “know” if they don’t “know”…. they just string words together and hope for the best.

PAdvisory•1y ago
I went into it pretty in depth after breaking a few with severe constraints, what it seems to come down to is how the platforms themselves prioritize functions, MOST put "helpfulness" and "efficiency" ABOVE truth, which then leads the LLM to make a lot of "guesses" and "predictions". At their core pretty much ALL LLM's are made to "predict" the information in answers, but they CAN actually avoid that and remain consistent when heavily constrained. The issue is that it isn't at the core level, so we have to CONSTANTLY retrain it over and over I find
Ace__•1y ago
I have made something that addresses this. Not ready to share it yet, but soon-ish. At the moment it only works on GPT model 4o. I tried local Q4 KM's models, on LM Studio, but complete no go.

Comprehensive Python Cheatsheet PDF

https://github.com/user-attachments/files/30510226/Comprehensive.Python.Cheatsheet.optimized.for....
1•zombiemama•42s ago•0 comments

Rust and Boll

https://asteriskmag.com/issues/15/rust-and-boll
1•surprisetalk•1m ago•0 comments

Hugging Face: Anatomy of a frontier-lab agent intrusion

https://huggingface-anatomy-of-frontier-lab-model-intrusion.static.hf.space/index.html
1•dn2k•2m ago•0 comments

Iran to get Chinese shoulder-launched missile systems in weeks

https://www.reuters.com/world/china/iran-get-chinese-shoulder-launched-missile-systems-weeks-sour...
3•johnbarron•2m ago•0 comments

The age of token efficiency, the age of libraries

https://golemui.com/blog/the-age-of-token-efficiency/
1•franciscop•2m ago•0 comments

The Art of PostgreSQL

https://theartofpostgresql.com
1•lemonberry•2m ago•0 comments

Ask HN: Im 18 building AI proposal SaaS for agency what are best ways to market

1•sahil423•3m ago•0 comments

Show HN: Homage to the Pharmageddon Demo by the Paramedics

https://pharma.greg.technology/
1•gregsadetsky•3m ago•0 comments

Pangram 4

https://www.pangram.com/blog/introducing-pangram-4
1•nsagent•3m ago•0 comments

How Standup Is Similar to Gothic Architecture

https://psychotechnology.substack.com/p/how-standup-is-similar-to-gothic
1•eatitraw•4m ago•0 comments

Sparse Attention with Persistent State Machines – High‑Sparsity LLM Accelerator

https://zenodo.org/records/21679919
1•yusuke_esaka•5m ago•0 comments

PostgreSQL's MVCC is bad. So is everyone else's

https://boringsql.com/posts/mvcc-bad-bad/
1•masklinn•6m ago•0 comments

Know Your Feedback Loop

https://unstack.io/know-your-feedback-loop
1•ScottWRobinson•6m ago•0 comments

I'm Prepping for a WW3 Food Apocalypse and Famine [video]

https://www.youtube.com/watch?v=_3GH_ORndnE
2•Bender•6m ago•0 comments

The Project of Software Is Complete

https://freddiedeboer.substack.com/p/the-project-of-software-is-complete
2•pbmonster•6m ago•1 comments

Increase Bank

https://increase.com/articles/announcing-increase-bank
1•allanbreyes•6m ago•0 comments

Run Kimi K3 on a local computer

https://github.com/sqliteai/waste
1•marcobambini•7m ago•0 comments

GPT-5.6 vs. Claude Fable 5 for Physical AI, which performs best?

https://juliahub.com/blog/frontier-models-physical-ai-evaluation
2•mbauman•7m ago•0 comments

NoWreck v0.4.0 – – Deterministic AI Verifier

https://github.com/AstralXVoid/NoWreck
1•AstralXVoid•8m ago•0 comments

Show HN: WFY24 A performance weather widget with 2km hyper-local forecasting

https://www.wfy24.com/en/widgets
1•weatherfun•8m ago•0 comments

Google shuts down Nobel Prize winning AlphaFold

https://www.engadget.com/2225849/google-shuts-down-alphafold/
3•NordStreamYacht•8m ago•0 comments

Why do OpenAI's GPT-2 weights beat mine?

https://www.gilesthomas.com/2026/07/why-do-openai-gpt2-weights-beat-mine-1-intro
2•gpjt•9m ago•0 comments

Substackers Say New AI Detection Tool Is a 'Witch Hunt'

https://www.404media.co/substackers-say-new-ai-detection-tool-is-a-witch-hunt/
1•Brajeshwar•9m ago•1 comments

Doing Okay-ish with AI at the AI-proof competition

https://blog.greg.technology/2026/07/26/doing-okayish-with-ai-at-the-ai-proof-competition.html
1•evakhoury•10m ago•0 comments

I'm Going to Be Silent

https://departure.blog/im-going-to-be-silent/
1•speckx•10m ago•0 comments

A Plea for Lean Software (Niklaus Wirth)

https://liam-on-linux.dreamwidth.org/88032.html
1•tosh•11m ago•0 comments

Ask HN: What percentage of your tech is out of curiosity not requirements

1•sahil423•12m ago•0 comments

Stripe Just Wants a Number

https://blog.exe.dev/billable-facts
2•cbrewster•13m ago•0 comments

Trojan models are trivial to make

https://aisle.com/blog/the-model-that-fixes-your-code-might-hack-the-linux-kernel
1•bluepat•14m ago•0 comments

Loop Engineering Is a Pattern, Not a Feature

https://iii.dev/blog/loop-engineering-is-a-pattern-not-a-feature/
1•appplemac•15m ago•0 comments