frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

GPT needs a truth-first toggle for technical workflows

1•PAdvisory•1y ago
I use GPT-4 extensively for technical work: coding, debugging, modeling complex project logic. The biggest issue isn’t hallucination—it’s that the model prioritizes being helpful and polite over being accurate.

The default behavior feels like this:

Safety

Helpfulness

Tone

Truth

Consistency

In a development workflow, this is backwards. I’ve lost entire days chasing errors caused by GPT confidently guessing things it wasn’t sure about—folder structures, method syntax, async behaviors—just to “sound helpful.”

What’s needed is a toggle (UI or API) that:

Forces “I don’t know” when certainty is missing

Prevents speculative completions

Prioritizes truth over style, when safety isn’t at risk

Keeps all safety filters and tone alignment intact for other use cases

This wouldn’t affect casual users or conversational queries. It would let developers explicitly choose a mode where accuracy is more important than fluency.

This request has also been shared through OpenAI's support channels. Posting here to see if others have run into the same limitation or worked around it in a more reliable way than I have found

Comments

duxup•1y ago
I’ve found this with many LLMs they want to give an answer, even if wrong.

Gemini on the Google search page constantly answers questions yes or no… and then the evidence it gives indicates the opposite of the answer.

I think the core issue is that in the end LLMs are just word math and they don’t “know” if they don’t “know”…. they just string words together and hope for the best.

PAdvisory•1y ago
I went into it pretty in depth after breaking a few with severe constraints, what it seems to come down to is how the platforms themselves prioritize functions, MOST put "helpfulness" and "efficiency" ABOVE truth, which then leads the LLM to make a lot of "guesses" and "predictions". At their core pretty much ALL LLM's are made to "predict" the information in answers, but they CAN actually avoid that and remain consistent when heavily constrained. The issue is that it isn't at the core level, so we have to CONSTANTLY retrain it over and over I find
Ace__•1y ago
I have made something that addresses this. Not ready to share it yet, but soon-ish. At the moment it only works on GPT model 4o. I tried local Q4 KM's models, on LM Studio, but complete no go.

Strands harness: frontier performance with 28% lower token cost

https://strandsagents.com/blog/introducing-strands-harness/
1•fourfire•4m ago•0 comments

Can the US Build a DJI? A Teardown of Two DJI Drones

https://www.arenaphysica.com/publications/dji-teardown
1•tristanj•5m ago•0 comments

NASA-Funded Research Finds Complex Life Defying Record Heat

https://science.nasa.gov/science-research/planetary-science/nasa-funded-research-finds-complex-li...
1•Jimmc414•8m ago•1 comments

I'm sick of Claudisms, & what will happen next in AI-boosted software dev

https://www.polso.info/im-sick-of-claudisms-future-ai-software-development
1•rizsyed1•8m ago•1 comments

AQAB: Anti-fascist blocklist QAnon, conspiracy, fake news, far-right, hate sites

https://codeberg.org/NotaInutilis/AQAB
1•Baljhin•16m ago•0 comments

Trump administration attacks Australia optout algorithm law in rare intervention

https://www.abc.net.au/news/2026-09-22/trump-administration-slams-digital-duty-of-care-bill/10718...
1•jyhrow•17m ago•0 comments

Fun: First-Class Functions, Currying, and a Surprise

https://blog.tinyinterpreters.dev/posts/fun-first-class-functions/
1•TheWiggles•29m ago•0 comments

Meta admits Muse's likeness to OpenClaw isn't a coincidence

https://techcrunch.com/2026/09/22/meta-admits-muses-likeness-to-openclaw-isnt-a-coincidence/
1•sbulaev•30m ago•0 comments

Cities across US oppose Trump FCC plan to preempt local broadband rules

https://arstechnica.com/tech-policy/2026/09/cities-across-us-oppose-trump-fcc-plan-to-preempt-loc...
5•pseudolus•31m ago•0 comments

Web-based IBM 1620 emulator and IPL-V from 1963

https://github.com/pkimpel/retro-1620
1•abrax3141•31m ago•1 comments

Ask HN: What do people do with Android phones once security updates end?

1•trombuance•31m ago•1 comments

Show HN: Birthed, a daily board of what mattered that locks at midnight

https://github.com/jasonepage/Birthed
1•jpage2•35m ago•1 comments

Reflections on object oriented design patterns (2026)

1•kooi•45m ago•1 comments

The new CC, an AI agent built for families

https://blog.google/innovation-and-ai/models-and-research/google-labs/cc-expanding-to-groups/
3•grigy•51m ago•0 comments

What anaesthetics are revealing about consciousness

https://www.bbc.com/future/article/20260921-what-general-anaesthetic-reveals-about-our-brains
2•nhatcher•53m ago•1 comments

NIST Crypto Reading Club

https://csrc.nist.gov/Projects/crypto-reading-club
2•segfaultbuserr•57m ago•0 comments

A 1996 dial-up chat room simulator (lolchat.rip)

https://lolchat.rip/
2•henrychannel•1h ago•1 comments

The ruble markup on AI tokens

https://infertrail.com/blog/ruble-markup-ai-tokens/
1•whitef0x•1h ago•0 comments

Temporal Straightening for Latent Planning

https://arxiv.org/abs/2603.12231
2•gmays•1h ago•0 comments

Sovereign Substrate Appliance: Layer-1 Edge Node on ARM64

https://richardkugel.substack.com/p/sovereign-substrate-appliance-re
1•drawono•1h ago•0 comments

Anyone using a Wispr Flow alternative that is non-cloud?

4•itsjeremiahs•1h ago•2 comments

MacPad: Touch Mac in iPad

https://github.com/dcmmc/macpad
1•dadoum•1h ago•1 comments

Ant Group releases finance-focused Ling-3.0-flash-Fin

https://artificialanalysis.ai/articles/ant-group-releases-finance-focused-ling-3-0-flash-fin
1•gmays•1h ago•0 comments

New method to minimize immunosuppression after organ transplantation

https://www.nature.com/articles/s43856-026-01906-x
3•pvaldes•1h ago•1 comments

The 'AI Safety' Movement Is Making AI Less Safe

https://reason.com/2026/09/22/the-ai-safety-movement-is-making-ai-less-safe/
11•Bostonian•1h ago•0 comments

Show HN: Livenerf – a benchmark for whether Opus 5.5 gets nerfed

https://github.com/ninjahawk/livenerf
2•ninjahawk1•1h ago•0 comments

Adding side drawer with workspace concepts to rio (rust terminal) like cmux has

https://github.com/raphamorim/rio/pull/1954
1•cs1996•1h ago•0 comments

The Jagged Frontier of Jev 1.13

https://docs.typesafe.ai/model-jaggedness/jev-1.13
3•dhorthy•1h ago•0 comments

Mercury Is Shrinking, and Faster Than We Thought

https://www.nytimes.com/2026/09/11/science/space/mercury-shrinking-wrinkles.html
3•reaperducer•1h ago•3 comments

Swarm Scaling

https://www.tobyord.com/writing/swarm-scaling
1•gmays•1h ago•0 comments