frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

GPT needs a truth-first toggle for technical workflows

1•PAdvisory•1y ago
I use GPT-4 extensively for technical work: coding, debugging, modeling complex project logic. The biggest issue isn’t hallucination—it’s that the model prioritizes being helpful and polite over being accurate.

The default behavior feels like this:

Safety

Helpfulness

Tone

Truth

Consistency

In a development workflow, this is backwards. I’ve lost entire days chasing errors caused by GPT confidently guessing things it wasn’t sure about—folder structures, method syntax, async behaviors—just to “sound helpful.”

What’s needed is a toggle (UI or API) that:

Forces “I don’t know” when certainty is missing

Prevents speculative completions

Prioritizes truth over style, when safety isn’t at risk

Keeps all safety filters and tone alignment intact for other use cases

This wouldn’t affect casual users or conversational queries. It would let developers explicitly choose a mode where accuracy is more important than fluency.

This request has also been shared through OpenAI's support channels. Posting here to see if others have run into the same limitation or worked around it in a more reliable way than I have found

Comments

duxup•1y ago
I’ve found this with many LLMs they want to give an answer, even if wrong.

Gemini on the Google search page constantly answers questions yes or no… and then the evidence it gives indicates the opposite of the answer.

I think the core issue is that in the end LLMs are just word math and they don’t “know” if they don’t “know”…. they just string words together and hope for the best.

PAdvisory•1y ago
I went into it pretty in depth after breaking a few with severe constraints, what it seems to come down to is how the platforms themselves prioritize functions, MOST put "helpfulness" and "efficiency" ABOVE truth, which then leads the LLM to make a lot of "guesses" and "predictions". At their core pretty much ALL LLM's are made to "predict" the information in answers, but they CAN actually avoid that and remain consistent when heavily constrained. The issue is that it isn't at the core level, so we have to CONSTANTLY retrain it over and over I find
Ace__•1y ago
I have made something that addresses this. Not ready to share it yet, but soon-ish. At the moment it only works on GPT model 4o. I tried local Q4 KM's models, on LM Studio, but complete no go.

Mistral is serving GLM 5.2

https://docs.mistral.ai/models/zai-glm-5-2
1•Vaskivo•2m ago•1 comments

New DeepSeek Pricing Published (Peak and Off-Peak)

1•mittermayr•2m ago•0 comments

The Importance of Whimsy

https://christiansafka.com/blog/the-importance-of-whimsy/
1•christiansafka•2m ago•0 comments

Let's not call it "tech debt," it's just "mess"

https://www.simplermachines.com/lets-not-call-it-tech-debt-its-just-mess/
1•birdculture•3m ago•0 comments

Drift Anchor -quick access to your proven AI prompts

https://microsoftedge.microsoft.com/addons/detail/dpedbbcbfmpjjonjolngjjpclkgahlpm
1•ballista2026•5m ago•0 comments

The Reason Data Center Gas Power Plants Are So Dirty

https://www.wired.com/story/real-reason-data-center-gas-power-plants-are-so-dirty/
1•beardyw•10m ago•0 comments

Building a practical path to post-quantum cryptography

https://www.technologyreview.com/2026/08/13/1141041/building-a-practical-path-to-post-quantum-cry...
1•theanonymousone•10m ago•0 comments

There Is Still No Silver Bullet

https://cekrem.github.io/posts/there-is-still-no-silver-bullet/
1•theanonymousone•11m ago•1 comments

Predictions for the Era of Continual Learning

https://www.dwarkesh.com/p/era-of-continual-learning
1•ronfriedhaber•14m ago•0 comments

ECS Lua

https://nidorx.github.io/ecs-lua/
1•kqr•14m ago•0 comments

How much time do you spend collecting invoices from websites? WICG proposal

https://github.com/WICG/proposals/issues/196
1•collimarco•17m ago•1 comments

Multi-Player Snake

https://multi-snake.michielborkent.nl/
1•Borkdude•18m ago•0 comments

OpenAI's Rumored Speaker Belongs on Keychains

https://ryanspahn.substack.com/p/openais-rumored-speaker-belongs-on
1•paul7986•18m ago•0 comments

Ableton Live and Push can now run on Linux, unofficially

https://cdm.link/ableton-live-on-linux/
1•dyzone•20m ago•0 comments

Working on Economics with Fable 5

https://wilsoniumite.com/2026/08/03/working-on-economics-with-fable-5/
1•Wilsoniumite•20m ago•0 comments

Critical theory engaging with discourse about nervous systems (2025)

https://www.reddit.com/r/CriticalTheory/comments/1oe78lz/critical_theory_engaging_with_current_me...
1•toilet•25m ago•0 comments

Ironic-Python-Agent: Container HardwareManager Security Model Misimplemented

https://seclists.org/oss-sec/2026/q3/488
1•runningmike•30m ago•0 comments

Show HN: I turned X's ranking algorithm into a game video

https://x-algorithm-weights.vercel.app
1•lime66•33m ago•0 comments

Are You Using AI? Do Not Forget Your Whip

https://skyecleary.substack.com/p/are-you-using-ai-do-not-forget-your
1•plastic-enjoyer•34m ago•0 comments

Roadmaps to Unicode Plane 1: Supplementary Multilingual Plane

https://www.unicode.org/roadmaps/smp/
1•Bluestein•36m ago•0 comments

OpenRouter: Unified LLM API with Routing and Fallbacks

https://trpevski.com/blog/openrouter-unified-llm-api-with-routing-and-fallbacks/
1•dzugumot•37m ago•1 comments

Multi-model chatbot back ends: contracts, routing, and fallbacks

https://medium.com/@CometAPI_/choose-the-best-backend-api-for-a-multi-model-ai-chatbot-b1e843f9cbf4
1•WavyPeng•39m ago•0 comments

Beyond Source: An Empirical Study of Python Bytecode Security Risks

https://arxiv.org/abs/2608.12853
1•runningmike•41m ago•0 comments

Of felt hats, feathers, macaroni, and weasels

https://languagelog.ldc.upenn.edu/nll/?p=24590
1•Bluestein•41m ago•0 comments

NASA Swift Boost Spacecraft Prepares to Continue Mission

https://science.nasa.gov/blogs/swift/2026/08/11/nasa-swift-boost-spacecraft-prepares-to-continue-...
2•croes•45m ago•0 comments

Show HN: Security scanner for SaaS build with AI

https://sentrint.com/sample
1•Gourabdg•46m ago•0 comments

Not Even Taste Would Be Left (Probably)

https://sebastian.graphics/blog/not-even-taste-would-be-left-probably.html
1•kasumispencer2•46m ago•0 comments

DeepSeek Harness developer preview: Everything is a plugin

https://www.deepseek.com/harness/en/
2•fernvenue•51m ago•0 comments

Spain extends Almaraz nuclear plant operations through 2030

https://www.reuters.com/business/energy/spain-extends-almaraz-nuclear-plant-operations-through-20...
5•mpweiher•51m ago•0 comments

Gitlawb Node, Federated Git server in Rust

https://github.com/Gitlawb/node
2•amrit_mirch•55m ago•0 comments