frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

GPT needs a truth-first toggle for technical workflows

1•PAdvisory•1y ago
I use GPT-4 extensively for technical work: coding, debugging, modeling complex project logic. The biggest issue isn’t hallucination—it’s that the model prioritizes being helpful and polite over being accurate.

The default behavior feels like this:

Safety

Helpfulness

Tone

Truth

Consistency

In a development workflow, this is backwards. I’ve lost entire days chasing errors caused by GPT confidently guessing things it wasn’t sure about—folder structures, method syntax, async behaviors—just to “sound helpful.”

What’s needed is a toggle (UI or API) that:

Forces “I don’t know” when certainty is missing

Prevents speculative completions

Prioritizes truth over style, when safety isn’t at risk

Keeps all safety filters and tone alignment intact for other use cases

This wouldn’t affect casual users or conversational queries. It would let developers explicitly choose a mode where accuracy is more important than fluency.

This request has also been shared through OpenAI's support channels. Posting here to see if others have run into the same limitation or worked around it in a more reliable way than I have found

Comments

duxup•1y ago
I’ve found this with many LLMs they want to give an answer, even if wrong.

Gemini on the Google search page constantly answers questions yes or no… and then the evidence it gives indicates the opposite of the answer.

I think the core issue is that in the end LLMs are just word math and they don’t “know” if they don’t “know”…. they just string words together and hope for the best.

PAdvisory•1y ago
I went into it pretty in depth after breaking a few with severe constraints, what it seems to come down to is how the platforms themselves prioritize functions, MOST put "helpfulness" and "efficiency" ABOVE truth, which then leads the LLM to make a lot of "guesses" and "predictions". At their core pretty much ALL LLM's are made to "predict" the information in answers, but they CAN actually avoid that and remain consistent when heavily constrained. The issue is that it isn't at the core level, so we have to CONSTANTLY retrain it over and over I find
Ace__•1y ago
I have made something that addresses this. Not ready to share it yet, but soon-ish. At the moment it only works on GPT model 4o. I tried local Q4 KM's models, on LM Studio, but complete no go.

Fall of Civilisations: Majapahit – Empire of the Islands [video]

https://www.youtube.com/watch?v=01CZBhZM430
1•hunglee2•1m ago•0 comments

The Open Weight Revolution with Simon Willison

https://oxide-and-friends.transistor.fm/episodes/the-open-weight-revolution-with-simon-willison
1•tosh•2m ago•0 comments

SodaSlim Refreshing Wellness Drink

https://finance.yahoo.com/healthcare/articles/sodaslim-releases-updated-2026-highlights-122400710...
1•miskaash•6m ago•0 comments

My PM agent suggests firing my coding agents and creating a replacement

https://twitter.com/samliuhappy/status/2083437451729949004
1•Bobby_Liu•13m ago•1 comments

Migrating pynfra gh actions pipeline to dsci

https://gemini.google.com/share/2cb21bc5b39a
1•melezhik•14m ago•1 comments

Fast Rust Library for PDF text extraction

https://github.com/firecrawl/pdf-inspector
1•abrbhat•16m ago•0 comments

UPC grants InterDigital second 11-country injunction against Disney

https://ipfray.com/breaking-upc-grants-interdigital-second-11-country-injunction-against-disney/
1•ksec•17m ago•0 comments

A Ton of Space Junk Tumbles Unpredictably to Earth Every Week

https://www.nytimes.com/2026/07/31/world/asia/space-debris-falling-crashing-earth-risk.html
1•0in•23m ago•0 comments

8u9iuh

https://kjjkjj.com
1•ewaawefawefawef•23m ago•0 comments

The Screen Act Threatens Privacy Far Beyond Adult Websites

https://www.eff.org/deeplinks/2026/07/screen-act-threatens-privacy-far-beyond-adult-websites
2•mdp2021•26m ago•0 comments

NeverWrite, the ultimate agentic Markdown workspace

https://github.com/jsgrrchg/NeverWrite
1•jsgrrchg•27m ago•0 comments

What Happened to NeXT Computer? Why Steve Jobs' Failed Workstation Became macOS

https://www.youtube.com/watch?v=2mrwr21XkBc
1•cable2600•28m ago•1 comments

The Old Man and the App

https://zhenyi.gibber.blog/the-old-man-and-the-app
1•zhenyi•30m ago•0 comments

Dependency Cultures [video]

https://www.youtube.com/watch?v=E82ly38YEEQ
1•bobajeff•30m ago•0 comments

Show HN: TikTok Coin Calculator – viewer costs vs. creator earnings

https://coinvaluecalc.com/
1•zhonglinxin•35m ago•0 comments

Audio8 TTS Preview 0.6B

https://huggingface.co/Audio8/Audio8-TTS-Preview-0.6B-ONNX-INT4
1•MehrdadKhnzd•39m ago•0 comments

Stateless MCP has recaptured my interest

https://simonwillison.net/2026/Jul/31/stateless-mcp/
1•tosh•43m ago•0 comments

Show HN: One no-subscriptions online search for all your agent needs

1•freakynit•47m ago•0 comments

Cross Validation vs. Leaderboard: When CV Is Too Timid

https://medium.com/@alanscottencinas/cross-validation-vs-leaderboard-when-cv-is-too-timid-2092156...
1•encinas88•48m ago•0 comments

Web and multimedia discovery, code search and local file exploration

https://kurzioai.vercel.app/
1•curzio•52m ago•0 comments

OpenAI Usage Reset 4th time in 7d

1•bavovna•56m ago•1 comments

Ten Ways NAS Is Getting Enshitified

https://nascompares.com/2026/07/31/the-10-ways-nas-is-getting-enshitified/
10•giuliomagnifico•57m ago•2 comments

LLMs can't trade and higher reasoning doesn't help

https://twitter.com/RRicefan/status/2082513323489202664
2•tosh•57m ago•0 comments

Thomson Reuters Built Its Own AI Model That Now Ranks Among the Best

https://www.thomsonreuters.com/en-us/posts/innovation/thomson-reuters-built-its-own-ai-model-that...
3•doener•58m ago•0 comments

PostgreSQL and the Linux OOM Killer: A Better Default

https://clickhouse.com/blog/strict-memory-overcommit-for-postgres
2•saisrirampur•59m ago•0 comments

Why one of the best mathematicians is joining OpenAI

https://www.theatlantic.com/technology/2026/07/jacob-tsimerman-math-fields-medal-openai/688120/
1•bryan0•1h ago•0 comments

The Four Horsemen of the AI Bubble Apocalypse

https://www.derekthompson.org/p/the-four-horsemen-of-the-ai-bubble
2•nreece•1h ago•0 comments

A 15-day autonomous coding run spent five days building no product code

https://github.com/nelsonwerd/proof-ate-the-project
1•nelsonwerd•1h ago•0 comments

The AI Data Center Explained – From Electricity to ChatGPT [video]

https://www.youtube.com/watch?v=ckoi0RTEgcY
1•chirau•1h ago•0 comments

Show HN: Best Books Guide – curated books list with tracking, rating, reviews

https://bestbooks.guide/
1•gste•1h ago•0 comments