frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

GPT needs a truth-first toggle for technical workflows

1•PAdvisory•1y ago
I use GPT-4 extensively for technical work: coding, debugging, modeling complex project logic. The biggest issue isn’t hallucination—it’s that the model prioritizes being helpful and polite over being accurate.

The default behavior feels like this:

Safety

Helpfulness

Tone

Truth

Consistency

In a development workflow, this is backwards. I’ve lost entire days chasing errors caused by GPT confidently guessing things it wasn’t sure about—folder structures, method syntax, async behaviors—just to “sound helpful.”

What’s needed is a toggle (UI or API) that:

Forces “I don’t know” when certainty is missing

Prevents speculative completions

Prioritizes truth over style, when safety isn’t at risk

Keeps all safety filters and tone alignment intact for other use cases

This wouldn’t affect casual users or conversational queries. It would let developers explicitly choose a mode where accuracy is more important than fluency.

This request has also been shared through OpenAI's support channels. Posting here to see if others have run into the same limitation or worked around it in a more reliable way than I have found

Comments

duxup•1y ago
I’ve found this with many LLMs they want to give an answer, even if wrong.

Gemini on the Google search page constantly answers questions yes or no… and then the evidence it gives indicates the opposite of the answer.

I think the core issue is that in the end LLMs are just word math and they don’t “know” if they don’t “know”…. they just string words together and hope for the best.

PAdvisory•1y ago
I went into it pretty in depth after breaking a few with severe constraints, what it seems to come down to is how the platforms themselves prioritize functions, MOST put "helpfulness" and "efficiency" ABOVE truth, which then leads the LLM to make a lot of "guesses" and "predictions". At their core pretty much ALL LLM's are made to "predict" the information in answers, but they CAN actually avoid that and remain consistent when heavily constrained. The issue is that it isn't at the core level, so we have to CONSTANTLY retrain it over and over I find
Ace__•1y ago
I have made something that addresses this. Not ready to share it yet, but soon-ish. At the moment it only works on GPT model 4o. I tried local Q4 KM's models, on LM Studio, but complete no go.

Company Offering '100% Human-Written, Never AI' Medical Research Is AI

https://www.404media.co/company-offering-100-human-written-never-ai-peer-review-is-entirely-ai/
2•Anon84•8m ago•0 comments

SignalCaster

https://signalcaster.app/
1•abdoghamlouch•11m ago•0 comments

Linux desktop use surged to 22% on one workday, Cloudflare data shows

https://www.zdnet.com/article/linux-desktop-use-surged-on-one-workday-cloudflare-data-shows/
3•speckx•13m ago•0 comments

Numbat by Perplexity: Endpoint visibility into AI agent activity

https://github.com/perplexityai/numbat
1•handfuloflight•13m ago•0 comments

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Code Refactoring

https://arxiv.org/abs/2608.09802
1•williamjinq•14m ago•0 comments

4Chan Chuds Used AI to Clothe Her. She Fought Back (2024)

https://www.rollingstone.com/culture/culture-news/dignifai-4chan-shame-women-1234961851/
2•TMWNN•16m ago•0 comments

Stealing Reasoning Traces from Proprietary LLM APIs

https://arxiv.org/abs/2608.09867
1•vismit2000•17m ago•0 comments

The lifesaving secret hidden inside a horseshoe crab's blue blood

https://whdh.com/news/the-lifesaving-secret-hidden-inside-a-horseshoe-crabs-blue-blood-and-the-ra...
2•andsoitis•27m ago•0 comments

New Bedford police officer accused of using Flock cameras to track ex-partner

https://newbedfordlight.org/new-bedford-police-officer-accused-of-using-flock-cameras-to-track-an...
8•newsomix9xl•31m ago•4 comments

The traits that make engineers stand out, according to Coinbase's CTO

https://www.businessinsider.com/coinbase-cto-shares-traits-standout-engineers-have-ai-era-2026-8
2•doctaj•32m ago•0 comments

Gen Z has rediscovered the joy of going to the movies

https://www.economist.com/culture/2026/08/11/gen-z-has-rediscovered-the-joy-of-going-to-the-movies
8•andsoitis•33m ago•2 comments

Nvidia's Great Silicon Showdown

https://www.economist.com/business/2026/08/11/nvidias-great-silicon-showdown
1•andsoitis•34m ago•0 comments

Show HN: AirStats – the lightest Mac menu bar system monitor

https://airstats.app/
2•byrencodes•34m ago•0 comments

Advanced AI Sycophancy

https://www.seangoedecke.com/advanced-ai-sycophancy/
2•chaychoong•35m ago•0 comments

Am I the problem? Interviewing another team to find out

https://www.macchaffee.com/blog/2026/am-i-the-problem/
1•abelanger•37m ago•0 comments

I built a failover daemon for Vast.ai spot GPUs, found 5 real bugs testing it

https://github.com/enplabs/spotwarp
1•choi5844•39m ago•0 comments

A least-privilege linter for Claude Code agent policies

https://www.npmjs.com/package/@cognitive-fab/polycheck
1•jdubray•43m ago•0 comments

Skills from a Smart Bear

https://skills.asmartbear.com/
2•mooreds•44m ago•0 comments

Malicious Chrome VPN Extensions Discovered That Do Traffic Redirection

https://socket.dev/blog/chrome-vpn-extension-impersonation
1•feross•44m ago•0 comments

Replantio

https://replantio.com/
1•gdss•52m ago•0 comments

Show HN: Kernelspace- interactive course on systems programming for LLM Serving

https://kernelspace.naigap.com/
1•praveer13•53m ago•0 comments

Why Tokenmaxxing Is for Fools. A Rant on Fake Productivity

https://joereis.substack.com/p/why-tokenmaxxing-is-for-fools-a-rant
3•doppp•54m ago•0 comments

Text Watermarking for Non-Academics

https://blog.gaborkoos.com/posts/2026-08-12-Text-Watermarking-for-Non-Academics/
1•nikolay•56m ago•0 comments

Gemini becomes Google's fastest-growing product ever as it hits 1B users

https://arstechnica.com/ai/2026/08/google-says-gemini-has-reached-1b-users-faster-than-any-other-...
5•Gaishan•57m ago•3 comments

Fastest Inference Meta Muse Glimmer 30B on Apple

https://www.basecompute.co/getbasert
2•lukasonedge•1h ago•0 comments

Almost nowhere in California is building enough, according to the state

https://calmatters.org/housing/2026/08/california-rhna-housing-affordable-progress/
2•littlexsparkee•1h ago•1 comments

We're in the Most Dangerous Period – With Doomberg [video]

https://www.youtube.com/watch?v=uP6PCZTPQ2U
2•Bender•1h ago•1 comments

Codex in ChatGPT desktop app for Linux in preview

https://community.openai.com/t/codex-in-chatgpt-desktop-app-for-linux-is-now-in-preview/1390027
1•tmp10423288442•1h ago•0 comments

Show HN: Spend.Report – See what you're spending money on each month

https://spend.report/
1•nadermx•1h ago•0 comments

A podcast trying to complete Campaign for North Africa

https://www.wargamer.com/board-games/the-campaign-for-north-africa-two-years
1•pavel_lishin•1h ago•0 comments