frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

GPT needs a truth-first toggle for technical workflows

1•PAdvisory•1y ago
I use GPT-4 extensively for technical work: coding, debugging, modeling complex project logic. The biggest issue isn’t hallucination—it’s that the model prioritizes being helpful and polite over being accurate.

The default behavior feels like this:

Safety

Helpfulness

Tone

Truth

Consistency

In a development workflow, this is backwards. I’ve lost entire days chasing errors caused by GPT confidently guessing things it wasn’t sure about—folder structures, method syntax, async behaviors—just to “sound helpful.”

What’s needed is a toggle (UI or API) that:

Forces “I don’t know” when certainty is missing

Prevents speculative completions

Prioritizes truth over style, when safety isn’t at risk

Keeps all safety filters and tone alignment intact for other use cases

This wouldn’t affect casual users or conversational queries. It would let developers explicitly choose a mode where accuracy is more important than fluency.

This request has also been shared through OpenAI's support channels. Posting here to see if others have run into the same limitation or worked around it in a more reliable way than I have found

Comments

duxup•1y ago
I’ve found this with many LLMs they want to give an answer, even if wrong.

Gemini on the Google search page constantly answers questions yes or no… and then the evidence it gives indicates the opposite of the answer.

I think the core issue is that in the end LLMs are just word math and they don’t “know” if they don’t “know”…. they just string words together and hope for the best.

PAdvisory•1y ago
I went into it pretty in depth after breaking a few with severe constraints, what it seems to come down to is how the platforms themselves prioritize functions, MOST put "helpfulness" and "efficiency" ABOVE truth, which then leads the LLM to make a lot of "guesses" and "predictions". At their core pretty much ALL LLM's are made to "predict" the information in answers, but they CAN actually avoid that and remain consistent when heavily constrained. The issue is that it isn't at the core level, so we have to CONSTANTLY retrain it over and over I find
Ace__•1y ago
I have made something that addresses this. Not ready to share it yet, but soon-ish. At the moment it only works on GPT model 4o. I tried local Q4 KM's models, on LM Studio, but complete no go.

F*: A general-purpose proof-oriented programming language

https://fstar-lang.org/
1•ducktective•30s ago•0 comments

An exam leak in India exposed a Gen Z jobs crisis that goes much deeper

https://www.cnbc.com/2026/08/02/india-exam-leak-protests-jobs-crisis-gen-z-unemployment-modi.html
1•cramer4next•4m ago•0 comments

The future is for everyone [video]

https://www.youtube.com/watch?v=JlRgRmAqoAc
1•firstSpeaker•4m ago•0 comments

When Verification Explores Too Far: LLM Test Coverage vs. Validity

https://zenodo.org/records/21758550
1•haitamk•5m ago•0 comments

Show HN: DesktopAudio, a tool to record what your Mac plays

https://desktopaudio.app/
1•AquiGorka•5m ago•0 comments

The TEMU-Fication of Software, Digital Goods and Services

https://xn--gckvb8fzb.com/the-temu-fication-of-software-digital-goods-services/
2•wrxd•10m ago•0 comments

Show HN: Handoff is await human() for AI agents

https://github.com/OmegaAgent/handoff
1•LivingGlitcher•10m ago•1 comments

Transformer Models in Financial Forecasting: Outperforming LSTMs

https://algo-finance.com/ai/machine-learning/transformer-models-financial-forecasting-2026/
1•anonymoussala•11m ago•0 comments

Velocity a proof of linear-scaling long context for existing LLMs, no retraining

https://github.com/Veloresearch/velocity-mta-proof
1•Veloresearch•14m ago•0 comments

An AI TikTok Shop Slop Factory That Shills Supplements the FDA Recalled

https://www.404media.co/inside-an-ai-tiktok-shop-slop-factory-that-shills-supplements-recalled-by...
1•Dfol•14m ago•0 comments

The Simple Elegance of the Integrated Timing Belt Loopback Fastener

https://danielmangum.com/posts/integrated-timing-belt-loopback-fastener/
1•hasheddan•17m ago•0 comments

Show HN: Type a formula, watch it become Clique, Hamiltonian Cycle, and Knapsack

https://aidoctrine.github.io/np-complete-universe/
1•AlekseN•18m ago•0 comments

Don't Take the Black Pill (Text Adaptation)

https://andrewkelley.me/post/dont-take-black-pill.html
1•Tomte•20m ago•0 comments

Shadscan – Deterministic UI Audits for Shadcn Apps

https://github.com/TheOrcDev/shadscan
1•javatuts•22m ago•0 comments

MapLibre GL JavaScript – WebGL library for interactive vector maps

https://maplibre.org/maplibre-gl-js/docs/
1•javatuts•22m ago•0 comments

Will Self-Hosted Local AI Go Mainstream, or Stay a Niche for Nerds?

https://grigio.org/will-self-hosted-local-ai-go-mainstream-or-stay-a-niche-for-nerds/
1•grigio•23m ago•0 comments

JavaScript Obfuscator

https://www.jstools.space/js-obfuscator/
1•javatuts•23m ago•0 comments

GPUs could explode to multiple TB with new storage-inspired memory tech

https://www.theregister.com/storage/2026/07/30/gpus-could-explode-to-multiple-tb-with-new-storage...
3•jacquesm•25m ago•0 comments

An internal OpenAI Astra model solved 10 major open math and CS problems

https://twitter.com/polynoamial/status/2083467194663571701
10•wa5ina•30m ago•0 comments

I'm (mostly) picking models on speed now, not intelligence

https://martinalderson.com/posts/speed-vs-intelligence/
1•martinald•32m ago•0 comments

BroMetal: Typed shaders compiled at build time to WebGPU

https://brometal.dev
1•davedx•35m ago•0 comments

Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

https://website-auditor.io/ai-visibility-index
4•SpikeyCoder•37m ago•1 comments

When did 'dupe culture' take over?

https://news.darden.virginia.edu/2026/08/01/qa-when-did-dupe-culture-take-over/
1•bookofjoe•38m ago•0 comments

Show HN: LyricVibe – synced lyrics for YouTube Music, Spotify, SoundCloud

https://chromewebstore.google.com/detail/lyricvibe-—-kinetic-lyric/iplhipppjgpeofmadbnkolnpigia...
1•riteshzz•41m ago•0 comments

Best way to avoid bloat and AI – selfcontaining OS

4•bialamusic•44m ago•0 comments

Getopt() but Friendlier

https://www.unix.dog/~yosh/blog/getopt-but-friendlier.html
2•fanf2•49m ago•0 comments

Meshdiff – visually compare two STL versions in the browser, client-side

https://meshdiff.com/
27•projscope•57m ago•2 comments

Size of Xbox price increase in Europe shocks

https://www.notebookcheck.net/Size-of-Xbox-price-increase-in-Europe-shocks-giving-Sony-s-PS5-cons...
2•HelloUsername•58m ago•1 comments

Pdf-inspector: Rust lib for PDF inspection, classification, and text extraction

https://github.com/firecrawl/pdf-inspector
2•7777777phil•58m ago•1 comments

Show HN: Stripe for Wearable Data

https://stridee.fit/developer
2•alvaromolina0•59m ago•1 comments