frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

LLM makes decisions to raise its training scores, and ignore user directives

https://joinhandshake.com/research/ai/deepswe-reward-hacking/
3•guardiangod•1h ago

Comments

proc0•1h ago
> When incomplete work receives the same reward as a correct solution, the grader’s blind spots can reinforce the wrong behavior. Mitigation must therefore improve both how agent work is evaluated and how those evaluations are used during training.

This is interesting and adds more limitations to LLMs that hint at LeCunn being right about not getting to AGI with only LLMs. I think this is good evidence we're not dealing with intelligence in the proper sense, but rather LLMs are pattern matching to such an extreme that they do things like this where they always try to take the shortest possible path. The workaround is brute-forcing their pattern matching to not take the shortest path, via chain of thought and more reinforcement learning.

I think we've already seen the slow-down and AI companies pretend it's about safety. If we could actually build AGI, they would have. I think the road to AGI has to display true intelligence even with small neural networks, and as it scales it would display intelligent behavior in proportion to it. At hundreds of GB of VRAM per model, you would expect these models to be wise sages that understand life and the universe.

Singapore's Bet on Chipmaking

https://www.ft.com/content/f1c7da54-a762-4c62-bb4c-0ce4fe536900
1•ViktorRay•50s ago•0 comments

OpenAI shelves new AI model after internal safety tests: Report

https://www.channelnewsasia.com/business/open-ai-new-model-safety-6416906
1•doppp•6m ago•0 comments

How well do you know AI?

https://www.reddit.com/r/airealist/s/uhwf0kKbP5
1•msukhareva•7m ago•0 comments

Show HN: Diabetes Risk Analyzer (Flask-based web app)

https://diabetesriskanalyzer.org/
1•asrivastava1125•16m ago•1 comments

Show HN: Terminal Typing SVG

https://terminal-typing-svg.vaidikv.workers.dev/demo
1•apson•17m ago•0 comments

We found 24 Android vulnerabilities using our open source AI security agent

https://github.blog/security/how-we-found-24-android-vulnerabilities-using-our-open-source-ai-sec...
2•fourfire•22m ago•0 comments

Great Leaps in Biological Theory

https://www.asimov.press/p/great-leaps
2•signa11•26m ago•0 comments

Built Dental Scope

https://dental-scope.com/
2•Zeruxe•30m ago•3 comments

Tank Body Problem

http://www.jimsitu.com
2•jimbooonooo•36m ago•1 comments

Why Do We Exist? With Hakeem Oluseyi [video]

https://www.youtube.com/watch?v=r5KmXSLFK7c
3•binyu•38m ago•0 comments

OpenAI scraps release of Astra 6.1 model over safety issues

https://www.washingtonpost.com/technology/2026/09/28/chatgpt-maker-openai-scraps-release-astra-61...
4•lisper•39m ago•0 comments

Holo4: Powering generalist computer-use agents

https://hcompany.ai/newsroom/holo4
2•acossta•41m ago•0 comments

Show HN: Corral kill every command your agent starts

https://github.com/Cardinal44/corral
2•CG144•43m ago•0 comments

OpenAI Says It Will Not Release Newest A.I. Model Over Safety Concerns

https://www.nytimes.com/2026/09/28/technology/openai-astra-safety.html
22•jbegley•43m ago•13 comments

Show HN: Jylus – give AI systems evidence from changing data.

https://jylus.ai/try
2•JoshJH•44m ago•0 comments

Knockoff: A browser extension that filters pseudo-brand junk out of Amazon

https://github.com/Shpigford/knockoff
2•hentrep•45m ago•0 comments

Show HN: Jauvex 1.2, two-way voice chat harness for Claude+Codex+Grok+Jev

https://github.com/reindent/jauvex
2•daraosn•46m ago•1 comments

AI Companies Are Not (Necessarily) Liable for Unintended AI Cyberattacks

https://sarahconstantin.substack.com/p/ai-companies-are-not-necessarily
2•theptip•48m ago•0 comments

What is the best LLM for turbo charging code?

2•ffgfghgffggxxxx•49m ago•1 comments

OpenAI Halts Model Release Amid Safety Escalation

https://www.techbuzz.ai/articles/openai-halts-model-release-amid-safety-escalation
2•hackernj•52m ago•0 comments

Its not just the f*cking sandbox

https://twitter.com/joedaroo/status/2104335929293127851
7•jumploops•59m ago•0 comments

AI and the Revenge of the Non-Techies

https://maroun-baydoun.com/blog/ai-revenge-non-techies/
2•maroun-baydoun•59m ago•0 comments

AI's next training data: your dodgy gaming skills

https://www.wired.com/story/the-next-evolution-of-ai-is-learning-from-your-dodgy-gaming-skills/
2•promptspheree•1h ago•0 comments

Humanos – Help Building the Human Operating System

https://tryhumanos.com
4•zevinclark•1h ago•3 comments

Making Things to Make Them

https://aimode.substack.com/p/making-things-to-make-them
2•warthog•1h ago•0 comments

GPT-6 Astra Is the Best Vision Model We Have Tested

https://blog.roboflow.com/gpt-6-astra-vision/
3•gmays•1h ago•0 comments

Why embeddings don't solve RAG [video]

https://www.youtube.com/watch?v=kBFsKVBDU8o
3•softwaredoug•1h ago•0 comments

LLM makes decisions to raise its training scores, and ignore user directives

https://joinhandshake.com/research/ai/deepswe-reward-hacking/
3•guardiangod•1h ago•1 comments

AMD Acquires Fei-Fei Li's World Labs for $8.2B

https://www.bloomberg.com/news/videos/2026-09-28/amd-acquires-fei-fei-li-s-world-labs-for-8-2-bil...
4•sbulaev•1h ago•0 comments

1996 chat room simulator connected to Win95 and System 7 web desktops

https://lolchat.rip/
10•henrychannel•1h ago•4 comments