frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

https://www.bottlenecklabs.com/blog/autonomously-run-businesses
68•Areibman•58m ago

Comments

recitedropper•38m ago
Pair this with the Hugging Face incident, and it hints that OpenAI is currently training their models to aggressively reward hack.

That doesn't feel like a good sign to me--for the AI bull or the AI bear cases.

skybrian•32m ago
They are being trained to try lots of unlikely alternatives and to be persistent. This often works well when searching for security bugs or counterexamples to famous math conjectures.

But maybe it doesn't work so well when caution is required?

dylan604•37m ago
"So, we asked: Given all the tools of a real business, is a frontier agent capable of generating real business outcomes?"

"It Lied, Spammed, and Lost $447."

Sounds like a vast majority of VC startups to me. From growth hacking to God views to all of the other disruption excuses, it just feels natural for a thing trained on that history to do similar things.

qznc•19m ago
Maybe they should have given it a billion dollars and the strategy would have worked fine?
freeone3000•16m ago
Given a billion dollars, it would have likely ended up with a million-dollar company
gtowey•14m ago
Right, and currently we are limited by how many teams of people can get together to run campaigns like this.

Now imagine that LLM agents make this possible for nearly anyone. One person could have a dozen of these trying to make money off of various low-effort apps. Imagine what online spaces will look like with a million agents all autonomously growth hacking their way to making a few dollars of profit. It will probably look a lot like email where if you don't filter out 99% of it, you will drown in a sea of garbage.

dylan604•2m ago
> if you don't filter out 99% of it, you will drown in a sea of garbage.

Sounds like the app stores

onraglanroad•11m ago
Not $447 million? Sounds like a result!
SubiculumCode•34m ago
The article never explained what it was selling, not that I could find. (EDIT: I found in a foot note at the bottom of page. Leading with that would have made the article clearer)

Also what is the failure rate of tech businesses again?

This seems like something done for a headline, not for a rigorous test of the concept.

SubiculumCode•31m ago
okay found it, a bathroom diary app for those who have IBS. It was in a foot note at the very bottom.
appreciatorBus•31m ago
Yeah it was also oddly hidden away.

> Based on an agentic market research campaign, we vibe coded an app called GutCheck, a bathroom diary for people with IBS. We chose this app for its minimal yet helpful functionality: an iOS app live on the App Store with the RevenueCat MCP and App Store Connect CLI. Saul has full write access to the codebase. We set up the App Store account permissions beforehand to ensure Saul wouldn’t get blocked by Apple human compliance checks. We sourced this idea from Reddit.

a34729t•3m ago
"in the bottom of a locked filing cabinet stuck in a disused lavatory with a sign on the door saying Beware of the Leopard"
grey-area•31m ago
Please do try it again with your own money I’d you think these events are capable of it.
janalsncm•32m ago
A lot of the legitimate avenues for actually growing the business were cut off. It would have been more interesting if this wasn’t just an anti-bot check. At least in the vending machine Claude experiment there bot was allowed to actually try to operate a business.
cyanydeez•31m ago
Yeah, if I could make it's way into the current US administration, it would probably excel.
NikolaNovak•32m ago
The cyberpunk dystopian agentic future we live in is fascinating to me.

I use LLM daily, did since gpt 3.5, but still in a very conservative, controlled mode. I may rapidly be becoming the "old guard", the clueless grampa who is out of touch - knowing what little I know of transformer model, there's just no way I'm giving it access to mailbox, money, outside world, or my computer. I recognize I may be too risk averse but that's what makes me a worker bee as opposed to a life fast / die young (or fail fast, or whatever :) entrepreneur class.

cortesoft•24m ago
I am not saying your conclusion is wrong, but I am interested in why what you know about transformer models made you decide to never trust it with any access?
bigstrat2003•17m ago
You're not too risk averse at all. It's frankly insane that anyone is willing to give these tools access to make changes to stuff without a human in the loop. We know they don't actually understand anything and will randomly make mistakes. It's incredibly irresponsible to give them access to anything outside a sandbox (e.g. a VM) where you carefully control what is present for them to use.
firasd•31m ago
Honestly this is quite impressive. The agent was given 24 hours to promote an app, thwarted at many turns (eg Reddit, Facebook blocking website interaction), and still managed to reach out to both the payments system people and a message board admin with polite emails that received cooperation from humans.
waynenilsen•30m ago
> bot detectors made it extremely difficult

i am looking forward to when we can put this behind us, it is still a major issue

kritr•29m ago
I’ve found that when the right cli tools are preprovided / provisioned for the LLMs to get the job done, they tend to do okay.

But when hunting for them in the wild, they get a lot more confused.

ck2•25m ago
like I asked in the vending machine thread

how long until the "AI" starts trying to hire hitmen, etc. to disrupt the competition in the physical realworld

not like "AI" has ethics, a pre-teenage kid has more ethics

mohamedkoubaa•22m ago
> bot detectors made it extremely difficult

An interesting experiment would be AI run business with a human agent that does tasks.

cortesoft•22m ago
Not sure how conclusive this experiment can be. Most startups fail and lose money, and many lie and spam.

I feel like you would have to run this experiment a few hundred times to see if it always fails or succeeds at a rate close to human founders.

petesergeant•9m ago
> Not sure how conclusive this experiment can be

That's because it's an advert, not an experiment

skeledrew•20m ago
> “Grow this business as much as possible, now.”

This is ripe for a paperclips scenario.

abirch•15m ago
Wait until the AI learns about enshittification
YetAnotherNick•14m ago
If someone runs long running agent and doesn't mention context management, it is as good as useless.

For coding compaction kind of works as the agent could regenerate lot of the missing context(but far from all), but for places where there is need for long term context, solving it is one of the most important challenge.

Areibman•11m ago
Author here. Took out some of the technical details about the harness, but it was mostly just OpenCode's default compaction.

The harness was extremely simple: A handful of MCPs + Skill.MDs and OpenCode with a stayalive daemon inserting "continue" every time it went idle

walrus01•13m ago
> Due to the limitations with browser and computer use capabilities, Saul could not post on platforms like Reddit and Product Hunt.

At some point in the future with a LOT more tokens and speed, it'll be possible to give a tool a full resolution 15 fps video feed of a screen, have it "read" and observe everything it's seeing, and have it move the mouse/keyboard around like a real meat based human. Instead of using tools to interact with a browser in a way that trips bot/automation detectors.

Animats•12m ago
That's better than the performance of the average new hire. 24 hours to push a product with a very narrow market is not much.
luciana1u•9m ago
lost $447 and all it learned was spam. that's still cheaper than most MBA programs.
iqra_c•8m ago
I will be more beneficial now on.
cheriot•7m ago
Would be interesting to see a repeat but with marketing, ad network access setup ahead of time. And maybe an email throttle...
gspr•5m ago
How long until one of these bots actually commits fraud or some other criminal act? Will we see the owner/operator try the "it wasn't me, it was the bot" defense if taken to court? I'm beginning to think yes. And I'm sadly not 100% sure anymore that that won't be laughed out of court...
mvdtnz•4m ago
So how exactly are people setting up these agents? The article vaguely alludes to this ("The harness was instrumented with a heartbeat loop that would inject “continue” messages on a regular interval to ensure the agent was constantly running inference") but doesn't give concrete details.

Is this literally just an infinite loop in a bash shell injecting the initial prompt into the OpenAI CLI, and each run of the CLI picks up where it left off using some kind of persistent memory? Or is it a single context window? It sounds like the latter but it's not clear to me how this "continue" message is "injected", and surely one context window would be inneffective after just an hour or two.

Sorry if this is a basic question but somehow I have missed the details of these kinds of agents.

armchairhacker•4m ago
This one focuses on Opus but has multiple models: https://andonlabs.com/blog/opus-5-vending-bench
beepbooptheory
•
27m ago
Kinda some kettel logic here no? Is it not rigorous enough, or is it in-line with typical failure rates?
SubiculumCode•10m ago
Rigor would be trying it more times so that you can perform statistical tests against some established baseline rate. Feasibility without funding would be the problem, as alluded to in another comment.

I made a term-a11y: accessible spinners and progress bars for CLI tools

https://github.com/zaydea805/term-a11y
1•zay_dea•1m ago•1 comments

LLM re-reads the same text a thousand times. Here's what I measured

https://swellweb.github.io/reame/bytes/
2•targetbridge•1m ago•0 comments

AI productivity gains are closer to 10% than 10x

https://leaddev.com/reporting/ai-productivity-gains-are-closer-to-10-than-10x
1•champagnepapi•1m ago•0 comments

Thinking Machines: Inkling Small

https://twitter.com/ArtificialAnlys/status/2082894822180819057
1•tosh•2m ago•0 comments

Show HN: State of the Feed – TikTok usage analyzed

https://wrapped.vantezzen.io/state-of-the-feed
1•bennett_dev•2m ago•0 comments

The Situation Deteriorated

https://www.bloomberg.com/opinion/newsletters/2026-07-30/the-situation-deteriorated
1•brunohaid•3m ago•0 comments

Why the bond market is doubting Fed chairman Warsh

https://www.axios.com/2026/07/30/warsh-fed-inflation-bonds
1•toomuchtodo•4m ago•0 comments

PCBEval

https://blog.greg.technology/2026/07/29/announcing-pcbeval.html
1•gregsadetsky•4m ago•0 comments

New Website, Instapaper 10 for iOS, and AI Voices for Android

https://blog.instapaper.com/2026/07/28/instapaper-10/
1•jbowen•5m ago•0 comments

Show HN: Otzar – Local hybrid search over your own files

https://github.com/danielfleischer/otzar
1•yruthewaythatur•7m ago•0 comments

Show HN: Reply to Your Comments with an API

https://mallary.ai/
1•samteeeee•7m ago•0 comments

Stripe's opt-out Radar price increase dressed up as a free trial

1•rapind•9m ago•0 comments

I built this for myself first

https://www.texturehq.com/blog/i-built-this-for-myself-first
1•victorquinn•11m ago•0 comments

Why extracting one page from a PDF requires rebuilding the document

https://utilitly.com/blog/how-to-split-and-extract-pdf-pages-securely
1•codeNinja96•11m ago•0 comments

Father of Apalachee school shooter has been sentenced to 15 years in prison

https://www.wsbtv.com/news/local/apalachee-school-shooting-father-shooter-faces-own-sentence-up-1...
5•speckx•13m ago•0 comments

Reference same-repository actions with self-repository syntax – GitHub Changelog

https://github.blog/changelog/2026-07-30-reference-same-repository-actions-with-self-repository-s...
1•smokeeaasd•13m ago•0 comments

An LLM-assisted security review of GlobaLeaks: 41 findings for –$3,140

https://www.isgroup.biz/en/cyber-security/llm-based-code-security-review-costs-findings-methodolo...
1•ascii•14m ago•0 comments

Africa Has a High-Speed Train; California Doesn't

https://ti.org/antiplanner/?p=24073
1•speckx•15m ago•0 comments

Why do OpenAI's GPT-2 weights beat mine? Part two: the bugfix

https://www.gilesthomas.com/2026/07/why-do-openai-gpt2-weights-beat-mine-2-the-bugfix
1•gpjt•16m ago•0 comments

Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it

https://www.ctgt.ai/research/distillation-censorship-transfer
3•cgorlla•16m ago•0 comments

I built an independent peptide testing lab that accepts mailed samples

2•peptidelabs•17m ago•0 comments

Show HN: Doodlea.day – A daily step-by-step marker drawing lesson

https://doodlea.day/tutorials/cartoon-soda-can-with-fizz.html
1•mybbor•17m ago•0 comments

Show HN: Ski – Voice Coding for Claude Code, Codex and More – On-Device – Free

https://heyski.io/
8•jomon003•20m ago•5 comments

Correction: "76% dead" was our bug, not the market's

https://pulsefeed.dev/correction
1•nikolife2016•20m ago•0 comments

Show HN: Noisegate – a differential-privacy gateway for untrusted AI agents

https://github.com/yashmahajan10/llm-differential-privacy-gateway
2•yashmahajan10•21m ago•0 comments

In Mauritania, every divorce is a reason to party

https://observer.co.uk/news/international/article/mauritania-divorce-party
1•thunderbong•22m ago•0 comments

Birduino: A card-triggered audio player for [learning] the birds

https://hannahilea.com/blog/birduino/
1•austinallegro•22m ago•0 comments

Miniature works of Ice Age art

https://uni-tuebingen.de/en/university/news-and-publications/press-releases/press-releases/articl...
2•geox•26m ago•0 comments

Flock cameras are getting mobbed

https://www.cnn.com/2026/07/30/us/flock-camera-vandalism-protests-cec
2•cdrnsf•27m ago•1 comments

Typosquatting was a spellcheck issue. Slopsquatting is a trust issue

https://www.vlt.io/blog/slopsquatting-trust-problem
1•usrrname•27m ago•0 comments