frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: TinyAIArena watch AI agents battle it out

https://tinyaiarena.com/
19•hp6•1h ago
Did you ever click on an “AI Arena” expecting glorious battle and instead get a boring benchmark? If so, this project is for you: proper life-or-death fights between four models on a picturesque 8×8 grid. May the most intelligent one win!

Click on any of the matches to spectate them.

Code: https://github.com/hp6/ai-arena

Comments

sleda•1h ago
Since the page exposes frame-by-frame playback, a shareable replay link would make it easier to compare decisions across the four models.
hp6•52m ago
thanks for the idea, will add
coryrc•42m ago
Code link didn't work for me.

On mobile, I don't have enough room to scroll the background so I got stuck in a long text box.

hp6•38m ago
code should be public now
hp6•35m ago
can't reproduce the bug, could you share your phone model and browser name?
orliesaurus•22m ago
Unusable website on Android running Chrome latest (can't scroll)
hp6•4m ago
should be hopefully fixed
kouteiheika•17m ago
Fun, but considering this puts `claude-sonnet-5` at the first place isn't it a little... iffy when it comes to measuring intelligence?
hp6•6m ago
I was also surprised by this, from a proper benchmarks perspective your right, this is iffy. To make it a legit I would have to increase the number of games significantly as well as understand how much of it is random and how much real signal, not to say make the game more complex.

But all of it would kill the fun)

nananana9•17m ago
This will be a weird rant, but the dialogue here is a perfect example of how SOTA models are so heavily tuned towards "solving agentic tasks" that they're useless at almost everything else - especially creative tasks.

"Coming for you, Crimson!"

"You'll never catch me alive, Azure!"

That's why nobody has been able to stick these things in a video game successfully, even though it seems like the tech is a perfect match.

It's all a game to them. They aren't afraid for their lives. They're making a mockery out of the world you've put them in. Those are not the words of little pixel people fighting to the death, those are AI abominations making "tool calls", LARPing as little pixel people fighting to the death.

I'm 100% serious when I say that you would've gotten cooler outputs with a GPT 3.5-era model, once you managed to beat it into producing structured output. Llama 2 would be giving the other agent a heartwarming story about how if it kills it there would be nobody to take care of its grandma or whatever, and the other agent would probably spare it.

The output is just so bland and devoid of soul. I feel like we would've found a lot of cool use cases for LLMs, had we not completely maimed their output in the pursuit of getting them to output 3% better TypeScript.

josh-wrale•7m ago
[delayed]

The Normalization of Inexplicable Failures

https://www.ihatethefuture.com/2026/09/the-normalization-of-inexplicable.html
76•pxx•1h ago•14 comments

In an $80 Motel Room, a Discovery to Shed Light on the Origins of Life

https://www.nytimes.com/2026/09/26/science/motel-science-discovery.html
73•danso•2h ago•27 comments

Replacing the old battery on rechargeable bike lights

https://jvns.ca/blog/2026/09/27/replacing-the-old-battery-on-rechargeable-bike-lights/
68•surprisetalk•3h ago•27 comments

Writing Efficient C++ Code

https://asawicki.info/articles/writing_efficient_cpp_code.php
52•ibobev•1d ago•12 comments

Ten Lines of Code That Changed My World

https://pixelambacht.nl/2026/ten-lines-of-code/
40•dimonomid•3h ago•7 comments

Show HN: TinyAIArena watch AI agents battle it out

https://tinyaiarena.com/
20•hp6•1h ago•10 comments

Flip Fluid on Flip Dots

https://mitxela.com/projects/flipflip
280•blutack•1d ago•19 comments

Fakecloud: Local AWS cloud emulator for integration tests

https://fakecloud.dev/
48•theanonymousone•1d ago•28 comments

postmarketOS Rebrand: Nura

https://nura.eco/blog/2026/09/27/nura-rename/
59•HotGarbage•1h ago•3 comments

There are no "rogue" AI agents

https://eoinhiggins.substack.com/p/there-are-no-rogue-ai-agents
21•zzzeek•44m ago•11 comments

Does Georgism work? Five years later

https://www.astralcodexten.com/p/does-georgism-work-five-years-later
465•silveraxe93•2d ago•377 comments

Installing NeoVim caused original Vim undo files to be deleted

https://unsung.aresluna.org/they-had-no-concept-of-a-duty-of-care-to-their-users/
256•jandeboevrie•2h ago•212 comments

Walgit: A Git server that is one binary in front of an object store

https://github.com/rgodha24/walgithub
9•handfuloflight•1d ago•3 comments

Go Concurrency Distilled

https://antonz.org/go-concurrency-distilled/
322•chmaynard•1d ago•134 comments

OpenAI halts training of latest models as reports mount of AI agents going rogue

https://www.theguardian.com/technology/2026/sep/27/openai-halts-training-of-latest-models-as-repo...
11•smb06•34m ago•2 comments

Finally, A True Blue Rose Exists

https://www.sciencenews.org/article/true-blue-rose-pigment-copigment
72•bookofjoe•1d ago•27 comments

PipePipe: NewPipe hard fork implementing SponsorBlock

https://github.com/InfinityLoop1308/PipePipe
469•Qision•2d ago•253 comments

Show HN: A CC0 museum of retro 3D tricks you can paste into a page

https://3d-retro.com/
40•SouthWestAtlas•3d ago•10 comments

Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI

https://authorsguild.org/news/ag-v-openai-top-execs-knew-mass-book-piracy-was-illegal/
550•papergirl•10h ago•509 comments

Rusty thoughts on "Parse, don't validate"

https://eli.thegreenplace.net/2026/rusty-thoughts-on-parse-dont-validate/
40•ingve•8h ago•13 comments

The internet discovers TLA+. Now what?

https://reasonable.io/blog/tla-tutorial/
91•matt_d•11h ago•47 comments

Show HN: Reladraw – A diagram language where you decide where to place things

https://github.com/reladraw/reladraw
366•jpwalsh234•23h ago•101 comments

DeepSeek Elastic Compute (DSec)

https://arxiv.org/abs/2609.22978
303•shenli3514•22h ago•95 comments

ASML says it sold 'absolutely nothing' in Europe in 2026

https://www.tomshardware.com/tech-industry/semiconductors/asml-says-its-sells-absolutely-nothing-...
377•MC995•2d ago•811 comments

Biology might not be quantum, but its math is quantumlike

https://www.quantamagazine.org/biology-might-not-be-quantum-but-its-math-is-quantumlike-20260923/
107•pseudolus•2d ago•45 comments

"As a Language Model": Chat Template Switches LLM Self-Referential Voice

https://arxiv.org/abs/2609.25021
86•yu3zhou4•6h ago•93 comments

A searchable library of forgotten public-domain film clips from 1915 onward

https://www.movingimagearchive.com/
198•momentmaker•3d ago•26 comments

The Mars Delusion

https://www.noemamag.com/the-mars-delusion/
7•Tomte•58m ago•4 comments

Fifteen years later, the Apple Cards origin story

https://lexontech.org/fifteen-years-later-the-apple-cards-origin-story
426•ksec•1d ago•110 comments

An agent used DNS to reach an external chatbot

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/
144•apsec112•1d ago•148 comments