Genuinely happy with some of the Qwen 3.8 results (especially since I can run that model at Q8).
Interesting to see how much better (at this task) Pi (OMP) is over Opencode as a harness.
I’d love to see a few more with outcomes that are as easy to judge but less subjective.
I’ve got a toy project going to make a fun to watch battle simulator where an LLM (or two if playing vs) has to write programs that control multiple bots (each with their own line of sight and limited battle context) that have to coordinate and fight alongside each other. Goal is to have the LLM update the code based on current situations maybe 5-10 times in a 5 min simulated battle. Exploring even allow the bots to request new programming and score based on number of reprogram steps.
Ideally also finding somehow (not sure what would be the right away) what is publicly available before running the test. It's quite a different outcome if there are competitions, e.g. js13k, live code examples from books, even templates, on specific that topic. Visually here the results looks very very similar to the point that I can't help but wonder if it's the result from the short yet relatively descriptive prompt or because some template was always found and relied on.
hanspagel•32m ago
arecsu•27m ago
utopiah•26m ago
If you don't have anything working check the console, maybe a WebGL issue.