llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle"
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...4.2266 cents, 38 seconds.
For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat.
UPDATE: Here's another one with five pelicans for each of the five Muse Spark 1.3 reasoning levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
The most expensive was reasoning level xhigh - 7.5 cents, 1m34s.
And I ran five pelicans at all reasoning levels for 1.2 as well, here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Definitely an upgrade over 1.2
The 2D / flat ground feels reasonable for a SVG, which implies a vector illustration.
Definitely shows how important a user data flywheel is for RL and model improvement.
Given this is Meta, my immediate assumptions that one is cheap because it lets me "be the product". I know I'm rushing to conclusions but there is zero trust here. The brain will do its thing. And the wording here is giving the brains a lot of wiggle room.
I don't see the wiggle room at all.
I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What I mean by that is that it did what I told it and if there was something unexpected in the code that it put out it was often because I gave it ambiguous or conflicting instructions. It didn't try to go above and beyond and just acted like a tool, which is what I want from a coding agent 90%+ of the time. I also felt that it did a much better job of following established patterns in my code than many of the other current models do. I'm a huge fan of OpenAI's models and Spark 1.2 is what I expected 5.6 Luna to be.
I'm curious and a little excited to use 1.3, but honestly a little worried that as Meta pushes for better benchmarks that Spark will start to fall into the trap of trying to be "helpful" in ways I don't want it to be.
Tangential, but when I first started using Spark 1.2, it made me realize how much I miss 5.3 Codex. That model was the peak of coding models, IMO, in that it knew how to write good code, but didn't try to overstep or be "helpful" in unexpected ways. That got me thinking about how the major labs seem to be stepping away from coding focused models toward more general purpose ones and how I can't help but feel like that's a mistake.
This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.
Good job Meta! Seriously. This is almost making me forget about the 18B$ lawsuit for children social media addiction.
Lmao. And their benchmark table only shows max reasoning.
I feel the same about Grok w/ Elon. I will pay extra to use someone else.
I'm not an Amodei stan, but of all of these people he seems to have the most ethical focus. Again, not everything done perfectly and I have my gripes, but of the leaders of frontier labs, I'll vote with my money.
And, yeah, I wouldn't trust sama to watch my bag while I went to the bathroom.
That's pretty much 90% of HN these days.
Apple releases a new iPhone? Here comes the flood of decade-old complaints about long-discontinued Mac butterfly keyboards and walled gardens.
Microsoft releases a new version of Windows? Here come the gripes about Azure.
Google changes something in GMail? Play Store!
It's like there's an army of bots out there determined to reduce the productivity of the Western tech bubble by diverting everyone into endless circular arguments about absolutely nothing of relevance to the topic at hand.
> Meta announces they have a new model, demonstrating its capabilities.
> Parent comment states „regardless of this model‘s specific capabilities, if I can avoid it I will.“
The model seems on par with Sol and Opus 5 on paper (admittedly on some older/saturated benchmarks, but very competitive for $).
Stats:
1M context, $0.10 input/$0.002 cached, $0.20 output (Mtok)
(It's probably going to be a bunch of repetitive batch jobs like web search that have no training value)
It's like when you ask your average person off the street to draw a house - it'll almost always be square with a triangle roof, one door, and two windows.
In the pelican/bike example, it's probably a bit of a self-perpetuating snowball too. If the earliest examples were bike left-to-right, flat ground, etc. then they are also being scraped up in future LLMs.
There is this scene in the HBO series Westworld where a "host" says some words in sequence which is shown on a display as she says it. Of course, even me thinking of this scene and connecting it to your comment was not original, someone else clearly had the same programming as me.
A medium blog post says
> Pair what with me?” — the moment Maeve (a humanoid android) uttered those words in Westworld (Season 1, Episode 6: “The Adversary”), something clicked. Not for the average viewer, but for me, a STEM educator and AI enthusiast who, just weeks earlier, had read Stephen Wolfram’s seminal essay, What Is ChatGPT Doing … and Why Does It Work?
(Why the drivetrain is on the right, I don't know. But most bike parts follow open standards so it's quite entrenched.)
While I'm sure this factors into things for advertisements for bike components, there is also just a general preference that westerners have for left-to-right motion. Not just in bike ads, but all ads with (or suggesting) movement. And also not just ads, but movies where directors believe left-to-right motion is associated with progression and right-to-left motion is regressive.
Because they're computers. They don't have an imagination and the ability to create things from whole cloth the way humans do.
Much like a mother pelican, they regurgitate what they've been fed.
Also 3X token use vs. 1.2
These benchmarks you guys invent for yourselves prove nothing.
A small model like Mistral 7b is just as useful to the end task as any larger model, if not more so because it’s faster.
Your big model may be able to draw pelicans or solve some esoteric nonsense but it cannot do real work in the real world.
These phoney benchmarks and experiments mean nothing.
None of the LLMs can replace a software engineer nor even a barista or car mechanic etc. not even close.
Instead of inventing fake benchmarks do something tangible and tell me how it performs.
Before laying off half the country and going full retard on AI
Thank you for doing this, I love your benchmark the most!
They all suck. Pick your poison.
Here's the order, from best to worst.
Amodei
SamA
Zuck
Elon
I think a less personal ranking would be, as a business owner, which of those providers is more dependable? As in, you don't care about evil, just your stuff working. I think maybe OpenAI?
frozenseven•1h ago
https://news.ycombinator.com/item?id=49541149