frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Are AI Labs Pelicanmaxxing?

https://dylancastillo.co/posts/pelicanmaxxing.html
60•dcastm•1h ago

Comments

dcchambers•43m ago
It's incredible that each model has it's own style that remains relatively consistent throughout all of the different generated examples.
andy99•42m ago
If an AI researcher was going to pelicanmaxx, they would almost certainly apply the augmentations mentioned in the article during training, e.g. randomly selecting animals and conveyances. You’d want a model that generalizes well, just sfting in that specific prompt would be pretty bush league for a frontier lab.

I don’t have any reason to believe they are gaming the benchmark, just saying. I do find the idea of a data labeller having to generate thousands of svgs of different animals on different modes of transportation quite funny though.

cute_boi•19m ago
At this point, I think there are so many pelican images in the pretraining data that drawing a pelican no longer makes sense as a model evaluation task.
johndough•36m ago
Another point for consideration: Specialized SVG models create way better looking pelicans riding a bicycle. (E.g. Refract V4: https://jumpshare.com/s/8liB7Aiuoo3yucbWGXjZ mirror: https://postimg.cc/McV70p84 )
solarkraft•21m ago
That’s an impressive image, but what a mistake it was to click the second link (on mobile without an ad blocker). I wouldn’t send it to anyone I respect ...
ACCount37•13m ago
The name is "Recraft V4", and from looking it up: yeah, it sure seems like whatever black magic they use for SVG generation kicks ass.
sbseitz•33m ago
I wish I could downvote this for Pelicanmaxxing lmao.
tomas789•32m ago
Having an objective score is quite difficult. Maybe it would be better to do a pairwise comparison and calculate ELO?
javier123454321•12m ago
If you want to, go ahead, but it seems to me the author already exceeded the energy expenditure that this question warranted.
Wowfunhappy•25m ago
> The more plausible story is SVGmaxxing

Exactly--and you have to ask yourself at this point what "maxxing" really means, since "get better at drawing SVGs" is a useful skill.

beering•6m ago
Really awful how the AI labs are skillmaxxing /s

Pelicans aside, we need to remember that benchmarks are the only good quantitative way we have of comparing models. If someone has complaints about “benchmaxxing”, please ask them to contribute a better benchmark! It is valuable work and very appreciated.

cute_boi•21m ago
https://playcode.io/blog/macbook-svg-benchmark

I think we should stop using pelican benchmark.

stusmall•19m ago
I'm glad someone ran the numbers on this. Every single Simon Willison post of an SVG is followed with someone dismissing it saying "I'm sure they train on it by now." This is despite a good blog post with sound logic on how easy that is to catch. [1] Glad to see someone took the time for a quantitative analysis of dumb little animals riding dumb little bikes.

1. https://simonwillison.net/2025/Nov/13/training-for-pelicans-...

jonatron•17m ago
OK, so we've done animal_vehicle, how about new SVG ideas each time? I just tried "make an SVG of a man sitting in a chair at a computer behind a desk" which gives more interesting results than the animalVehicle test.
ninju•3m ago
There probably good set of images of that description already so it does exercise the inference capability of the model
j45•12m ago
The models definitely seem to pay attention to the tests.

Since the tests can be generally gamed with directing descriptions at it non-deterministically, there's a greater chance the questions solution can be found.

Of course, hopefully the models are instead adding patterns and types of questions as well and it makes the models more capable, but it may be limited in how it transfers to other types of questions in breadth or depth.

simonw•4m ago
This is fantastic

I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations.

Catching a lab cheating specifically on my one dumb benchmark would be really funny.

Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than anything I was considering.

His conclusion:

> Nothing jumped out at me. I couldn’t find a case where the pelican-bicycle images looked noticeably better than the rest of that model’s grid.

Rooster61•4m ago
I find it humorous that the animal + plane combo appears to be such an outlier. I assume this is due to the models assuming the user mean plain and misspelled it in the prompt.

Terrence Tao's ChatGPT Conversation about the Jacobian Conjecture Counterexample

https://chatgpt.com/share/6a5fdc7a-d6f8-83e8-bbea-8deb42cfed56
219•gmays•1h ago•106 comments

GigaToken: ~1000x faster Language model tokenization

https://github.com/marcelroed/gigatoken/
122•syrusakbary•1h ago•24 comments

Show HN: Bento - An entire PowerPoint in one HTML file (edit+view+data+collab)

https://bento.page/slides/
429•starfallg•3h ago•91 comments

Making

https://beej.us/blog/data/ai-making/
166•erikschoster•3h ago•65 comments

The startup's Postgres survival guide

https://hatchet.run/blog/postgres-survival-guide
210•abelanger•6h ago•111 comments

Are AI Labs Pelicanmaxxing?

https://dylancastillo.co/posts/pelicanmaxxing.html
61•dcastm•1h ago•18 comments

Nobody knows what a used GPU cluster is worth

https://ciphertalk.substack.com/p/nobody-knows-what-a-used-gpu-cluster
50•rbanffy•1w ago•39 comments

Can a MUD evaluate LLMs? A $99 proof of concept

https://cruciblebench.ai/
55•Davisb135•3h ago•27 comments

Mechanical light bulb from 1675 [video]

https://www.youtube.com/watch?v=0Y-9GbsS9Fg
40•kadohg•1w ago•14 comments

Launch HN: Unlayer (YC W22) – Add email and document builders to your app

https://unlayer.com
27•adeelraza•3h ago•19 comments

Perlin's Noise Algorithm

https://blog.jaysmito.dev/blog/02-perlins-noise-algorithm/
48•ibobev•4h ago•8 comments

Show HN: HN Hall of Fame – browse 3,100 legendary Hacker News links

https://www.orangecrumbs.com/hall/
144•oyster143•3h ago•32 comments

10 REM"_(C2SLFF4

https://beej.us/blog/data/mystery-comment/
129•ingve•7h ago•35 comments

“We have information that Moonshot distilled Fable for the development of K3”

https://twitter.com/mkratsios47/status/2079933645888880708
119•softwaredoug•4h ago•263 comments

OpenNode – Bitcoin Payment Processor

https://opennode.com/
89•gurjeet•4h ago•76 comments

Introduction to Formal Verification with Lean Part 1

https://hashcloak.com/blog/tutorial-introduction-to-formal-verification-with-lean-(part-1)
206•badcryptobitch•3d ago•40 comments

Neo Radar: A browser-based orbital mechanics engine with 41k real asteroids

https://neoradar.space
44•daviazpen•3h ago•10 comments

Show HN: DeepSQL – A self-hostable DBA agent for Postgres and MySQL

https://deepsql.ai/
19•venkat971•2d ago•15 comments

Passkeys were invented by engineers with zero understanding of consumer brain

https://twitter.com/nikitabier/status/2079787406300266743
293•ksec•4h ago•381 comments

I tried to record music onto a cassette tape using modern tech

https://swiftrocks.com/i-tried-to-record-music-onto-a-cassette-tape-using-modern-tech
15•rockbruno•5d ago•7 comments

Ghost Cut – or why Cut and Paste is broken everywhere

https://ishmael.textualize.io/blog/ghost-cut/
61•willm•4h ago•48 comments

Everyone Should Know SIMD

https://mitchellh.com/writing/everyone-should-know-simd
13•WadeGrimridge•1h ago•0 comments

Hologram works. Elixir runs in the browser

https://hologram.page/blog/backing-hologram
66•lawik•4h ago•8 comments

Show HN: Web swing through midtown NYC

https://www.swingnyc.com/
38•shahahmed•2h ago•15 comments

When Is NVLink Worth It?

https://platform-fools.com/posts/2026-04-27-nvlink/
31•ak_t•4h ago•5 comments

Intel Starts Shipping High-NA EUV Silicon

https://morethanmoore.substack.com/p/intel-starts-shipping-high-na-euv
213•zdw•3d ago•90 comments

Cornell's Interactive Wall of Birds

https://academy.allaboutbirds.org/features/wallofbirds/?_hsmi=428996456
84•yareally•3d ago•31 comments

Airbus Full Scale Foldable Wing Extensions

https://www.airbus.com/en/newsroom/press-releases/2026-07-airbus-launches-new-flight-test-program...
51•r2sk5t•4h ago•45 comments

Original Apollo 11 Guidance Computer source code for command and lunar modules

https://github.com/chrislgarry/Apollo-11
160•noteness•13h ago•50 comments

Most Americans say "not in my backyard" to AI data centers

https://www.redfin.com/news/ai-data-centers-opposition-education-benefit/
103•toomuchtodo•4h ago•216 comments