frontpage.
newsnewestaskshowjobs

Made with ♥ by @iamnishanth

Open Source @Github

Open in hackernews

Show HN: DesignArena – crowdsourced benchmark for AI-generated UI/UX

https://www.designarena.ai/
46•grace77•5h ago
I’ve been using AI to generate some repetitive frontend (guilty), and while most outputs felt vibe-coded, some results were surprisingly good. So I cleaned it up and made a ranking game out of it with friends, and you can check it out here: https://www.designarena.ai/vote

/vote: Your prompt will be answered by four random, anonymous models. You pick the one you prefer and crown the winner, tournament-style.

/leaderboard: See the current winning models, as dictated by voter preferences.

/play: Iterate quickly by seeing four models respond to the same input and pressing space to regenerate the results you don’t lock-in.

We were especially impressed with the quality of DeepSeek and Grok, and variance between categories (To judge by the results so far, OpenAI is very good for game dev, but seems to suck everywhere else).

We’ve learned a lot, and are curious to hear your comments and questions. Excited to make this better!

Comments

coryvirok•4h ago
This is really good! It would be really cool to somehow get human designs in the mix to see how the models compare. I bet there are curated design datasets with descriptions that you could pass to each of the models and then run voting as a "bonus" question (comparing the human and AI generated versions) after the normal genAI voting round.
grace77•4h ago
wow this is a super interesting idea, and the team loves it — we'll fast follow-through and follow-up here when we add it, thanks for the suggestion!
debesyla•1h ago
This would be extra interesting for unique designs - something more experimental, new. As as for now even when you ask AI to break all rules it still outputs standard BS.
a2128•4h ago
I tried the vote and both results always suck, there's no option to say neither are winners. Also it seems from the network tab you're sending 4 (or 5?) requests but only displaying the first two that respond, which biases it to the small models that respond more quickly which usually results in showing two bad results
ethan_smith•4h ago
Adding a "neither is good" option would improve data quality by preventing forced choices between two poor designs.
grxxxce•4h ago
this is a great note — will be sure to add!
grace77•3h ago
Yes — great point. We originally waited for all model responses and randomized the vote order, but that made it a very bad user experience -- some models, especially open-source ones, took over 4 minutes to respond, leading to a high voter drop-off rate.

To preserve the voter experience without introducing bias, our current approach waits for the slowest model within each binary comparison — so even if one model is faster, we don’t display until both are ready. You're right that this does introduce some bias for the two smallest models, and we'd love to hear suggestions for how to make this better!

As for the 5th request: we actually kick off one reserve model alongside the four randomly selected for the tournament. This backup isn’t shown unless one of the four fails — it’s not the fastest or lowest-latency model, just a randomly selected fallback to keep the system robust without skewing results.

adi_hn07•3h ago
Awesome web app !! This is loads of fun and gives an idea about which models are best for which type of UI design.

Do launch on https://www.superlaun.ch for more traffic and exposure for your web app.

grace77•3h ago
thank you! posting now :)
adi_hn07•2h ago
Thanks !!
justusm•2h ago
nice! Training models using reward signals for code correctness is obviously very common; I'm very curious to see how good things can get using a reward signal obtained from visual feedback
grace77•2h ago
As are we, seems like the natural next step
muskmusk•1h ago
This is a surprisingly good idea. The model vs model is fun, but not really that useful.

But this could be a legitimate way to design apps in general if you could tell the models what you liked and didn't like.

grace77•1h ago
yes! that is the hope — /play is our first attempt at building out utility, would love your feedback and will ship hard to make it happen!
butz•20m ago
What is this new trend of building websites without scrollbars? Was this website also made by GenAI?
grace77•2m ago
I wish—just added them back

The UX Psychology Glossary

https://builtformars.com/ux-glossary
1•pipase•2m ago•0 comments

Wine 10.12 (Dev) – Run Windows Applications on Linux, BSD, Solaris and macOS

https://gitlab.winehq.org/wine/wine/-/releases/wine-10.12
1•neustradamus•3m ago•0 comments

We Should Stop Making AI Look Human

https://medium.com/@ravelantunes/why-we-should-stop-making-ai-look-human-4c78cdeee806
1•ravelantunes•6m ago•0 comments

Long Google

https://loeber.substack.com/p/27-long-google
1•wahlrus•7m ago•0 comments

Senolytic Update

https://www.science.org/content/blog-post/senolytic-update
1•EA-3167•8m ago•0 comments

Autonomous robot surgeon removes organs with 100% success rate

https://newatlas.com/robotics/worlds-first-robot-surgery/
1•AlexDragusin•8m ago•0 comments

Most (ly Dead) Influential Programming Languages (2020)

https://www.hillelwayne.com/post/influential-dead-languages/
1•azhenley•10m ago•0 comments

Bcachefs Lands Fixes in Linux 6.16 for Some "High Severity" Regressions

https://www.phoronix.com/news/Bcachefs-Fixes-Linux-6.16-rc6
2•Bender•19m ago•0 comments

Show HN: Planking for penguins, a real-time exercise tracking game

https://twitter.com/measure_plan/status/1944127683241078937
1•getToTheChopin•20m ago•0 comments

FairLight TV #127, $D011 Mayhem [video]

https://www.youtube.com/watch?v=0KhH4YrUsXo
1•sagacity•20m ago•0 comments

First U.S. Rare Earth Mine in 70 Years Opens in Wyoming

https://cowboystatedaily.com/2025/07/11/first-u-s-rare-earth-mine-in-70-years-opens-in-wyoming/
1•Bender•21m ago•0 comments

5 big EV takeaways from Trump’s “One Big Beautiful Bill”

https://www.wired.com/story/5-big-ev-takeaways-one-big-beautiful-bill/
1•Bender•23m ago•0 comments

Using AMD MI300X for High-Throughput, Low-Cost LLM Inference

https://www.herdora.com/blog/the-overlooked-gpu
2•technoabsurdist•29m ago•0 comments

Explaining 6 Levels of Automated Driving and Which Ones Are Actually on US Roads

https://www.jalopnik.com/1903079/self-driving-levels-explained/
1•rntn•33m ago•0 comments

Show HN: I Built a Stick-On Wireless Lamp That Installs in 30 Seconds

https://www.shopinfinitylamp.store/product/infinity-wall-lamp
1•kingvyn•34m ago•0 comments

Corn after soy: New study quantifies rotation benefits and trade-offs

https://phys.org/news/2025-06-corn-soy-quantifies-rotation-benefits.html
2•PaulHoule•34m ago•0 comments

Scientists Are Sneaking Passages into Research Papers to Trick AI Reviewers

https://www.msn.com/en-us/news/us/scientists-are-sneaking-passages-into-research-papers-designed-to-trick-ai-reviewers/ar-AA1IrpuI
1•galaxyLogic•34m ago•1 comments

YouTube Piano – Play It with Your Computer Keyboard

https://www.youtube.com/watch?v=3gZC5763wYk
1•busymom0•34m ago•0 comments

Shipping Linear Drafts

https://mufeezamjad.com/blog/linear-drafts
1•ishan0102•35m ago•0 comments

Aeron: Efficient reliable UDP unicast, UDP multicast, and IPC message transport

https://github.com/aeron-io/aeron
2•todsacerdoti•35m ago•0 comments

Super Easy* 2-Stage Git Deployment

https://ratfactor.com/cards/super-easy-2-stage-git-deployment
1•ingve•36m ago•0 comments

Cooling a Raspberry Pi Device [pdf]

https://pip.raspberrypi.com/categories/685-whitepapers-app-notes/documents/RP-003608-WP/Cooling-a-Raspberry-Pi-device.pdf
1•sandwichsphinx•37m ago•0 comments

Mcbot McHacked

https://captaincompliance.com/education/mcdonalds-ai-hiring-bot-hacked-with-123456-password-exposing-millions-of-job-seekers/
1•richartruddie•38m ago•1 comments

Ask HN: Can "Pull Request to Get Hired" Replace Traditional Tech Hiring?

1•jerawaj740•39m ago•2 comments

My Foray into Vlang

https://kristun.dev/posts/my-foray-into-vlang/
2•hggh•40m ago•0 comments

Show HN: Build web forms in rich text

https://kameo.dev
1•nistuley•41m ago•0 comments

Student Wins $250k Prize in Regeneron Science Talent Search

https://pasadenanow.com/main/pasadena-high-school-student-wins-250000-top-prize-in-regeneron-science-talent-search
1•geox•42m ago•0 comments

AI Agent Marketplace

https://aetheragentforge.org
1•cgibson2025•48m ago•0 comments

Angr (open-source binary analysis platform for Python)

https://angr.io/
2•Bogdanp•50m ago•0 comments

They Fled War in Ethiopia. Then American Bombs Found Them

https://www.nytimes.com/2025/07/12/world/middleeast/ethiopia-yemen-american-bombs.html
1•perihelions•55m ago•0 comments