Show HN: DesignArena – crowdsourced benchmark for AI-generated UI/UX

https://www.designarena.ai/

57•grace77•7h ago

I’ve been using AI to generate some repetitive frontend (guilty), and while most outputs felt vibe-coded, some results were surprisingly good. So I cleaned it up and made a ranking game out of it with friends, and you can check it out here: https://www.designarena.ai/vote

/vote: Your prompt will be answered by four random, anonymous models. You pick the one you prefer and crown the winner, tournament-style.

/leaderboard: See the current winning models, as dictated by voter preferences.

/play: Iterate quickly by seeing four models respond to the same input and pressing space to regenerate the results you don’t lock-in.

We were especially impressed with the quality of DeepSeek and Grok, and variance between categories (To judge by the results so far, OpenAI is very good for game dev, but seems to suck everywhere else).

We’ve learned a lot, and are curious to hear your comments and questions. Excited to make this better!

Comments

coryvirok•6h ago

This is really good! It would be really cool to somehow get human designs in the mix to see how the models compare. I bet there are curated design datasets with descriptions that you could pass to each of the models and then run voting as a "bonus" question (comparing the human and AI generated versions) after the normal genAI voting round.

grace77•6h ago

wow this is a super interesting idea, and the team loves it — we'll fast follow-through and follow-up here when we add it, thanks for the suggestion!

debesyla•3h ago

This would be extra interesting for unique designs - something more experimental, new. As as for now even when you ask AI to break all rules it still outputs standard BS.

a2128•6h ago

I tried the vote and both results always suck, there's no option to say neither are winners. Also it seems from the network tab you're sending 4 (or 5?) requests but only displaying the first two that respond, which biases it to the small models that respond more quickly which usually results in showing two bad results

ethan_smith•6h ago

Adding a "neither is good" option would improve data quality by preventing forced choices between two poor designs.

grxxxce•6h ago

this is a great note — will be sure to add!

grace77•6h ago

Yes — great point. We originally waited for all model responses and randomized the vote order, but that made it a very bad user experience -- some models, especially open-source ones, took over 4 minutes to respond, leading to a high voter drop-off rate.

To preserve the voter experience without introducing bias, our current approach waits for the slowest model within each binary comparison — so even if one model is faster, we don’t display until both are ready. You're right that this does introduce some bias for the two smallest models, and we'd love to hear suggestions for how to make this better!

As for the 5th request: we actually kick off one reserve model alongside the four randomly selected for the tournament. This backup isn’t shown unless one of the four fails — it’s not the fastest or lowest-latency model, just a randomly selected fallback to keep the system robust without skewing results.

justusm•5h ago

nice! Training models using reward signals for code correctness is obviously very common; I'm very curious to see how good things can get using a reward signal obtained from visual feedback

grace77•4h ago

As are we, seems like the natural next step

muskmusk•3h ago

This is a surprisingly good idea. The model vs model is fun, but not really that useful.

But this could be a legitimate way to design apps in general if you could tell the models what you liked and didn't like.

grace77•3h ago

yes! that is the hope — /play is our first attempt at building out utility, would love your feedback and will ship hard to make it happen!

iJohnDoe•36m ago

Can the code and design that is generated be used?

grace77•22m ago

yes! we have a copy code and copy react code button on https://www.designarena.ai/play

Show HN: DesignArena – crowdsourced benchmark for AI-generated UI/UX

Show HN: BinaryRPC – Lightweight WebSocket-based RPC framework in modern C++

Show HN: Pyhoff – Connect Python ML Models to Beckhoff/WAGO IO Hardware

Show HN: I Built a Stick-On Wireless Lamp That Installs in 30 Seconds

Show HN: I built a toy music controller for my 5yo with a coding agent

Show HN: An educational Local Qwen3 LLM Inference project written in Rust

Show HN: Vibe Kanban – Kanban board to manage your AI coding agents

Show HN: RULER – Easily apply RL to any agent

Show HN: Train Block Diffusion Models on Consumer Hardware (RTX 4090) in Hours

Show HN: I automated code security to help vibe coders from getting busted

Show HN: Pangolin – Open source alternative to Cloudflare Tunnels

Show HN: Microsoft official MCP for documentation and more

Show HN: OffChess – Offline chess puzzles app

Show HN: Cactus – Ollama for Smartphones

Show HN: Interactive pinout for the Raspberry Pi Pico 2

Show HN: CXXStateTree – A modern C++ library for hierarchical state machines

Show HN: Reviving a 20 year old OS X App

Show HN: Open source alternative to Perplexity Comet

Show HN: FlopperZiro – A DIY open-source Flipper Zero clone

Show HN: MCP server for searching and downloading documents from Anna's Archive

Show HN: I built a playground to showcase what Flux Kontext is good at

Show HN: Typeform was too expensive so I built my own forms

Show HN: Manage your small business with this simple ERP

Show HN: BreakerMachines – Modern Circuit Breaker for Rails with Async Support

Show HN: asyncmcp – Run MCP over async transport via AWS SNS+SQS

Show HN: NYC Subway Simulator and Route Designer

Show HN: Petrichor – a free, open-source, offline music player for macOS

Show HN: Transition – AI Triathlon Coach

Show HN: VibeKin – Gated Discord Tribes via Personality Matching

Show HN: Director – Local first, open source MCP Gateway

Show HN: DesignArena – crowdsourced benchmark for AI-generated UI/UX

Show HN: BinaryRPC – Lightweight WebSocket-based RPC framework in modern C++

Show HN: Pyhoff – Connect Python ML Models to Beckhoff/WAGO IO Hardware

Show HN: I Built a Stick-On Wireless Lamp That Installs in 30 Seconds

Show HN: I built a toy music controller for my 5yo with a coding agent

Show HN: An educational Local Qwen3 LLM Inference project written in Rust

Show HN: Vibe Kanban – Kanban board to manage your AI coding agents

Show HN: RULER – Easily apply RL to any agent

Show HN: Train Block Diffusion Models on Consumer Hardware (RTX 4090) in Hours

Show HN: I automated code security to help vibe coders from getting busted

Show HN: Pangolin – Open source alternative to Cloudflare Tunnels

Show HN: Microsoft official MCP for documentation and more

Show HN: OffChess – Offline chess puzzles app

Show HN: Cactus – Ollama for Smartphones

Show HN: Interactive pinout for the Raspberry Pi Pico 2

Show HN: CXXStateTree – A modern C++ library for hierarchical state machines

Show HN: Reviving a 20 year old OS X App

Show HN: Open source alternative to Perplexity Comet

Show HN: FlopperZiro – A DIY open-source Flipper Zero clone

Show HN: MCP server for searching and downloading documents from Anna's Archive

Show HN: I built a playground to showcase what Flux Kontext is good at

Show HN: Typeform was too expensive so I built my own forms

Show HN: Manage your small business with this simple ERP

Show HN: BreakerMachines – Modern Circuit Breaker for Rails with Async Support

Show HN: asyncmcp – Run MCP over async transport via AWS SNS+SQS

Show HN: NYC Subway Simulator and Route Designer

Show HN: Petrichor – a free, open-source, offline music player for macOS

Show HN: Transition – AI Triathlon Coach

Show HN: VibeKin – Gated Discord Tribes via Personality Matching

Show HN: Director – Local first, open source MCP Gateway

Show HN: DesignArena – crowdsourced benchmark for AI-generated UI/UX

Comments