I voted on a few "Which one is funnier?" choices, but honestly none of them were funny. The em dashes in every joke already set the "another LLM slop" mood. I would suggest re-writing the prompts to drop dashes.
Suggestion #2: introduce basic quality checks. There was this one that talked about a 2020 event as if it was still indeterminate: "The date for Superbowl 2020 has been announced as Sunday, February 2 ... They haven't yet announced who the Patriots will be playing."
Suggestion #3: slider instead of A/B single choice. Sometimes I leaned towards one choice, but it still wasn't that funny to select it; a more nuanced scale would have helped there.
Thanks for building this. From the first glance, we humans are safe from machine AGI judging by joke quality today.
yakshithk_•55m ago
LLMs take three tests:
explain why jokes work (or don't), write jokes under shared premises or predict which jokes humans prefer
The finding so far that surprised me: every model aces explaining real jokes (95%+) but drops hard on explaining why a failed joke fails (81–92%).
happy to answer anything about the eval system and open to any sort of feedback!