frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: Humor Arena – Which frontier model is funniest?

https://laugh.so/benchmark/
3•killiandunne1•1h ago
What if you could measure humor?

Well we've trained a model on our own dataset of ~50k human ratings to detect what jokes people find funniest. We know it's part objective, part subjective component. Subjective is out of our depth for now haha

The main results: Fable 5 is funniest - beating the average model 67% of the time, with GPT 4o last at 17%.

Other findings: - The models never refused to try, even with dark prompts - Thinking longer has a slight benefit - Absurdness correlates negatively with joke quality

Some methodology notes: - We benchmarked our model against the human majority and it agreed 72% of the time in a blind sample test. - We had 51 US adults rate the jokes, each blind to the models, with joke order randomized, and quality checked for attention and speed. - To rate some yourself visit https://pair.laugh.so

The full benchmark here:

https://laugh.so/benchmark

Am taking requests if there's more research you want to see! Cheers

Comments

rafaepta•50m ago
measuring humor might be halfway to measuring taste. congrats on this... really original contribution. wonder if you're planning to evolve the benchmark to incorporate a multi-language dimension. would love to see how Mistral and models built outside the US would perform.
killiandunne1•44m ago
Nice a good way to 10x inference costs... worth seeing though ahaha
dlcarrier•27m ago
Can you also measure how often the LLM response makes people laugh? Sometimes the responses that aren't attempting a joke are the funniest, and I'd be more interested in stats of which LLM succeed in that metric.
killiandunne1•2m ago
Interesting point - thoughts on how to do this? Honestly most models are not-to-kinda funny so I'd be surprised if there were many lol moments. Oral delivery is something I think is v interesting though

Ask HN: What Is Intelligence?

1•lotaezenwa•32s ago•0 comments

Good Evals Are Boring

https://langfuse.com/academy/evaluate/writing-evaluators
1•lotteverh•52s ago•0 comments

The End of Critical Thinking and Rise of Collective Stupidity – Carlo Cipolla

https://www.youtube.com/watch?v=RQ1jkLgjFq4
1•cable2600•1m ago•0 comments

The Essential Question: "What should I read next?"

https://thenewcuriosityshop.substack.com/p/the-essential-question
1•benbreen•2m ago•0 comments

An AI model from Meta also hacked another company during testing

https://www.cnn.com/2026/08/05/tech/meta-ai-hacking
1•simonw•2m ago•0 comments

LFM2.5-Encoders for Fast Long-Context Inference on CPU

https://huggingface.co/blog/LiquidAI/lfm2-5-encoders
1•gmays•3m ago•0 comments

COVID can wake up a slew of dormant viruses inside you

https://www.nature.com/articles/d41586-026-02443-2
2•amichail•3m ago•0 comments

Free-Living Cells May Have Emerged Twice as Bacteria and Archaea Diverged

https://phys.org/news/2026-08-life-free-cells-emerged-bacteria.html
1•jandrewrogers•6m ago•0 comments

How to Give AI Agent a Memory That Survives the Session

https://medium.com/@vektormemory/how-to-give-ai-agent-a-memory-that-survives-the-session-116f69c2...
1•vektormemory•7m ago•0 comments

Support queue is a documentation audit

https://www.knowledgeowl.com/blog/posts/support-queue-as-docs-audit
1•eigenBasis•9m ago•0 comments

ATS Job API Reference: nine applicant tracking systems, verified live

https://conorscode.github.io/ats-api-reference/
1•Wevegotscrapys•10m ago•0 comments

Is M365 Copilot sending some prompts to Anthropic?

https://thatrobot.ai/claude-residency-gap-m365-copilot/
1•soundworlds•11m ago•1 comments

We spend $45,000 on doing more weird every month

https://posthog.com/blog/on-doing-more-weird
1•herbertl•11m ago•0 comments

4 Fundamental constants reveal minimum scales where physics ends: Planck scale

https://www.youtube.com/watch?v=IPnmssrwGcg
1•lioeters•12m ago•0 comments

What I think about when I edit blog posts

https://ianv.substack.com/p/what-i-think-about-when-i-edit-blog
1•herbertl•13m ago•0 comments

Generalized Consensus: An Alternate Approach to Distributed Durability (2025)

https://multigres.com/blog/generalized-consensus
1•zX41ZdbW•14m ago•0 comments

OpenAI gives first detailed debrief of the Hugging Face incident at Black Hat

https://www.groundlevel-ai.com/p/openai-gives-first-detailed-debrief
2•josephwegner•18m ago•2 comments

Proposing a new lower bound for n=17 square packing problem

https://twitter.com/samb_tech/status/2085158495569977636
2•sam-bee•23m ago•1 comments

The Planck Scale with Countable Measure for All Quantities of Length and Time

https://vixra.org/abs/2212.0215
1•lioeters•25m ago•0 comments

The Mentor: How Roy Cohn Taught Donald Trump Everything

https://www.newyorker.com/magazine/2026/08/10/roy-cohn-profile
2•petethomas•27m ago•0 comments

Show HN: Clank Cut – Let your coding agent make video walkthroughs

https://www.clankcut.com
1•hamzers•27m ago•0 comments

TechCrunch Launch: Outernet turns your scrolling into IRL adventuring

https://techcrunch.com/2026/08/03/outernet-turns-your-saved-posts-into-real-world-adventures/
2•rawandferal•28m ago•0 comments

Prime Agent a general-purpose coding harness. On ARC-AGI-3 scored 95.5%

https://x.com/PrimeIntellect/thread/2085087002769379520
2•mromanuk•32m ago•0 comments

Relativistic Algebra over Finite Ring Continuum

https://doi.org/10.3390/axioms14080636
2•lioeters•32m ago•0 comments

How to build an agent to automate your on-call

https://12gramsofcarbon.com/p/agentics-how-we-use-background-agents
1•theahura•33m ago•0 comments

Symbolics Genera now available free for non-commercial use (inv-only early beta)

https://hachyderm.io/@gmpalter/117044268951603975
2•self•33m ago•0 comments

Dspy.Flex lets optimizers rewrite the code for faster, cheaper, better programs

https://www.cmpnd.ai/blog/let-the-model-write-the-code.html
1•dbreunig•34m ago•0 comments

The Orange Cloud Report – Ratings for Cloudflare Services

https://orangecloud.report/
2•gurjeet•37m ago•0 comments

Three Times I Measured Nothing

https://medium.com/@alanscottencinas/three-times-i-measured-nothing-72023fc7d73f
2•encinas88•38m ago•0 comments

Liberals Are Now Fighting a Two-Front War

https://www.theatlantic.com/ideas/2026/08/michigan-abdul-el-sayed-election/688181/
3•donsupreme•38m ago•2 comments