4 ms·
Show HN: Humor Arena – Which frontier model is funniest?
What if you could measure humor?
Well we've trained a model on our own dataset of ~50k human ratings to detect what jokes people find funniest.
We know it's part objective, part subjective component.
Subjective is out of our depth for now haha
The main results: Fable 5 is funniest - beating the average model 67% of the time, with GPT 4o last at 17%.
Other findings:
- The models never refused to try, even with dark prompts
- Thinking longer has a slight benefit
- Absurdness correlates negatively with joke quality
Some methodology notes:
- We benchmarked our model against the human majority and it agreed 72% of the time in a blind sample test.
- We had 51 US adults rate the jokes, each blind to the models, with joke order randomized, and quality checked for attention and speed.
- To rate some yourself visit https://pair.laugh.so https://pair.laugh.so
The full benchmark here:
https://laugh.so/benchmark https://laugh.so/benchmark
Am taking requests if there's more research you want to see! Cheers
- deleted 2mo ago[deleted]
- maxsich 2mo ago[dead]
- rafaepta 2mo agomeasuring humor might be halfway to measuring taste. congrats on this... really original contribution. wonder if you're planning to evolve the benchmark to incorporate a multi-language dimension. would love to see how Mistral and models built outside the US would perform.
- killiandunne1 2mo agoNice a good way to 10x inference costs... worth seeing though ahaha
- dlcarrier 2mo agoCan you also measure how often the LLM response makes people laugh? Sometimes the responses that aren't attempting a joke are the funniest, and I'd be more interested in stats of which LLM succeed in that metric.
- killiandunne1 2mo agoInteresting point - thoughts on how to do this? Honestly most models are not-to-kinda funny so I'd be surprised if there were many lol moments. Oral delivery is something I think is v interesting though
- dlcarrier 2mo agoYou'll probably still get plenty of smirks, smiles, and silent chuckles. OpenCV or a similar machine vision system should be able to pick those out from a webcam aimed at a reviewer.
- deleted 2mo ago[deleted]
- maxzhdev 2mo agoI can’t imagine how to correctly assess the ability to humor, because this is a very subjective and relative phenomenon. The work ahead is serious
- killiandunne1 2mo agoWe don't wanna be too serious ;)