3 ms·
hey, builder here. The short version: LLMs take three tests: explain why jokes work (or don't), write jokes under shared premises or predict which jokes human
by yakshithk_ 23d ago
hey, builder here. The short version:
LLMs take three tests:
explain why jokes work (or don't), write jokes under shared premises or predict which jokes humans prefer
The finding so far that surprised me: every model aces explaining real jokes (95%+) but drops hard on explaining why a failed joke fails (81–92%).
happy to answer anything about the eval system and open to any sort of feedback!