3 ms·
Someone should make a censorship/alignment (whatever you want to call it) benchmark for LLMs.
by CrypticShift 3y ago
Someone should make a censorship/alignment (whatever you want to call it) benchmark for LLMs.
- thierrydamiba 3y agohttps://tatsu-lab.github.io/alpaca_eval/ https://tatsu-lab.github.io/alpaca_eval/ Such a leaderboard exists, AlpacaEval Leaderboard ranks LLMs on the ability to follow user instructions.