4 ms·
Author here - I'm planning to create game versions of this benchmark, as well as my other multi-agent benchmarks (https://github.com/lechmazur/step_game https:/
by zone411 2y ago
Author here - I'm planning to create game versions of this benchmark, as well as my other multi-agent benchmarks (https://github.com/lechmazur/step_game https://github.com/lechmazur/step_game, https://github.com/lechmazur/pgg_bench/ https://github.com/lechmazur/pgg_bench/, and a few others I'm developing). But I'm not sure if a leaderboard alone would be enough for comparing LLMs to top humans, since it would require playing so many games that it would be tedious. So I think it would be just for fun.
- michaelgiba 2y agoI was inspired by your project to start making similar multi-agent reality simulations. I’m starting with the reality game “The Traitors” because it has interesting dynamics. https://github.com/michaelgiba/survivor https://github.com/michaelgiba/survivor (elimination game with a shoutout to your original) https://github.com/michaelgiba/plomp https://github.com/michaelgiba/plomp (a small library I added for debugging the rollouts)
- zone411 2y agoVery cool!