3 ms·
I've developed a benchmark that I think should be resistant to saturation, is easily verifiable, and anecdotally correlates with desirable behavior (ability to
by erikwiffin 2mo ago
I've developed a benchmark that I think should be resistant to saturation, is easily verifiable, and anecdotally correlates with desirable behavior (ability to not get confused while generating text with state).
I think it's interesting, I think other people would find it useful, but I don't want to spend a bunch of money running it against all the frontier models.
What's the best way to reach out to labs like yours to collaborate on something like that? Are there any labs that are more open to submissions from internet randos?
- gertlabs 2mo agoSlimeBallBench? Looks cool, love the implementation! You would most likely need a LOT of environments like these if your goal was selling them to labs
- erikwiffin 2mo agoSlimeBallBench was not me, but it is cool! and at least partially what inspired me to build my own. Mine is less fun, but hopefully more useful https://erikwiffin.github.io/memory-reasoning-eval/ https://erikwiffin.github.io/memory-reasoning-eval/ Selling to labs is more than I'm looking for. I'm aiming for a couple hundred dollars so I don't have to finance a Fable vs Sol run out of my own pocket. It would be cool to have my benchmark be one of the ones referenced in a model card!