3 ms·
This is a very interesting benchmark, and I think it has a lot of potential to make the speed of a model quantifiable. I'm often asking myself is it better to
by alembic_fumes 13d ago
This is a very interesting benchmark, and I think it has a lot of potential to make the speed of a model quantifiable.
I'm often asking myself is it better to use higher or lower effort levels, or to maybe drop down to a "dumber" but faster model. And so using a real-time based competition as a benchmark could shed some light on this, I think.
In this vein, here are what I would love to see added in this benchmark:
- Include Google's Gemini models. I keep hearing Gemini being praised for its speed, and I would like to see whether that gives it a big enough edge over the bigger but slower models.
- How does a Cerebras-accelerated open source model fare against a much larger but much slower frontier model?
I also feel like in general there is a lot of very low-hanging fruit to start benchmarking models across the spectrum of real-time vs batch-style workloads. Perhaps Brood War sits somewhere quite near the "real-time" end of the spectrum, but what about something like a game of speed chess, or a turn-based game with time limits?
I think what I would like to see the most is for someone to come up with a benchmark that supports tuning the "real-timeliness" of the benchmark, and then running a sweep of a model across the whole spectrum. That could get result in real nice graphs with multiple models on the pareto-frontier, varying based on the hosting provider and the model dimensions.