4 ms·
I disagree. Not saying the other benchmarks are better. It just depends on your use case and application. For my use of the chat interface, I don't think lmsys
by __jl__ 2y ago
I disagree. Not saying the other benchmarks are better. It just depends on your use case and application.
For my use of the chat interface, I don't think lmsys is very useful. lmsys mainly evaluates relatively simple, low token count questions. Most (if not all) are single prompts, not conversations. The small models do well in this context. If that is what you are looking for, great. However, it does not test longer conversations with high token counts.
Just saying that all benchmarks, including lmsys, have issues and are focused on specific use cases.