4 ms·
Their currently available offerings are "meh", but they're definitely top 5 (Currently they're number 4). This leaderboard is the only one most people trust as
by declaredapple 3y ago
Their currently available offerings are "meh", but they're definitely top 5 (Currently they're number 4).
This leaderboard is the only one most people trust as it uses an ELO system with humans blindly comparing and rating them in the chatbot arena.
https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboard https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
Current leaders are OpenAI, Anthropic, Mistral, Google, 01-ai, and then various Llama finetunes and Llama itself (Meta)
- Jackson__ 3y agoThis leaderboard is so funny when you look at Anthropic's Claude. Every version After 1.0 gets a worse score. It seems their vision of "alignment" does not align with the users vision of a competent assistant even remotely.
- declaredapple 3y agoYeah it's mystifying to me that they because the #1 commercial OpenAI competitor and then decided to go all "Helpless and Harmless"
- frognumber 3y agoI don't think that's their business model. I would never use Claude for personal use. However, my employer wants to make sure nothing reputation-harming happens. That's a lot more important than model quality; everything is light years ahead of where we were four years ago. For a lot of applications which face customers / citizens / students / ..., NEVER screwing up is a lot more important than quality. For my own use, I prefer interacting with soul, humanity, attitude, and edge. For my employer's use, it's different. I think there's space for both. As an investor, I'd be bullish on both Anthropic and something with no safety built in. I'd be a lot less bullish on what's in between.
- declaredapple 3y agoI totally understand the need for an "on-rails" model. However I've found that using it as a personal assistant and for some automation tasks it suffers. The fact that it has only gotten worse as judged by humans doesn't inspire confidence for me. I could totally understand offering two options, but they're going so far it's actually unusable for many tasks. The famous case it refused to "kill a python process".
- frognumber 3y agoHere's the thing about business models in this domain: You don't win by generality or working for as many tasks as possible. You win by being the best at something. I will pick the best tool for the job I'm doing, be that writing product descriptions, conversational agent, or tech support. Runner-up has a chance -- for example, by lowering margins, or simply by being subsidized by investors in hopes of moving into #1. However, the difference between "unusable" and "fourth-best" is negligible in terms of business returns. I won't pick your product. There isn't a snowball's chance of "safe AI" being #1, or even #5, for what you want to use it for. It might as well be unusable. It needs to be #1 in the niche it's targeting. (The above isn't universal; there are places where bundling many types of functionality has synergy; this just isn't one of them).