3 ms·
This leaderboard is so funny when you look at Anthropic's Claude. Every version After 1.0 gets a worse score. It seems their vision of "alignment" does not ali
by Jackson__ 3y ago
This leaderboard is so funny when you look at Anthropic's Claude. Every version After 1.0 gets a worse score.
It seems their vision of "alignment" does not align with the users vision of a competent assistant even remotely.
- declaredapple 3y agoYeah it's mystifying to me that they because the #1 commercial OpenAI competitor and then decided to go all "Helpless and Harmless"
- frognumber 3y agoI don't think that's their business model. I would never use Claude for personal use. However, my employer wants to make sure nothing reputation-harming happens. That's a lot more important than model quality; everything is light years ahead of where we were four years ago. For a lot of applications which face customers / citizens / students / ..., NEVER screwing up is a lot more important than quality. For my own use, I prefer interacting with soul, humanity, attitude, and edge. For my employer's use, it's different. I think there's space for both. As an investor, I'd be bullish on both Anthropic and something with no safety built in. I'd be a lot less bullish on what's in between.
- declaredapple 3y agoI totally understand the need for an "on-rails" model. However I've found that using it as a personal assistant and for some automation tasks it suffers. The fact that it has only gotten worse as judged by humans doesn't inspire confidence for me. I could totally understand offering two options, but they're going so far it's actually unusable for many tasks. The famous case it refused to "kill a python process".
- frognumber 3y agoHere's the thing about business models in this domain: You don't win by generality or working for as many tasks as possible. You win by being the best at something. I will pick the best tool for the job I'm doing, be that writing product descriptions, conversational agent, or tech support. Runner-up has a chance -- for example, by lowering margins, or simply by being subsidized by investors in hopes of moving into #1. However, the difference between "unusable" and "fourth-best" is negligible in terms of business returns. I won't pick your product. There isn't a snowball's chance of "safe AI" being #1, or even #5, for what you want to use it for. It might as well be unusable. It needs to be #1 in the niche it's targeting. (The above isn't universal; there are places where bundling many types of functionality has synergy; this just isn't one of them).