3 ms·
These models do pretty well in benchmarks and real world so I'm highly suspicious of this article. Further more, in the original report, the examples of bad ans
by simianwords 5d ago
These models do pretty well in benchmarks and real world so I'm highly suspicious of this article. Further more, in the original report, the examples of bad answers are from Haiku - at least 7 out of 10. Anyone who knows anything about LLMs know that haiku shouldn't be used for anything pretty much.
There's no reproducible set either. I'm not gonna trust this report.
- stymaar 5d agoMost people[1] interacting with chatbots don't have a paid subscription and they do interact with the free-tier LLMs that are Luna and Haiku, so I still think it's relevant. [1]: not on HN obviously, but IRL, and probably among FT's readership as well.
- cillian64 5d agoA free claude account with no subscription gets you access to sonnet and I believe uses it by default over haiku