3 ms·
How accurate do you think that leaderboard is? It puts LLaMA2 Chat 70B at 92.66% and GPT-4 at 95.28%, only a ~3% difference.
by capableweb 3y ago
How accurate do you think that leaderboard is?
It puts LLaMA2 Chat 70B at 92.66% and GPT-4 at 95.28%, only a ~3% difference.
- bhouston 3y agoI don't know exactly, and all benchmarks are subject to problems. But in my experience the ordering in this list is roughly right for the models I've tried.