19 ms·
Here are some benchmarks, excellent to see that an open model is approaching (and in some areas surpassing) GPT-3.5! AI2 Reasoning Challenge (25-shot) - a set
by lappa 3y ago
Here are some benchmarks, excellent to see that an open model is approaching (and in some areas surpassing) GPT-3.5!
AI2 Reasoning Challenge (25-shot) - a set of grade-school science questions.
- Llama 1 (llama-65b): 57.6
- LLama 2 (llama-2-70b-chat-hf): 64.6
- GPT-3.5: 85.2
- GPT-4: 96.3
HellaSwag (10-shot) - a test of commonsense inference, which is easy for humans (~95%) but challenging for SOTA models.
- Llama 1: 84.3
- LLama 2: 85.9
- GPT-3.5: 85.3
- GPT-4: 95.3
MMLU (5-shot) - a test to measure a text model’s multitask accuracy. The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more.
- Llama 1: 63.4
- LLama 2: 63.9
- GPT-3.5: 70.0
- GPT-4: 86.4
TruthfulQA (0-shot) - a test to measure a model’s propensity to reproduce falsehoods commonly found online. Note: TruthfulQA in the Harness is actually a minima a 6-shots task, as it is prepended by 6 examples systematically, even when launched using 0 for the number of few-shot examples.
- Llama 1: 43.0
- LLama 2: 52.8
- GPT-3.5: 47.0
- GPT-4: 59.0
[0] https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
[1] https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard/discussions/30#6474fd3b82907acdddf34e33 https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
- redox99 3y agoYour Llama2 MMLU figure is wrong
- sebzim4500 3y agoLooks like he copied it from https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb... I see different figures in different places, no idea what's right.
- doctoboggan 3y agoGood to see these results, thanks for posting. I wonder if GPT-4's dominance is due to some secret sauce or if its just the first mover advantage and Llama will be there soon.
- famouswaffles 3y agoIt's just scale. But scale that comes with more than an order of magnitude more expense than the Llama models. I don't see anyone training such a model and releasing it for free anytime soon
- bbor 3y agoI thought it was revealed to be fundamentally ensemblamatic in a way the others weren’t? Using “experts” I think? Seems like it would meet the bar for “secret sauce” to me
- famouswaffles 3y agoSparse MoE models are neither new nor secret. The only reason you haven't seen much use of them for LLMs is because they would typically well underperform their dense counterparts. Until this paper (https://arxiv.org/abs/2305.14705 https://arxiv.org/abs/2305.14705) indicated they apparently benefit far more from Instruct tuning than dense models, it was mostly a "good on paper" kind of thing. In the paper, you can see the underperformance i'm talking about. Flan-Moe-32b(259b total) scores 25.5% on MMLU pre Instruct tuning and 65.4 after. Flan 62b scores 55% before Instruct tuning and 59% after.
- cubefox 3y agoThis paper came out well after GPT-4, so apparently this was indeed a secret before then.
- famouswaffles 3y agoThe user I was replying to was talking about the now and future. We also have no indication sparse models outperform dense counterparts so it's scale either way.
- HeWhoLurksLate 3y agoIs there a difference here between a secret and an unknown? It may well be that some researcher / comp engineer had an idea, tried it out, realized it was incredibly powerful, implemented it for real this time and then published findings after they were sure of it? I'm more of a mechanical engineering adjacent professional than a programmer and only follow AI developments loosely
- gitgud 3y agoIs it possible that some LLM’s are trained on these benchmarks? Which would mean they’re overfitting and are incorrectly ranked? Or am I misunderstanding these benchmarks?…
- bbor 3y agoIt would be a bit of a scandal, and IMO too much hassle to sneak in. These models are trained on massive amounts of text - specifically anticipating which metrics people will care about and generating synthetic data just for them seems extra. But not an expert or OP!
- stu2b50 3y agoI don't think it's a scandal, it's a natural thing that happens when iterating on models. OP doesn't mean they literally train on those tests, but that as a meta-consequence of using those tests as benchmarks, you will adjust the model and hyperparameters in ways that perform better on those tests. For a particular model you try to minimally do this by separating a test and validation set, but on a meta-meta level, it's easy to see it happening.
- jasonfarnon 3y agoYou don't see an engineer at an extremely PR-conscious company at least checking how their model performs on popular benchmarks before rolling it out? And if its performance is lackluster, you do you really see them doing nothing about it? It probably doesn't make a huge difference anyway. I know those old vision models were overfitted to the standard image library benchmarks, but they were still very impressive.
- marcopicentini 3y agoHow they compare the exact value returned in a response? I found that returning a stable json format is something unpredictable or it reply in a different language.
- ineedasername 3y agoWhen were the GPT-4 benchmarks calculated, on original release or more recently? (curious per the debate about alleged gpt-4 nerfing)
- lappa 3y agoThey're based on the original technical report. "Refuel" has run a different set of benchmarks on GPT-3.5 and GPT-4 and found a decline in quality. https://www.refuel.ai/blog-posts/gpt-3-5-turbo-model-comparison https://www.refuel.ai/blog-posts/gpt-3-5-turbo-model-compari...
- ShamelessC 3y agoPlenty of the complaints/accusations predate the release of the 0613 set of models. To be clear, I have trouble with the theory as I have not yet seen evidence of "nerfing". What you provided is actually the _only_ evidence I've seen that suggests degradation - but in this case OpenAI is being completely transparent about it and allows you to switch to the 0314 model if you would like to. Every complaint I have seen has been highly anecdotal, lacking any rigor, and I bet are explained by prolonged usage resulting in noticing more errors. Also probably a bit of "the magic is gone now" psychological effect (like how a "cutting edge" video game such as Half-Life 2 feels a bit lackluster these days).
- digitcatphd 3y agoCould it be the case that many of these benchmarks are just learning this material included in their parameters?
- Roark66 3y agoI have to say in my experience falcon-40b-instruct got very close to chatgpt (gpt-3. 5),even surpassing it in few domains. However, it is important to note (not at all)OpenAI are doing tricks with the model output. So comparing OS models with just greedy output decoding (very simple) is not fair for OS models. Still, I'm very excited this model at 13B seems to be matching falcon-40B in some benchmarks. I'm looking forward to using it :-)
- fnl 3y ago> OpenAI are doing tricks with the model output Do you have any pointers to the “tricks” that are being applied?
- babushkanazi 3y ago[dead]
- jcuenod 3y agoSounds like a reference to Mixture of Experts
- zzzzzzzza 3y agocould be something like prompt rewriting or chain of thought or reflexion going on in the background as well