4 ms·
For others that were confused by the Gemini versions: the main one being discussed is Gemini Ultra (which is claimed to beat GPT-4). The one available through B
by m3at 3y ago
For others that were confused by the Gemini versions: the main one being discussed is Gemini Ultra (which is claimed to beat GPT-4). The one available through Bard is Gemini Pro.
For the differences, looking at the technical report [1] on selected benchmarks, rounded score in %:
Dataset | Gemini Ultra | Gemini Pro | GPT-4
MMLU | 90 | 79 | 87
BIG-Bench-Hard | 84 | 75 | 83
HellaSwag | 88 | 85 | 95
Natural2Code | 75 | 70 | 74
WMT23 | 74 | 72 | 74
[1] https://storage.googleapis.com/deepmind-media/gemini/gemini_1_report.pdf https://storage.googleapis.com/deepmind-media/gemini/gemini_...
- nathanfig 3y agoThanks, I was looking for clarification on this. Using Bard now does not feel GPT-4 level yet, and this would explain why.
- dkarras 3y agonot even original chatgpt level, it is a hallucinating mess still. Did the free bard get an update today? I am in the included countries, but it feels the same as it has always been.
- Traubenfuchs 3y agoformatted nicely: Dataset | Gemini Ultra | Gemini Pro | GPT-4 MMLU | 90 | 79 | 87 BIG-Bench-Hard | 84 | 75 | 83 HellaSwag | 88 | 85 | 95 Natural2Code | 75 | 70 | 74 WMT23 | 74 | 72 | 74
- carbocation 3y agoI realize that this is essentially a ridiculous question, but has anyone offered a qualitative evaluation of these benchmarks? Like, I feel that GPT-4 (pre-turbo) was an extremely powerful model for almost anything I wanted help with. Whereas I feel like Bard is not great. So does this mean that my experience aligns with "HellaSwag"?
- tarruda 3y agoI get what you mean, but what would such "qualitative evaluation" look like?
- carbocation 3y agoI think my ideal might be as simple as a few people who spend a lot of time with various models describing their experiences in separate blog posts.
- tarruda 3y agoI see. I can't give any anecdotal evidence on ChatGPT/Gemini/Bard, but I've been running small LLMs locally over the past few months and have amazing experience with these two models: - https://huggingface.co/mlabonne/NeuralHermes-2.5-Mistral-7B https://huggingface.co/mlabonne/NeuralHermes-2.5-Mistral-7B (general usage) - https://huggingface.co/deepseek-ai/deepseek-coder-6.7b-instruct https://huggingface.co/deepseek-ai/deepseek-coder-6.7b-instr... (coding) OpenChat 3.5 is also very good for general usage, but IMO NeuralHermes surpassed it significantly, so I switched a few days ago.
- carbocation 3y agoThanks! I’ve had a good experience with the deepseek-coder:33b so maybe they’re on to something.
- fasttransients 3y agoThank you for the suggestions – really helpful for my hobby project. Can't run anything bigger than 7B on my local setup, which is a fun constraint to play with.
- p_j_w 3y ago>Like, I feel that GPT-4 (pre-turbo) was an extremely powerful model for almost anything I wanted help with. Whereas I feel like Bard is not great. So does this mean that my experience aligns with "HellaSwag"? It doesn't mean that at all because Gemini Turbo isn't available in Bard yet.
- teleforce 3y agoExcellent comparison, it seems that GPT-4 is only winning in one dataset benchmark namely HellaSwag for sentence completion. Can't wait to get my hands on Bard Advanced with Gemini Ultra, I for one welcome this new AI overlord.
- aroo 3y agoHorrible comparison given one score was achieved using 32-shot CoT (Gemini) and the other was 5-shot (GPT-4).
- throwaway287391 3y agoCoT@32 isn't "32-shot CoT"; it's CoT with 32 samples (or rollouts) from the model, and the answer is taken by consensus vote from those rollouts. It doesn't use any extra data, only extra compute. It's explained in the tech report here: > We find Gemini Ultra achieves highest accuracy when used in combination with a chain-of-thought prompting approach (Wei et al., 2022) that accounts for model uncertainty. The model produces a chain of thought with k samples, for example 8 or 32. If there is a consensus above a preset threshold (selected based on the validation split), it selects this answer, otherwise it reverts to a greedy sample based on maximum likelihood choice without chain of thought. (They could certainly have been clearer about it -- I don't see anywhere they explicitly explain the CoT@k notation, but I'm pretty sure this is what they're referring to given that they report CoT@8 and CoT@32 in various places, and use 8 and 32 as the example numbers in the quoted paragraph. I'm not entirely clear on whether CoT@32 uses the 5-shot examples or not, though; it might be 0-shot?) The 87% for GPT-4 is also with CoT@32, so it's more or less "fair" to compare that Gemini's 90% with CoT@32. (Although, getting to choose the metric you report for both models is probably a little "unfair".) It's also fair to point out that with the more "standard" 5-shot eval Gemini does do significantly worse than GPT-4 at 83.7% (Gemini) vs 86.4% (GPT-4).
- dragonwriter 3y ago> I'm not entirely clear on whether CoT@32 uses the 5-shot examples or not, though; it might be 0-shot? Chain of Thought prompting, as defined in the paper referenced, is a modification of few-shot prompting where the example q/a pairs used have chain-of-thought style reasoning included as well as the question and answer, so I don't think that, if they were using a 0-shot method (even if designed to elicit CoT-style output) they would call it Chain of Thought and reference that paper.
- make3 3y agothe numbers are not at all comparable, because Gemini uses 34 shot and variable shot vs 5 for gpt 4. this is very deceptive of them.
- bitshiftfaced 3y agoYes and no. In the paper, they do compare apples to apples with GPT4 (they directly test GPT4's CoT@32 but state its 5-shot as "reported"). GPT4 wins 5-shot and Gemini wins CoT@32. It also came off to me like they were implying something is off about GPT4's MMLU.
- tiziano88 3y agoPermanent link to the result table contents: https://static.space/sha2-256:ea7e5d247afa8306cb84cbbd4438fd6e58a3109781758099cde123d4f6b44517 https://static.space/sha2-256:ea7e5d247afa8306cb84cbbd4438fd...
- deleted 3y ago[deleted]