4 ms·
According to my understanding the blog post, FreeWilly2 performs near or above ChatGPT4 for most test-cases. Is this true? Am I misunderstanding this? Is this
by courseofaction 3y ago
According to my understanding the blog post, FreeWilly2 performs near or above ChatGPT4 for most test-cases. Is this true?
Am I misunderstanding this? Is this not a big deal?
- SparkyMcUnicorn 3y agoWhere are you seeing GPT-4? All I see is "compares favorably with GPT-3.5 for some tasks".
- dan-g 3y agoThe post mentions GPT4All benchmarks—maybe that’s where the confusion lies? https://gpt4all.io/ https://gpt4all.io/
- bbatsell 3y agoThe benchmarks towards the bottom are of ChatGPT-4, according to the footnotes.
- deleted 3y ago[deleted]
- rnosov 3y agoThe report cites both GPT-3.5 and GPT-4 scores on page 7 [1]. I've checked the numbers and they compare FreeWilly2 to GPT-3.5. For example, HellaSwag score of 85.5% corresponds to GPT-3.5. [1] https://arxiv.org/pdf/2303.08774v3.pdf https://arxiv.org/pdf/2303.08774v3.pdf
- courseofaction 3y agoThe footnoted papers mention GPT-4 in the title. Unless they're citing GPT3.5 results from papers with GPT-4 the name, which seems confusing.
- lgas 3y agoThere are no actual footnote marks that connect any statements in the post to the footnotes, so there are no specific claims referenced. But if you read the actual text of the page, they say it compares favorably to 3.5 for some tasks. Which means it falls short of 3.5 for the rest and GPT-4 for all of them, or else they surely would have mentioned that as well.
- deleted 3y ago[deleted]
- emadm 3y agoIt beats GPT 3.5 in some benchmarks, the first open model to do so I believe. Versions being worked on now will do much better. GPT 4 is far better and will likely not be beaten by any current open models and approaches but maybe an ensemble of them.