6 ms·
Would really want to see some benchmarks against ChatGPT / GPT-4. The improvements in the given benchmarks for the larger models (Llama v1 65B and Llama v2 70B
by cheeseface 3y ago
Would really want to see some benchmarks against ChatGPT / GPT-4.
The improvements in the given benchmarks for the larger models (Llama v1 65B and Llama v2 70B) are not huge, but hard to know if still make a difference for many common use cases.
- jmiskovic 3y agoThen why not read their paper? "The largest Llama 2-Chat model is competitive with ChatGPT. Llama 2-Chat 70B model has a win rate of 36% and a tie rate of 31.5% relative to ChatGPT."
- capableweb 3y agoDo they specify which GPT version they used? Could Llama 2 really beat GPT-4?
- jmiskovic 3y agoThe 70B Llama2 model ties in with 173B ChatGPT-0301 model. The GPT-4 still stands unchallenged.
- sebzim4500 3y agoSource on the 173B parameters?
- jmiskovic 3y agoIt's actually 175B. https://arxiv.org/pdf/2005.14165.pdf https://arxiv.org/pdf/2005.14165.pdf
- exradr 3y agoThe wikipedia article for GPT-4 has this as its source: https://the-decoder.com/gpt-4-architecture-datasets-costs-and-more-leaked/ https://the-decoder.com/gpt-4-architecture-datasets-costs-an...
- davidkunz 3y agoThey used ChatGPT-0301, it can't beat GPT-4.
- chaxor 3y agoIt would be nice to see 6 of them trained for different purposes by combining 5 of their outputs together and 1 trained to summarize for the most complete and correct output. If we are to trust the leaks about GPT-4, this may be a more fair comparison, even if it is only ~10-20% of the size or so.
- pedrovhb 3y agoIsn't that essentially beam sampling?
- majorbadass 3y ago"In addition to open-source models, we also compare Llama 2 70B results to closed-source models. As shown in Table 4, Llama 2 70B is close to GPT-3.5 (OpenAI, 2023) on MMLU and GSM8K, but there is a significant gap on coding benchmarks. Llama 2 70B results are on par or better than PaLM (540B) (Chowdhery et al., 2022) on almost all benchmarks. There is still a large gap in performance between Llama 2 70B and GPT-4 and PaLM-2-L."
- gentleman11 3y agoit's not open source
- illnewsthat 3y agoThe paper[1] says this in the conclusion: > [Llama 2] models have demonstrated their competitiveness with existing open-source chat models, as well as competency that is equivalent to some proprietary models on evaluation sets we examined, although they still lag behind other models like GPT-4. It also seems like they used GPT-4 to measure the quality of responses which says something as well. [1] https://ai.meta.com/research/publications/llama-2-open-foundation-and-fine-tuned-chat-models/ https://ai.meta.com/research/publications/llama-2-open-found...
- janejeon 3y agoIn the paper, I was able to find this: > In addition to open-source models, we also compare Llama 2 70B results to closed-source models. As shown in Table 4, Llama 2 70B is close to GPT-3.5 (OpenAI, 2023) on MMLU and GSM8K, but there is a significant gap on coding benchmarks. Llama 2 70B results are on par or better than PaLM (540B) (Chowdhery et al., 2022) on almost all benchmarks. There is still a large gap in performance between Llama 2 70B and GPT-4 and PaLM-2-L.