5 ms·
PaLM 2 on HumanEval coding benchmark (0 shot): 37.6% success GPT-4: 67% success Not even close, gpt4 miles ahead
by technics256 3y ago
PaLM 2 on HumanEval coding benchmark (0 shot):
37.6% success
GPT-4:
67% success
Not even close, gpt4 miles ahead
- orpheansodality 3y agoGPT-4 is a fine-tuned model (likely first fine-tuned for code, then for chat on top of that like gpt-3.5-turbo was[0]), while PaLM2 as reported is a foundational model without any additional fine-tuning applied yet. I would expect its performance to improve on this if it were fine-tuned, though I don't have a great sense of what the cap would be. [0] https://platform.openai.com/docs/model-index-for-researchers https://platform.openai.com/docs/model-index-for-researchers
- macrolime 3y agoThey also write about Flan-PaLM2 which is instruction fine-tuned, but still some ways off GPT-4.
- heliophobicdude 3y agoHumaneval needs careful consideration though. In the GPT-4 technical report, they reported contamination of humaneval data in the training data. They did measure against a "non-contaminated" training set but no idea if that can still be trusted. https://cdn.openai.com/papers/gpt-4.pdf https://cdn.openai.com/papers/gpt-4.pdf
- hipmanbro 3y agoThere were a few reasoning benchmarks that I noticed think they omitted a direct comparison since they weren't as competitive compared to GPT-4, and instead opted to just show the benchmarks comparing itself to other versions of PaLM or other language models HellaSwag: GPT-4: 95.3%, PaLM 2-L: 86.8% MMLU: GPT-4: 86.4%, Flan-PaLM 2-L: 81.2% ARC: GPT-4: 96.3%, PaLM 2-L: 89.7% (from: GPT-4 paper: https://arxiv.org/pdf/2303.08774.pdf https://arxiv.org/pdf/2303.08774.pdf)