4 ms·
Hi. It's nice to see these fixes. I got a question after checking results on the open LLM leaderboard[1]. Comparing the result of NyxKrage/Microsoft_Phi-4 and
by RandyOrion 2y ago
Hi. It's nice to see these fixes.
I got a question after checking results on the open LLM leaderboard[1].
Comparing the result of NyxKrage/Microsoft_Phi-4 and microsoft/phi-4 or unsloth/phi-4, I can see fixing both the tokenizer and chat template causes the performance of both IFEval and BBH to increase. However, the performance on MATH, GPQA and MUSR degrades A LOT.
Is there any explanation on why this is happening?
[1] https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard#/?search=phi-4 https://huggingface.co/spaces/open-llm-leaderboard/open_llm_...
- danielhanchen 2y agoYep that is something I've been dumbfounded by as well - the official Microsoft phi4 upload also suffers on MATH so at least we can rule out it's because I did something wrong. I thought of two possibilities: 1. 509 does better on MATH but absolutely terribly on IFEVAL because it does not use a chat template - whilst others so use the chat template. 2. I think HF uses exact matching I think so maybe that's the culprit. I can test 1. by resubmitting without using the chat template!
- RandyOrion 2y agoThanks for the reply. Would like to see whether the chat template is cause of this strange behavior.