5 ms·
"Yes it improves performance!" proceeds to show the most unconvincing stats ever you can probably blow on your GPU and get a similar performance change
by make3 2y ago
"Yes it improves performance!" proceeds to show the most unconvincing stats ever
you can probably blow on your GPU and get a similar performance change
- refulgentis 2y agoI'm sorry, I don't understand what you mean. I checked the original article again too. As it stands, my understanding is you are claiming: - blowing on a GPU (which I take to mean doing roughly nothing) - gets roughly the same perf change - as moving from fp16 to q4
- danielhanchen 2y agoAre you referring to the finetuning part? The multiple bug fixes are separate from the finetuning sections - Unsloth itself makes finetuning 2x faster and use 70% less memory - the bug fixes are totally detached from finetuning - ie you can take the fixed version we uploaded at https://huggingface.co/unsloth/phi-4 https://huggingface.co/unsloth/phi-4, and use it in any framework or inference engine. Apologies I'm confused on the comment sorry. If you're questioning the credibility of the bug fixes - we fixed 8 bugs in Gemma https://x.com/danielhanchen/status/1765446273661075609 https://x.com/danielhanchen/status/1765446273661075609, multiple bugs in Llama, Mistral, Qwen, a gradient accumulation bug https://x.com/danielhanchen/status/1846235913443262891 https://x.com/danielhanchen/status/1846235913443262891 and much more
- grumpopotamus 2y ago2x faster than what?
- danielhanchen 2y agoOh 2x faster and uses >70% less memory than Hugging Face + Flash Attention 2! I did a CUDA / GPU Mode talk about it here: https://www.youtube.com/watch?v=hfb_AIhDYnA https://www.youtube.com/watch?v=hfb_AIhDYnA Also to the PyTorch team here: https://www.youtube.com/watch?v=MQwryfkydc0 https://www.youtube.com/watch?v=MQwryfkydc0 and the PyTorch Conference here: https://www.youtube.com/watch?v=PdtKkc5jB4g https://www.youtube.com/watch?v=PdtKkc5jB4g
- kouteiheika 2y ago> Oh 2x faster and uses >70% less memory than Hugging Face + Flash Attention 2! Is this doing the same type of fine-tuning, or are you comparing full bf16 fine-tuning in HF with 4-bit QLoRA in Unsloth (in which case it's not really an apples-to-apples comparison)? If it's the latter then do you have a comparison of the former?
- danielhanchen 2y agoOh I compared 4bit QLoRA HF+FA2 with Unsloth 4bit QLoRA. 16bit LoRA have similar boosts in performance! Full bf16 full finentuning is not yet supported, but it'll come out soon!
- danielhanchen 2y agoUpdate - the Phi-4 team is working on adding all our fixes to the original model! https://huggingface.co/microsoft/phi-4/discussions/21 https://huggingface.co/microsoft/phi-4/discussions/21
- make3 2y agohey this is great work, I'm sorry I complained, I'm thankful for what you're doing here
- danielhanchen 2y agoNo worries at all! :)
- danielhanchen 2y agoI uploaded our fixed versions to https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard#/?search=phi-4 https://huggingface.co/spaces/open-llm-leaderboard/open_llm_... which show the difference in scores. I agree it's not super convincing, so I provided anecdotal evidence as well - I'll work with the Phi-4 team to upstream these fixes! PS for further credibility, we also fixed 8 bugs in Gemma 1 - see https://x.com/danielhanchen/status/1765446273661075609 https://x.com/danielhanchen/status/1765446273661075609 , multiple bugs in Llama, Mistral, Qwen and other models