Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
danielhanchen
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
31.
▲
by
danielhanchen
5mo ago
Oh my apologies I didn't respond - if only HN had a notifier haha Oh yes we added a custom folder button which can pull .gguf files for now from any folder - it supports LM Studio and Ollama ones - but afreed it's still a mess. On
32.
▲
by
danielhanchen
6mo ago
We made Unsloth Studio which should help :) 1. Auto best official parameters set for all models 2. Auto determines the largest quant that can fit on your PC / Mac etc 3. Auto determines max context length 4. Auto heals tool calls, prov
33.
▲
by
danielhanchen
6mo ago
Haha :)
34.
▲
by
danielhanchen
6mo ago
Haha :) We had some issues with Kimi-2.6 since it was int4 and we were investigating how to handle it :)
35.
▲
by
danielhanchen
6mo ago
We also made some dynamic MLX ones if they help - it might be faster for Macs, but llama-server definitely is improving at a fast pace. https://huggingface.co/unsloth/Qwen3.6-27B-UD-MLX-4bit
36.
▲
by
danielhanchen
6mo ago
Yes sadly CUDA 13.2 is broken - NVIDIA will push a fix in CUDA 13.3
37.
▲
by
danielhanchen
6mo ago
Yes we have started doing diffusion GGUFs but it's in it's infancy :) But yes we do generate images to test quants out!
38.
▲
by
danielhanchen
6mo ago
Love the JPEG analogy :)
39.
▲
by
danielhanchen
6mo ago
Yes so chat templates and the actual implementations
40.
▲
by
danielhanchen
6mo ago
Thank you! Agreed on chat template issue
41.
▲
by
danielhanchen
6mo ago
Oh yes! This only applies if one uses hf download / snapshot_download - other normal download methods sadly won't have XET
42.
▲
by
danielhanchen
6mo ago
Probably yes
43.
▲
by
danielhanchen
6mo ago
Yep agreed at least 1 week is a good idea :) We do get early access to nearly all models, and we do find the most pressing issues sometimes. But sadly some issues are really hard to find and diagnose :(
44.
▲
by
danielhanchen
6mo ago
Oh that is pretty good! And the SVG one!
45.
▲
by
danielhanchen
6mo ago
They sometimes do! Qwen, Google etc do them!
46.
▲
by
danielhanchen
6mo ago
Sadly it's not always chat template fixes :( But yes we now split the first shard as pure metadata (10MB) for huge models - these include the chat template etc - so you only need to download that. For serious fixes, sadly we have to re
47.
▲
by
danielhanchen
6mo ago
Oh for multi files? Hmm ok let me check that out
48.
▲
by
danielhanchen
6mo ago
Hey thanks - yes agreed - for now we do: 1. Split metadata into shard 0 for huge models so 10B is for chat template fixes - however sometimes fixes cause a recalculation of the imatrix, which means all quants have to be re-made 2. Add HF di
49.
▲
by
danielhanchen
6mo ago
Oh hey - we're actually the 4th largest distributor of OSS AI models in GB downloads - see https://huggingface.co/unsloth https://unsloth.ai/docs/basics/unsloth-dynamic-2.0-ggufs is what migh
50.
▲
by
danielhanchen
6mo ago
Yep we can do that probs add a table - in general be post in discussions of model pages - for eg https://huggingface.co/unsloth/MiniMax-M2.7-GGUF/discussions... HF also provides SHA256 for eg https://hu
51.
▲
by
danielhanchen
6mo ago
Thanks!
52.
▲
by
danielhanchen
6mo ago
Oh thanks haha :) We try our best to get model releases out the door! :) Hope you're doing great!
53.
▲
by
danielhanchen
6mo ago
Yes this is fair - we try our best to communicate issues - I think we're mostly the only ones doing the communication that model A or B has been fixed etc. We try our best as model distributors to fix them on day 0 or 1, but 95% of iss
54.
▲
by
danielhanchen
6mo ago
We re-uploaded Gemma4 4 times - 3 times were due to 20 llama.cpp bug fixes, which we helped solve some as well. The 4th is an official Gemma chat template improvement from Google themselves, so these are out of our hands. All providers had
55.
▲
by
danielhanchen
6mo ago
No Bartowski's are more affected - (38% NaN) than ours (22%) - for MiniMax 2.7 see https://www.reddit.com/r/LocalLLaMA/comments/1slk4di/minimax... We already fixed ours. Bart hasn't yet but is
56.
▲
by
danielhanchen
6mo ago
1. Gemma-4 we re-uploaded 4 times - 3 times were 10-20 llama.cpp bug fixes - we had to notify people to upload the correct ones. The 4th is an official Gemma chat template improvement from Google themselves. 2. Qwen3.5 - we shared our 7TB r
57.
▲
by
danielhanchen
6mo ago
No it's not our fault - re our 4 uploads - the first 3 are due to llama.cpp fixing bugs - this was out of our control (we're llama.cpp contributors, but not the main devs) - we could have waited, but it's best to update when
58.
▲
by
danielhanchen
6mo ago
Yes we collab with them!
59.
▲
by
danielhanchen
6mo ago
Oh appreciate you trying out Unsloth Studio :)
60.
▲
Gemma 4 Fine-Tuning Guide
(unsloth.ai)
2 points
by
danielhanchen
6mo ago
|
0 comments
More ›