3 ms·
They'd still have strong models without distillation, but strong enough to challenge frontier models and to claim the meaningful market share that they have? Pr
by giwook 1mo ago
They'd still have strong models without distillation, but strong enough to challenge frontier models and to claim the meaningful market share that they have? Probably not.
- truncate 1mo agoOr, they can figure out something else out? I recall couple years ago when China didn't have enough GPUs (still don't?), DeepSeek team figured out how to train with less computing. IIRC they made Mixture of Experts mainstream and made really optimized kernels and clever use of PTX instruction set.
- ACCount37 1mo ago"China does it in a cave with a box of scraps" is a myth. Chinese labs play the shell game to get their hands on a lot of compute outside China. Tricks like distillation save compute in the RL leg of the process - where a lot of the frontier labs puts their own training run compute.
- truncate 1mo agoI'm sure they do, but its not 0 or 1 thing.