2 ms·
Classic knowledge distillation! I’d even argue that we won’t need 8x7b for fine-tuning here. Soon enough, phi-2 or phixtral models will be sufficiently powerfu
by fzysingularity 3y ago
Classic knowledge distillation! I’d even argue that we won’t need 8x7b for fine-tuning here. Soon enough, phi-2 or phixtral models will be sufficiently powerful after fine tuning for these domains.