3 ms·
To me it seems that a plug-and-play integration with HuggingFace transformers is the killer feature here. If this simply allows to fine tune a model on your loc
by startupsfail 3y ago
To me it seems that a plug-and-play integration with HuggingFace transformers is the killer feature here. If this simply allows to fine tune a model on your local GPU, that you otherwise can’t train, this is pretty interesting.
It seems that it also might work for 4bit-quantized models. And removing the step of model quantization and acceleration in terms of TOPS for quantized inference can maybe compensate for the need to do more steps.
Also, a good question of how well it works in a distributed/batched setting.