3 ms·
I have lots of experience with both, and I use both together for different use cases. SpaCy fills the need of predictable/explainable pattern matching and NER
by binarymax 6y ago
I have lots of experience with both, and I use both together for different use cases. SpaCy fills the need of predictable/explainable pattern matching and NER - and is very fast and reasonably accurate on a CPU. Huggingface fills the need for task based prediction when you have a GPU.
- danieldk 6y agoHuggingface fills the need for task based prediction when you have a GPU. With model distillation, you can make models that annotate hundreds of sentences per second on a single CPU with a library like Huggingface Transformers. For instance, one of my distilled Dutch multi-task syntax models (UD POS, language-specific POS, lemmatization, morphology, dependency parsing) annotates 316 sentences per second with 4 threads on a Ryzen 3700X. This distilled model has virtually no loss in accuracy compared to the finetuned XLM-RoBERTa base model. I don't use Huggingface Transformers, but ported some of their implementations to Rust [1], but that should not make a big difference since all the heavy lifting happens in C++ in libtorch anyway. tl;dr: it is not true that tranformers are only useful for GPU prediction. You can get high CPU prediction speeds with some tricks (distillation, length-based bucketing in batches, using MKL, etc.). [1] https://github.com/tensordot/syntaxdot/tree/main/syntaxdot-transformers/src/models https://github.com/tensordot/syntaxdot/tree/main/syntaxdot-t...
- binarymax 6y agoInteresting. Did you start from a Distilled base model (like DistilRoBerta), or did you distill your fine-tuned model?
- danieldk 6y agoSorry for the late reply. I distilled from my own finetuned XLM-RoBERTa model.
- aabhay 6y agoIs there a standard template for creating a distilled model? I didn’t see a public hugging face implementation, just the models