3 ms·
For me, the most important thing about this version is the reduced memory usage. Previously the smallest english model took 1GB of RAM, making it troublesome to
by kamac 9y ago
For me, the most important thing about this version is the reduced memory usage. Previously the smallest english model took 1GB of RAM, making it troublesome to run it on any cloud instances. If v2 is to take ~200mb instead, that's a huge improvement.
- syllogism 9y agoThe thing that always bothered me about v1 was that it was fast, but in many ways not that scaleable. I really under-estimated the importance of Pickle support for instance, because I didn't appreciate that that's how multiprocessing works in Python. You might find this method particularly useful for meeting memory constraints: https://spacy.io/api/vocab#prune_vectors https://spacy.io/api/vocab#prune_vectors . This lets you reduce a large word vectors table to a small one by remembering the nearest neighbours for the words you prune out. So if you have a rare word like 'biophysicist', you can map it to the vector for a word like 'scientist', and get a close-enough word vector for it.
- Vaskivo 9y agoDoes that mean that it can run on a Raspberry Pi?
- kamac 9y agoUnless RAM usage hasn't significantly increased beyond 200mb since alpha, it should run.