7 ms·
SpaCy: Industrial-Strength Natural Language Processing (NLP) in Python
- bobosha 1y agoSpaCy is awesome - we have used it in a number of enterprise-grade applications and found it to hold up well.
- erikqu 1y agoI figured this project died post-chatgpt, I <3 spacy, learned a ton on this platform back in the day
- giantg2 1y agoWhat are the key differences from other NLP Python libraries?
- jihadjihad 1y agoSpeed (the C in spaCy). A decade ago it was hard to find anything actually production grade for NLP, most packages had an academic bent or were useful for prototyping. SpaCy really changed the game by being able to run performant NLP on standard hardware.
- esafak 1y agonltk was slow.
- EagnaIonat 1y agonltk was never intended for production, it was for built for teaching.
- bratao 1y agoI'm really curious about the history of spaCy. From my PoV: it grew a lot during the pandemic era, hiring a lot of employees. I remember something about raising money for the first time. It was very competitive in NLP tasks. Now it seems that it has scaled back considerably, with a dramatic reduction in employees and a total slowdown of the project. The v4 version looks postponed. It isn't competitive in many tasks anymore (for tasks such as NER, I get better results by fine-tuning a BERT model), and the transformer integration is confusing.
- binarymax 1y agoI’ve had success with fine tuning their transformer model. The issue was that there was only one of them per language, compared to huggingface where you have a choice of many of quality variants that best align with your domain and data. The SpaCy API is just so nice. I love the ease of iterating over sentences, spans, and tokens and having the enrichment right there. Pipelines are super easy, and patterns are fantastic. It’s just a different use case than BERT.
- cantdutchthis 1y agoformer employee here, Matt wrote a blogpost with pretty much all of the details here: https://honnibal.dev/blog/back-to-our-roots https://honnibal.dev/blog/back-to-our-roots
- microtonal 1y ago:wave: Also: https://explosion.ai/blog/back-to-our-roots-company-update https://explosion.ai/blog/back-to-our-roots-company-update (Interesting tidbit: I got hired by Explosion after a HN comment on model distillation :))
- binarymax 1y agoI’ve been a user of SpaCy since 2016. I haven’t touched it in years and I just picked it up again to develop a new metric for RAG using part of speech coverage. The API is one of the best ever, and really set the bar high for language tooling. I’m glad it’s still around and getting updates. I had a bit of trouble integrating it with uv, but nothing too bad. Thanks to the explosion team for making such an amazing project and keeping it going all these years. To the new “AI” people in the room: checkout SpaCy, and see how well it works and how fast it chews through text. You might find yourself in a situation where you don’t need to send your data to OpenAI for some small things. Edit: I almost forgot to add this little nugget of history: one of Huggingfaces first projects was a SpaCy extension for conference resolution. Built before their breakthrough with transformers https://github.com/huggingface/neuralcoref https://github.com/huggingface/neuralcoref
- deleted 1y ago[deleted]
- jehejej 1y ago*coreference resolution.
- ok_dad 1y agoWhat’s great about the API that you enjoy and do you have anything you hate about it? I’m writing a small library at work for some NLP tasks and I haven’t got a whole lot of experience in writing libraries for NLP, so I’m interested in what would make my library the best for the user.
- binarymax 1y agoThe thing about spaCys API is that it perfectly aligns with how NLP worked at the time with actual programming paradigms and allows you to be very pythonic. For example, you can use list comprehension to get all the nouns from a document in a one liner. These days NLP is quite different, because we look for outcomes rather than iterating over tokens. What does your NLP library need to do? The way I design APIs is I write the calling code that I want to exist, and then I write the API to make it work. Here’s an example I’ve worked on for LLM integration. I just wanted to be able to get simple answers from an LLM and cast the answer to a type: https://www.npmjs.com/package/llm-primitives https://www.npmjs.com/package/llm-primitives
- skeptrune 1y agoSpaCy is criminally underrated. I expect to see it experience a new wave of growth as folks new to AI start to realize all of the language tooling they need to build more reliable "traditional" ML pipelines. API surface is designed well and it's still actively maintained almost 10 years after it initially went public.
- chpatrick 1y agoIs there any use case for "traditional" NLP in the age of LLMs?
- lyu07282 1y agoI used to work a lot with those pipelines, I think the truth is that LLMs (and LLM embeddings) have surpassed pretty much all traditional NLP. I guess if speed is more important than accuracy? but even then, like with small embedded LLMs they still outperform "traditional NLP" on pretty much every task probably. So it doesn't make a lot of sense to not use it nowadays.
- skeptrune 1y agoMost definitely! LLMs are amazing tools for generating synthetic datasets that can be used alongside traditional NLP to train things like decision trees with libraries like cat/xgboost. I have a search background so learning to rank is always top of mind for me, but there other places like sentiment analysis, intent detection, and topic classification where it's great too.
- chpatrick 1y agoBut for the analysis use cases you mentioned, can't you just ask an LLM to read the text and output the answer as JSON, and you're done? Is it just because running LLMs is expensive?
- skeptrune 1y agoNo, it's just slow and less accurate. Wrong tool for the job when you care a lot about understanding the reasoning and internals of what the model is caring the most about.
- roadside_picnic 1y agoA friend, who also has a background in NLP, was asking me the other day "Is there still even a need for traditional NLP in the age of LLMs?" This is one of the under-discussed areas of LLMs imho. For anything that would have have required either word2vec embeddings of a tf-idf representation (classification tasks, sentiment analysis, etc) there are rare exceptions where it wouldn't just be better to start with a semantic embedding from an LLM. For NER and similar data extraction tasks, the only advantage of traditional approaches is going to be speed, but my experience in practice is that accuracy is often much more important than speed. Again, I'm not sure why not start with an LLM in these cases. There are still a few remaining use cases (PoS tagging comes to mind), but honestly, if I have a traditional NLP task today, I'm pretty sure I'm going to start with an LLM as my baseline.
- coder68 1y agoI have been working on text classification tasks at work, and I have found that for my particular use-case, LLMs are not performing well at all. I have spent a few thousand dollars trying, and I have tried everything from few-shot to asking simple binary yes/no questions, and I have had mixed success. I have stopped trying to use LLMs for this project and switched to discriminative models (Logistic Regression with TFIDF or Embeddings), which are both more computationally efficient and more debuggable. I'm not entirely sure why, but for anything with many possible answers, or to which there is some subjectivity, I have not had success with LLMs simply due to inconsistency of responses. For VERY obvious tasks like: "is this store a restaurant or not?" I have definitely had success, so YMMV.
- leobg 1y agoIf I have 1,000 labeled examples for a classification task, I’ll expand that into a training dataset using augmentation, and then finetune a small model like RoBERTa. It’s fast, cheap, accurate — and predictable. Others have had success with SetFit as the training framework and Ettin as the base model.
- coder68 1y ago
- patrickhogan1 1y agoSpaCy was my go to library for NER before GPT 3+. It was 10x better than regex (though you could also include regex within your pipelines. Its annotation tooling was so far ahead. It is still crazy to me that so much of the value in the data annotation space went to Scale AI vs tools like SpaCy that enabled annotation at scale in the enterprise.
- renegat0x0 1y agoI use spacy in my raspberry pi project. I am not sure I want to use LLM for analyzing words in it.
- joshdavham 1y agoI’ve been using SpaCy for many of my projects for 5 years now. The library has incredible ergonomics and allows you to reuse the same API across languages as different as French and Japanese! I also appreciate that they allow you to install different model sizes (I usually go with small).
- robotswantdata 1y agoSpaCy is the OG, nothing but praise for the devs. Built a lot of very powerful legal apps with it pre GPT , very useful today for NER where you want something “small”, fast and reliable. Used it again recently and the dev experience is 1000x that of wrangling LLMs.
- ur-whale 1y agoAt the risk of asking a naive question ... why would anyone still do traditional NLP today?
- nutjob2 1y agoThere is a need for good and easy NLP "structured" interfaces to traditional structured data software. This is a gaping hole in the tech right now. Most other NLP tasks can be handled by ML approaches but this one is a poor fit for those. I'm sure LLM true believers will disagree.
- apprentice7 1y agoCould you give a couple specific examples? I'm trying to get into traditional NLP but everything I find is AI related and I don't know if it's worth going the traditional route long-term.
- jftuga 1y agoI recently wrote an open source Python module to deidentify people's names and gender specific pronouns. It uses spaCy's Named Entity Recognition (NER) capabilities combined with custom pronoun handling. See the screenshot in the README.md file. * https://github.com/jftuga/deidentification https://github.com/jftuga/deidentification * https://pypi.org/project/text-deidentification/ https://pypi.org/project/text-deidentification/