6 ms·
Advanced NLP with spaCy v3
- artembugara 5y agoWe've been using spaCy a lot for the past few months. Mostly for non-production use cases, however, I can say that it is the most robust framework for NLP at the moment. V3 added support for transformers: that's a killer feature as many models from https://huggingface.co/docs/transformers/index https://huggingface.co/docs/transformers/index work great out of the box. At the same time, I found NER models provided by spaCy to have a low accuracy while working with real data: we deal with news articles https://demo.newscatcherapi.com/ https://demo.newscatcherapi.com/ Also, while I see how much attention ML models get from the crowd, I think that many problems can be solved with rule-based approach: and spaCy is just amazing for these. Btw, we recently wrote a blog post comparing spaCy to NLTK for text normalization task: https://newscatcherapi.com/blog/spacy-vs-nltk-text-normalization-comparison-with-code-examples https://newscatcherapi.com/blog/spacy-vs-nltk-text-normaliza...
- artembugara 5y agoAlso I have an article about spaCy NER: https://newscatcherapi.com/blog/named-entity-recognition-with-spacy https://newscatcherapi.com/blog/named-entity-recognition-wit... The conclusion I came up with: "A few notes on my Spacy NER accuracy with "real world" data Low accuracy with sentences without a proper casing 1. Low accuracy overall, even with a large model 2. You'd need to fine-tune your model if you want to use it in production 3. Overall, there's no open-source high accuracy NER model that you can use out-of-a-box"
- wyldfire 5y agoI assume your product does some kind of entity disambiguation and/or link to an ontology? Spacy doesn't provide this out of the box either, AFAICT. Can you share more info about how you do it?
- artembugara 5y agoWe don't provide entity disambiguation out of a box. It's more of a on request for Enterprise clients. But overall, entity disambiguation is one of the most useful and difficult tasks in the NLP. SpaCy supports entity linking via knowledge base: https://spacy.io/api/entitylinker https://spacy.io/api/entitylinker
- Vetch 5y ago> Overall, there's no open-source high accuracy NER model that you can use out-of-a-box" Part of it is most underestimate the complexity of NER and the rest of it, in my opinion, is that NER is not well-defined as a classification problem. At least in my experience, having a specific battery of questions to query documents, first by transformer based semantic search and narrowed by Q/A models, removed the need for explicit NER, entity linking or relation extraction. For the case of entities as features for rule systems, shallow models and using all label predictions instead of just selecting argmax has been sufficiently robust. Using big transformers for classification doesn't pay enough to be worth it there.
- pantsforbirds 5y agoWe use spaCy at work for (mostly) news articles as well. We've been pretty impressed with it overall for detecting larger trends using the NER models. I've been contemplating whether it might be useful to make a spaCy module that uses a Count-Min Sketch to track the top N of each of the NER categories partitioned on a daily (or weekly etc.) time. Think it could be an interesting use case to get sort of similar results to Google's search trends.
- artembugara 5y agoI'd really love to chat about that. Any chance to connect? email in bio
- Xenoamorphous 5y agoI don’t know how it compares with other paid alternatives (like Google’s or Amazon’s) but spaCy’s NER was pretty close to the (paid) service we were using (IBM) to the point we ditched IBM. Also for news articles. But yeah disambiguation/entity linking would be nice.
- artembugara 5y agoI'd be happy to chat more if you want.
- brd 5y agoI really appreciate how accessible SpaCy has made NLP work but their NER is definitely low accuracy. Where stem/lem felt critical to successful NLP processing a few years ago, we've found stem/lem work to be much less important for downstream tasks when transformer based models are involved. For topic extraction stem/lem still seems to do a lot to improve accuracy and for rules based approaches I can still see how it would facilitate more efficient processing at scale. I'd be curious to hear your experience fine tuning and/or training new models after stem/lem processing with transformers, we've admittedly done little testing to see how transformers actually performer if properly tuned to post-processed data.
- artembugara 5y agoDid you try something like autoNLP by huggingface?
- brd 5y agoNo, we've got our own fine tuning pipeline and initial tests showed better performance without traditional stem/lem processing so we dropped it from our classification pipelines and haven't seen a need to revisit.
- kulikalov 5y agoAre you using the high accuracy eng model for NER? I’ve been very happy with orgs recognition, it actually did way better than any other open source model in my case.
- artembugara 5y agoTry it on a sentence where all tokens are lower/upper case. It just doesn’t really work.
- PeterisP 5y agoWell, caseless text is a special scenario and not the default scenario. Case is a very strong signal for NER disambiguation, so if you want to support that, then you should apply a special model for that - because if the default model would include support for caseless text, then it would harm the accuracy for all the majority of scenarios where text actually is cased properly. In essence, the current approaches are targeted for one domain of text over another. You can have a model that works reasonably in one scenario, or a model that works reasonably in another scenario, or an universal model that works poorly in all scenarios and thus is useless unless you really don't know what you're going to be analyzing. You can support non-literary slang, but that comes at a cost for accuracy on literary languages. You can support multiple variants of language (e.g. for English - British, Indian, AAVE and non-AAVE American) but that comes at a cost of accuracy on any particular variant. You can support text ridden with typos, grammatical mistakes and chat-abbreviations, but that comes at a cost on correct text. The same applies for word casing. So for all of these things you try to support them if and only if you think you need them, since you don't have much of an "accuracy reserve" to sacrifice; the systems generally are barely sufficient for their use for your target domain, and they become not sufficient if you try to make them more general than you need to. It would be nice if the default models would explicitly list their assumptions, though. Like, a model trained only on correct literary text of one language variant in proper case and not on anything else should clearly state that.
- Eridrus 5y agoI feel like NER is a poorly designed task in general. You're eventually trying to link the entities to some kind of KB, so you should be injecting that entity information into your system for detecting mentions.
- robbedpeter 5y agoRule based processing can augment transformers by both filtering out bad input and by parsing good input into a form that plays to the strengths of a model. You can do some fantastic things with BERT and spaCy, or gpt-neo/J/3, or combinations as needed. Expert systems and ontological tools and things like nltk, spaCy, and LinkGrammar are excellent complements to an ai workflow. Use the fast, "dumb" tools to do the fast, dumb tasks, and only use the huge smart models when you need it. GPT-3 shouldn't be used if you're just doing tagging or NER, but you can get higher quality nuanced extrapolation or summarization if you run things through a mad libs style prompt generator that leans into prompts that work really well.
- minimaxir 5y agoA relatively underdiscussed quirk of the rise of superlarge language models like GPT-3 for certain NLP tasks is that since those models have incorporated so much real world grammar, there's no need to do advanced preprocessing and can just YOLO and work with generated embeddings instead without going into spaCy's (excellent) parsing/NER features. OpenAI recently released an Embeddings API for GPT-3 with good demos and explanations: https://beta.openai.com/docs/guides/embeddings https://beta.openai.com/docs/guides/embeddings Hugging Face Transformers makes this easier (and for free) as most models can be configured to return a "last_hidden_state" which will return the aggregated embedding. Just use DistilBERT uncased/cased (which is fast enough to run on consumer CPUs) and you're probably good to go.
- new_stranger 5y agoI imagine it being very useful to understand what you just said
- hooande 5y agolol. a rough translation is that the new super language models are good enough that you don't have to keep track of specific parts of speech in your programming. if you look at the arrays of floating point weights that underlie gpt-3 etc, you can use them to match present participle phrases with other present participle phrases and so forth this is of course a correct and prescient observation. minimaxir is kind of an NLP final boss, so I wouldn't expect most people to be able to follow everything he says
- minimaxir 5y agoI don't think it's more of a final boss thing: IMO working with embeddings/word vectors is easier, even in the basest case such as word2vec/GloVe, to understand than some of the more conventional NLP techniques (e.g. bag of words/TF-IDF). The spaCy tutorials in the submission also have a section on word vectors.
- 5y ago
- master_yoda_1 5y ago
- Ldorigo 5y agoAh, yes. The tried-and-true method of "just selling the hype" with an open source library that everyone can use for free.
- dang 5y ago"Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something." https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- Der_Einzige 5y agoAs usual, dang is wrong and not moderating effectively. This is not a shallow comment but a legitimate concern about spaCy, and to a lesser extent other NLP tools such as NLTK. Most of the tooling around them that people end up using really is nothing more than wrappers around other tools. See the default tokenizers or models utilized by these tools. And yes, even if spaCy is not making money itself, you can bet that the other paid for tools that they sell are.
- master_yoda_1 5y agospaCy makes lots of money. from https://explosion.ai/about https://explosion.ai/about "In August 2021, we sold 5% of Explosion to SignalFire for $6 million. Employees are given a stake in Explosion using a virtual share bonus program."
- dang 5y agoActually if the GP had posted this critique instead of a shallow, reductionist internet dismissal ("just want to sell the hype"), that would have been fine. Thoughtful critique is welcome—it just requires higher-quality comments than that.
- 41209 5y agoI really love spaCy, it's trivial to throw up a server which handles basic NLP. No complaints here, very happy to see it still being updated