4 ms·
I work with full text search where this is common. Here is some points. Stemming: Reducing words to their base or root form (e.g., “working,” “worked” becoming
by tmikaeld 2y ago
I work with full text search where this is common. Here is some points.
Stemming: Reducing words to their base or root form (e.g., “working,” “worked” becoming “work”).
Lemmatization: Similar to stemming, but more sophisticated, accounting for context (e.g., “better” lemmatizes to “good”).
Token normalization: Standardizing tokens, such as converting “wrk” to “work” through predefined rules (case folding, character replacement).
Fuzzy matching: Allowing approximate matches based on edit distance (e.g., “wrk” matches “work” due to minimal character difference).
Phonetic matching: Matching words that sound similar, sometimes used to match abbreviations or common misspellings.
Thesaurus-based search: Using a predefined list of synonyms or alternative spellings to expand search queries.
Most of these are open and free lists you can use, check the sources on manticore search for example.
- soared 2y agoPorter stemming is currently widely used in adtech for keywords.
- thaumasiotes 2y ago> Lemmatization: Similar to stemming, but more sophisticated, accounting for context (e.g., “better” lemmatizes to “good”). I don't understand. How is that different from stemming? What's the base form of "better" if not "good"? The nature of the relationship between "better" and "good" is no different from that between "work" and "worked".
- authorfly 2y agoStemming is basically rules based on the characters. It came first. This is because most words in most languages follow patterns of affixes/prefixes (e.g. worse/worst, harder/hardest), but not always (good/better/best) The problem was that word/term frequency based modelling would inappropriately not linked terms that actually had the same route (stam or stem). Stemming removed those affixes so it turned "worse and worst" into "wor and wor" and "harder/hardest" into "hard", etc. However it failed for cases like good/better. Lemmatizing was a larger context and built up databases of word senses linking such cases to more reliably process words. So lemmatizing is rules based, plus more.
- thaumasiotes 2y ago> So lemmatizing is rules based, plus more. Fundamentally, the rule of lemmatizing is that you encounter a word, you look it up in a table, and your output is whatever the table says. There are no other rules. Thus, the lemma of seraphim is seraph and the lemma of interim is interim. (I'm also puzzled by your invocation of "context", since this is an entirely context-free process.) There has never been any period in linguistic analysis or its ancestor, philology, in which this wasn't done. The only reason to do it on a computer is that you don't have a digital representation of the mapping from token to lemma. But it's not an approach to language processing, it's an approach to lack of resources.
- mannykannot 2y agoI see your point about context-free table lookup, but it looks to me as though authorfly's distinctions would apply to how the tables get written in the first place.
- authorfly 2y agoWe don't disagree. A look up table with exact rules is a rules system to me from an NLP/GOFAI perspective. I was aware of how the libraries tend to work because I had often used things like looking up lemmas/word sense/pos in NLTK and Spacy in the past, and I know the libraries code fairly well. Context today may mean more (e.g. the whole sentence, or string, or the prompt context for an LLM), and obviously context has a meaning in computational linguistics (e.g. "context free grammar"), but the point here is stemmers arbitrary follow the same process without a second stage. If a stemmer encounters "best" and "good" it by definition does not have a stage to use the same lemma for them. Context is just one of those overloaded terms unfortunately. Lemmatizing, in terms of how it works on simple scenarios (lets imagine reviews) helps to lump those words together and correctly identify the proportion of term frequencies for words we might be interested in more consistently than stemming can. It's still limited by using word breaks like spaces or punctuation ofcourse.