4 ms·
This is a much easier problem as the input is large but streamable and the database is small. However, you would still need to use some of the tokenisation ste
by andyjpb 6y ago
This is a much easier problem as the input is large but streamable and the database is small.
However, you would still need to use some of the tokenisation steps that he talks about such as finding word boundaries and stemming.
The performance of your problem would be bounded by the size of your unseen text.
Something like the The Knuth-Morris-Pratt string-search algorithm would be useful.
- zemo 6y agothanks this is a good reference, this was the type of term I was looking for. I can implement this off of a paper. it's a bit tougher than regular string matching because of homoglyph attacks (i.e., people trying to subvert the matching by writing f_ck or f0ck or fȗck or whatever) but I can probably do a pretty straightforward tokenization at ingest time to clean that up.
- 082349872349872 6y agoIf your users don't ever discuss places such as Scunthorpe, England or Fucking, Austria, the tokenisation may be effective.