4 ms·
Your script has what looks to be a pretty big inefficiency here[1] that's slowing it down. Looping over a set kills the constant-time presence-checking that you
by haikuginger 9y ago
Your script has what looks to be a pretty big inefficiency here[1] that's slowing it down. Looping over a set kills the constant-time presence-checking that you get from using that data structure; it will likely be much faster to do something like this:
tags = list(set(descr.split()) & categories)
The following would also work:
tags = [x for x in descr.split() if x in categories]
[1]https://gist.github.com/28mm/9820bd8b6eb27555efe9d6f46dd95a81#file-gistfile1-txt-L97 https://gist.github.com/28mm/9820bd8b6eb27555efe9d6f46dd95a8...
- 28mm 9y agoAh, interesting observation. I’ll look at changing it to something more like what you’ve suggested. If memory serves, the reason it doesn’t first split the description on white space is that some categories contain whitespace, and would never match. Thanks!
- haikuginger 9y agoIf you need to check against multiword tags, I'd suggest a utility function to expand a list of words into each possible one-or-more-word subset. Should still be substantially faster than the current state, and you can improve it even more by limiting it to phrases with no more words than the tag with the maximum number of words. def get_all_phrases(descr): words = descr.split() if len(words) == 1: return words phrases = [] for i in range(2, len(words) + 1): phrases += get_phrases_of_len(i, words) return words + phrases def get_phrases_of_len(length, words): return [' '.join(words[i:i+length]) for i in range((len(words) - length) + 1)]