3 ms·
> For a novice it looks a pretty simple job of using some Fuzzy string matching tools and get this done. Nice. That annoying quip aside, as with many things i
by CorvusCrypto 8y ago
> For a novice it looks a pretty simple job of using some Fuzzy string matching tools and get this done.
Nice.
That annoying quip aside, as with many things in data processing, it's case to case. In TF-IDF you lose ordering information by definition. This is probably fine for this use case but it does mean that if ordering does matter since a set of stores share the same words.in different ordering, this will fail to resolve the difference. The author says he did due diligence on the data but there are other ways this can fall short. For example ["Walmart", "5280"] compared to ["Store", "5280"] is going to not be so similar as one would want due to the down-weighting of the identifying number in TF-IDF. So imo the disadvantage mentioned for using BoW over TF-IDF is actually not a disadvantage sometimes. As with everything it depends on your problem and data.
To the author I would hope in the future you remove statements like "to a novice, it seems easy to use X". There is nothing novice about going into a problem with an idea and trying it if it seems to fit the use case.