3 ms·
I’ve found phonetic similarly algorithms like this to be very useful for data cleaning work, especially when your data generating process involves transliterati
by edraferi 7y ago
I’ve found phonetic similarly algorithms like this to be very useful for data cleaning work, especially when your data generating process involves transliteration. That is, if your users know Language A but your data involves Language B, your users will record terms from Language B with highly varied spellings that nonetheless sound very similar when pronounced. Phonetic similarity algorithms help you cluster these spellings together by sound so you can map them to a canonical term from Language B.