3 ms·
It is important to distinguish two use cases of phonetic algorithms: - Recovering English spelling from English pronunciation. Soundex and Metaphone are good a
by mci 7y ago
It is important to distinguish two use cases of phonetic algorithms:
- Recovering English spelling from English pronunciation. Soundex and Metaphone are good at this.
- Undoing the torments of Eastern-European sibilants in various transliterations of surnames, e.g. Tchaikovsky, Tschaikowski, Tchaïkovski, Ciajkovskij, Tsjaikovski, Tjajkovskij, Tsjajkovskij, Csajkovszkij, and Czajkowski. Here, Double Metaphone or Daitch–Mokotoff Soundex do a better job.
[1] https://en.wikipedia.org/wiki/Metaphone https://en.wikipedia.org/wiki/Metaphone
[2] https://en.wikipedia.org/wiki/Daitch%E2%80%93Mokotoff_Soundex https://en.wikipedia.org/wiki/Daitch%E2%80%93Mokotoff_Sounde...
- mhd 7y agoOr Beider-Morse. Solr, for example, supports quite a few phonetic matching algorithms by default[1]. [1]: https://lucene.apache.org/solr/guide/7_4/phonetic-matching.html https://lucene.apache.org/solr/guide/7_4/phonetic-matching.h...
- mtts 7y agoThe second category essentially concerns spelling variations. Years and years ago I built a tools for searching old documents and to deal with all the spelling variants these contain I implemented something called the Gloria Guts algorithm, which ranks words based on how much their spelling differs. As I recall it worked much better than sounded for our data set