5 ms·
Fast and accurate language identification using fastText
- wyldfire 9y agoI think it would be cool to see how easily they could create a WASM/asm.js target.
- alexott 9y agoI would try to make comparison with Google's CLD tomorrow
- microcolonel 9y agoBearing in mind that fastText supports many more languages than CLD.
- alexott 9y agodepends on the mode, but I've compared with only ~60 languages
- rasmussondk 9y agoPlease let us know how it goes.
- alexott 9y agoI just posted a link to blogpost
- alexott 9y agohttp://alexott.blogspot.de/2017/10/evaluating-fasttexts-models-for.html http://alexott.blogspot.de/2017/10/evaluating-fasttexts-mode...
- allan_s 9y agoNice the author created this based on tatoeba.org data, I used to be the main developer and for tatoeba I created a language detector (because it's was painful for people to have to input a sentence AND the language, especially for polyglots), so it's more likely the language data used for this language detector was made itself by a language detector, funny when you think about it :) https://github.com/allan-simon/Tatodetect https://github.com/allan-simon/Tatodetect (I should rewrite it in Rust some days) , it's a simple N-gram detector.
- visarga 9y agoWhy is it just 93% accurate on Wikipedia? Is it that hard to identify languages?
- microcolonel 9y agoI suspect it's due to mixed-language content on Wikipedia. A lot of Wikipedia articles talk about foreign language art and culture, this is one of the largest (if not the largest) single categories of content on non-English Wikipedias.
- alexott 9y agoYes, it's not so good on the samples with several languages
- matthberg 9y agoReally fascinating from a linguistics perspective, I'm curious as to how this works and if it is possible to abstract away to help with the cataloguing of dying languages.