3 ms·
Unfortunately the Tatoeba data set is fundamentally flawed, as they have assumed that there is only one instance of any particular language, when that is clearl
by scoot 2y ago
Unfortunately the Tatoeba data set is fundamentally flawed, as they have assumed that there is only one instance of any particular language, when that is clearly not the case for all languages.
For example, Brazilian Portuguese is very different to European Portuguese in pronunciation, idioms, and to varying degree grammar and vocabulary, to the extent that spoken European Portuguese is largely unintelligible to Brazilian Portuguese speakers. Even Brazilian Portuguese has significant regional difference, and there is no "standard" Brazilian, but at least differentiating at the country level would make the dataset somewhat more useful.
To top it off, they use national flags to represent languages, for example the UK flag to represent English, despite the spoken examples being overwhelmingly American English, or some mashup for example the Brazilian and Portuguese flags slapped together.
- mtalantikite 2y agoYeah, I don't disagree, there are problems with it for sure. But at least for language learning in most languages, as a beginner you're probably just trying to get yourself to a B1-ish level in your target language before switching over to (easy) native content. The quirks of regional pronunciation and idioms can come later. And having some resources is better than having no resources, particularly when we're talking about languages without much in regards to learning materials. That's where a dataset like Tatoeba could be helpful for training ML models or being used as a base for learning apps like Clozemaster. For a language like Portugese, though, you already have Assimil courses and many levels of Pimsleur, etc in both Brazilian and European dialects to get you going. There looks to even be some FSI courses, if you're coming from English. Again, I do agree that it's weird they have both European and Brazilian Portugese mixed on Tatoeba, whereas for Arabic, where many dialects are also not mutually intelligible, there are different categories. I wonder why that happened.