3 ms·
Hey there! I'm the first author and I'm very shocked to see this on HN. To answer your question -- the distribution change that the models should be robust to i
by ilovenlp 4y ago
Hey there! I'm the first author and I'm very shocked to see this on HN. To answer your question -- the distribution change that the models should be robust to is the multilingual lexicon, as discussed with the Singlish example. We know that a speaker's linguistic background can influence how they communicate in Creole. So with the Singlish example, a sentence with the same meaning can be really different if its a person who also speaks Mandarin, versus someone who speaks Malay. Maybe in a training set, we'll see that a dataset is actually biased towards examples containing Mandarin, and it then becomes important to be robust towards the lesser represented Malay then, for example.
Singlish and Nigerian Pidgin are a really good examples of Creoles where we would expect robustness to multilingual vocabulary to be important, as these Creoles are are "linguae francae" within their respective countries (although, these two languages have VERY different levels of acceptance within their respective societies). Looking back on this work, Haitian Creole makes less sense, as Haiti is largely a monolingual country, with only a few also speaking French. But with the history of the language, the vocabulary is still mixed between French origin and like Akan/Igbo/Yoruba etc.
Hope this helps, thanks for taking a look :-)