3 ms·
This is probably not entirely invalidating the result, but the language samples in the dataset seem to be extremely badly translated from english, with unnatura
by fvdessen 3y ago
This is probably not entirely invalidating the result, but the language samples in the dataset seem to be extremely badly translated from english, with unnatural, verbose and grammatically wrong sentences. That would not help with good tokenisation.
For example the english text
> please add milk to the grocery list
Is compared to the french text
> s'il vous plaît ajouter du lait à la liste d' épicerie
But a native would say
> veuillez ajouter du lait à la liste de courses
- yenniejun111 3y agoThis is a really good point! I also noticed that some of the translations were not good or very stilted for the languages I do speak. However, this is a limitation of the dataset of this size and breadth