15 ms·
No Language Left Behind
- btheshoe 4y agoI'm not entirely sure why low resource languages are seen as such a high priority for AI research. It seems that by definition there's little payoff to solving translation for them.
- dunefox 4y agoSmall data, big meaning is much more important than big data, little meaning. Much closer to real intelligence.
- tehsauce 4y agoI think the reason low resource languages are prioritized is to compensate for the fact that AI research normally has a tendency to marginalize these languages.
- btheshoe 4y agoyes, but what principles justify the importance placed on low resource languages?
- froskur 4y agoLow resource in this context means that there are few resources available to train a neural network with, not that there are few speakers. Although many low resource languages have relatively few speakers, there are also ones with tens of millions of speakers. The reason for emphasis is in my opinion twofold: 1) Allowing these people to use the fancy language technology in their own language is good in and of itself. 2) Training neural networks on fewer resources is more difficult than using more resources and therefore a fun and interesting challenge.
- macintux 4y agoPlus presumably we learn more from solving harder problems, and we prepare for one day needing to translate some alien language in a hurry.
- munificent 4y agoCynical answer: It's good PR.
- Jabbles 4y agoSurely the fact that they did all the high-resource languages first and are only now getting round to the less-popular ones demonstrates that that is not, in fact, the case?
- wilde 4y agoThe point is that there are lots of humans who speak these languages and use tech. They just don’t use Wikipedia so getting a good translation corpus going was harder.
- gwern 4y agoAnd it's both cumulative across all those languages (see above), cheap/amortized (if you can do a good multilingual NMT for 50 languages, how hard can 50+1 languages be?), and many of those languages are likely to grow both in terms of sheer population and in GDP. (Think about South Asian or African countries like Indonesia or Nigeria.) The question isn't why are FB & Google investing so much in powerful multilingual models which handle hundreds of languages, but why aren't other entities as well?
- ausbah 4y agowhat other entities would really have access to the text resources that FB & Google? outside of a few other large companies I can't imagine many
- albertzeyer 4y agoI don't really remember the exact numbers anymore, but covering only the top 5 languages will cover maybe 40% of the world population, while covering the top 200 languages (many of them low resource) will cover maybe 90% of the world population. Some numbers (but you can not exactly infer from them such accumulated numbers): https://en.wikipedia.org/wiki/List_of_languages_by_total_number_of_speakers https://en.wikipedia.org/wiki/List_of_languages_by_total_num... Some more numbers from here: https://www.sciencedirect.com/science/article/pii/S0167639313000988?casa_token=Iy1BsJ3AtVUAAAAA:PWHeMDDaC6azNiqWahtWewxYfrqla4JbZgs5MMRCaYRNFX4PoqR_8c3Uwb83fkokwWX5YLqOQi0 https://www.sciencedirect.com/science/article/pii/S016763931... "96% of the world’s languages are spoken by only 4% of its people." Although this statement is more about the tail from the approx 7000 languages.
- jefftk 4y agoIt doesn't sound like you're considering that people are very often fluent in a major language in addition to their regional one?
- albertzeyer 4y agoI am. That's why I mentioned that you can not infer my statements directly from the numbers you find on Wikipedia etc. You can not simply add up those numbers.
- quink 4y agoThe examples given are, with native speaker numbers, Assamese (15 million), Catalan (4 million) and Kinyarwanda (10 million). These alone are more than an Australia. Furthermore, Facebook considers the internet to consist of Facebook and Wikipedia (Zero). I view this as just another extension of their Next Billion initiative, an effort to ensure that another billion people are monopolised by Facebook. That's the payoff.
- goodside 4y ago"Low-resource language" isn't just a euphemism for "language almost nobody speaks". There are many languages that are widely spoken but nonetheless are hard to obtain training data for. Getting something like Wikipedia going for a minority language can be a difficult chicken-and-egg problem because users will use English for its completeness/recency, despite their limited fluency, and the native-language Wikipedia remains neglected. So you can end up in a situation where users use one language for social media and another for news/research, and Facebook is in a unique position to care about the former.
- onurcel 4y agohi @btheshoe, I work on this project in the data part. As others mentioned, the amount of data available for a language is not correlated to the number of speakers of that language, which explains the potential impact of focusing on these.
- jw4ng 4y agoWe think it's important for AI to truly support everyone in the world. A world where AI only serves a subset of the population is not ideal. In machine translation, this means supporting as many language as possible at high quality. We also imagine a future where anyone will be able to communicate with anyone else seamlessly; this also means solving translations for all languages.
- daniel-cussen 4y agoWouldn't that also entail a bot speaking in any language?
- bobsmooth 4y agoText to speech is a separate problem.
- cyphar 4y agoAside from the fact that being able to generalise a model with very little training data is an important AI research problem to solve, language death is a serious concern and is being accelerated due to the fact that many languages are not supported at all by modern technology (leading to "prestige language" pressures that are a known cause of historical language death). For instance, Icelandic is not supported by any modern smartphone platform which has lead to Icelandic natives communicating with each other in English and very little information is translated to Icelandic[1,2]. That being said, I am worried that having translations that are "too good" could also act to accelerate language death as the importance of keeping languages alive will seem less significant (to non-language-nerds) if we can translate works written in that language to any other language with very small datasets. Luckily I'm not convinced that AI models will be able to produce convincing and consistent translations for a long time -- languages are so different in so many ways that I can't see how adding more dimensions and parameters to a model would account for them. [1]: https://youtu.be/qYlmFfsyLMo?t=141 https://youtu.be/qYlmFfsyLMo?t=141 [2]: https://www.nytimes.com/2017/04/22/world/europe/iceland-icelandic-language-linguistics.html https://www.nytimes.com/2017/04/22/world/europe/iceland-icel...
- pesenti 4y agoBlog post: https://ai.facebook.com/blog/nllb-200-high-quality-machine-translation/ https://ai.facebook.com/blog/nllb-200-high-quality-machine-t... Paper: https://research.facebook.com/publications/no-language-left-behind/ https://research.facebook.com/publications/no-language-left-... Github: https://github.com/facebookresearch/fairseq/tree/nllb/ https://github.com/facebookresearch/fairseq/tree/nllb/
- robocat 4y agoAlso note comments from hello_im_angela (= Angela Fan) and jw4ng (= Jeff Wang). Those are the HN accounts for Angela and Jeff from No Languages left Behind.
- kwhitefoot 4y agoWhat is a "low resource language"?
- pesenti 4y agohttps://datascience.stackexchange.com/questions/62868/high-low-resources-language-what-does-it-mean https://datascience.stackexchange.com/questions/62868/high-l...
- jw4ng 4y agohey there, I work on this project. We categorize a language as low-resource if there are fewer than 1M publicly available, de-duplicated bitext samples. also see section 3, table 1 in the paper: https://research.facebook.com/publications/no-language-left-behind/ https://research.facebook.com/publications/no-language-left-...
- maestrae 4y agohey, this sounds silly but I can't seem to find a link of all the languages covered in the 200 hundred languages. I've looked at the website and the blogpost and neither have a readily available link. Seems like a major oversight. There is of course a drop down in both but the languages there are a lot less than 200. I'm particularly interested in a list of the 55 African languages for example.
- hello_im_angela 4y agoWe have a full list here (copy pastable): https://github.com/facebookresearch/flores/tree/main/flores200 https://github.com/facebookresearch/flores/tree/main/flores2... and Table 1 of our paper (https://research.facebook.com/publications/no-language-left-behind/ https://research.facebook.com/publications/no-language-left-...) has a complete list as well.
- maestrae 4y agothank you!
- TaupeRanger 4y agoSo they have a system that can translate to languages for which there isn't as much data as English, Spanish, etc. Waiting for a Twitter thread from a native speaker of one of these "low resource languages" to let us know how good the actual translations are. Cynically, I'd venture that they hired some native speakers to cherry pick their best translations for the story books. But mostly this just seems like a nice bit of PR (calling it a "breakthrough", etc.). I can't imagine this is going to help anyone who actually speaks a random, e.g., Nilo-Saharan language.
- alexott 4y agoTwitter may not be representative imho because of the short text. It should first come to a problem of reliable language detection, and Twitter is quite often wrong there
- hello_im_angela 4y agoIf you're curious to try the system yourself, it's actually being used to help Wikipedia editors write articles for low-resource language Wikipedias: https://twitter.com/Wikimedia/status/1544699850960281601 https://twitter.com/Wikimedia/status/1544699850960281601
- netol 4y agoHow is the license of the models (CC NC) compatible with licenses used in Wikipedia? Did you sign an special agreement with the Wikimedia Foundation?
- onurcel 4y agoin this work we tried to rely not only on automated evaluation scores but also on human evaluation for exactly this reason: we wanted to have a better understanding of how our model actually performs and how it correlates to automated scores.
- mikewarot 4y agoThe analogy I like the most is that they've found the "shape" of languages in high dimensions, and if you rotate the shape for English the right way, you get an unreasonably good fit for the shape of Spanish, again for all the other languages. We're at a point where it's now possible to determine the shape of every language, provided there are enough speakers of the language left who are both able and willing to help. <Snark> Once done, Facebook can then commodify their dissent, and sell it back to them in their native language. </Snark>
- goldemerald 4y agoThe shape analogy doesn't really apply with modern language models. Each word gets its own context dependent high dimensional point. With everything being context dependent, simple transformations like rotations are impossible. A more accurate perception is that any concept expressible in language now has its own high dimensional representation, which can then be decoded into any other language.
- gfaster 4y agoAnyone who knows or is learning another language can easily tell you that the "warping" methodology of MTL is insufficient. There was a really good video by Tom Scott [1] that talked about this but the short version is that there is critical bits of language in context and inferred by speakers. Any accurate MTL needs nearly full context both on the page and in the cultural moment, in addition to probably needing to ask questions of the author. [1]: https://www.youtube.com/watch?v=GAgp7nXdkLU https://www.youtube.com/watch?v=GAgp7nXdkLU
- mikewarot 4y agoSo, if I had a corpus of all the literature from 1800-1850 digitized, the context would be sufficiently different as to be a new language? It seems to me that the happy accident of doing this research at the start of getting all human knowledge digitized is part of the unreasonable effectiveness of this overall technique. Had it happened in 200 years, it might not have worked, right?
- adrianN 4y ago
- jw4ng 4y agoJeff Wang here with my fellow Meta AI colleague Angela Fan from No Languages left Behind, seeing the comments flowing through. If you want to ask us anything, go for it!
- pesenti 4y agoAre all the 200x200 translations going directly or is English (or another language) used as an intermediate for some of them?
- jw4ng 4y agoAll translation directions are direct from language X to language Y, with no intermediary. We evaluate the quality through 40,602 different translation directions using FLORES-200. 2,440 directions contain supervised training data created through our data effort, and the remaining 38,162 are zero-shot.
- dangom 4y agoWhat is the greatest insight you gained and could share with non-experts from working on this project?
- jw4ng 4y agoI gained a deeper understanding of what it truly means to be inclusive. Every language is unique just like everybody and making sure content works for all and including as many people as possible is really really hard, but through this project i'm hopeful we are taking it one step further
- Jabbles 4y ago> Every language is unique just like everybody TBH it just sounds like you've redefined the word "unique".
- mike8889 4y ago[flagged]
- Etheryte 4y agoI'll believe it when I actually see it. I'm a native of a reasonably small language spoken by about a million people and never have I ever seen a good automatic translation for it. The only translations that are good are the ones that have been manually entered, and those that match the structure of the manually entered ones. I think the sentiment is laudable and wish godspeed to the people working on this, but for the time being I don't see it becoming a reality yet. When Google Translate regularly struggles even with big pairs such as German-English-German, I have reservations about someone making it work for languages where datasets are orders of magnitude smaller.
- hello_im_angela 4y agoIt's an extremely difficult problem indeed. A lot of people on the team speak low-resource languages too (my native language as well!), so definitely resonate with what you're saying. My overall feeling is: yeah it's hard, and after decades we can't even do German translation perfectly. But if we don't work on it, it's not gonna happen. I really hope that people who are excited about technology for more languages can use what we've open sourced.
- azinman2 4y ago> But if we don't work on it, it's not gonna happen. That’s exactly right. There’s too much bias in society that if something isn’t perfect, then why bother? Nothing is perfect, so with that attitude there can be no progress. Thank you for doing important work!
- Gigachad 4y agoPersonally I'm hoping that globalisation prunes out as many languages as possible before we end up with brain implants automatically translating everything for us and no one can communicate without these chips.
- pmontra 4y agoBecoming bilingual is one thing. Completely extinguishing a language is a totally different matter. It is usually associated with migrating away from the geographic area of the language and/or physically losing speakers (old age, wars, genocides, etc.) You can check the list at https://en.wikipedia.org/wiki/List_of_languages_by_time_of_extinction https://en.wikipedia.org/wiki/List_of_languages_by_time_of_e... To make an example and be blunt: I do not expect any European country official language to get extinct anytime during our lifespan unless that country gets destroyed, which obviously won't be a good thing. As for brain implants, I won't hold my breath.
- Tabular-Iceberg 4y agoMy concern with this is that in low resource languages the unavoidable biases of the ML models might overpower their own organic development. We shrug off all the little quirks of machine translated text because it usually gets the point across, and we recognize them as quirks because most of what we read was written by real people with no such quirks. But when most of what you read contain those quirks, I fear those will quickly become the standard way of writing and even speaking in those languages.
- texaslonghorn5 4y agoIn a worst case you can end up with the Scots Wikipedia situation, where some power editor created a bunch of pages using an entirely fabricated, overly stereotypical language and that influenced what people thought Scots actually was.
- onurcel 4y agoThis is one of the examples we keep in mind and that's also why we can't 100% trust public dataset labels. This motivated us to train a Language IDentification system for all the languages we wanted to handle in order to build the monolingual dataset. More details in the paper ;) Or here, if you have questions
- protomyth 4y agoI think it will interesting when it runs into a language (e.g. Dakota) where the women and men speak differently. Should be an interesting test.
- zen_1 4y agoDoesn't seem to be a big issue for Arabic, where verbs are gendered (so in the sentence "I am going to the store", the verb "to go" will be either masculine or feminine, reflecting the speaker's gender).
- nemothekid 4y ago
- vjerancrnjak 4y agoWhat are hardware requirements to run this? I see the mixture model is ~ 300 GB and was trained on 256 GPUs. I assume distilled versions can easily be run on one GPU.
- hello_im_angela 4y agoWe release several smaller models as well: https://github.com/facebookresearch/fairseq/tree/nllb/examples/nllb/modeling https://github.com/facebookresearch/fairseq/tree/nllb/exampl... that are 1.3B and 615M parameters. These are usable on smaller GPUs. To create these smaller models but retain good performance, we use knowledge distillation. If you're curious to learn more, we describe the process and results in Section 8.6 of our paper: https://research.facebook.com/publications/no-language-left-behind/ https://research.facebook.com/publications/no-language-left-...
- mdda 4y ago"All models are licensed under CC-BY-NC 4.0" : So, to clarify, does this mean that companies cannot use these models in the course of business, or is it more about selling the translation results directly?
- bvanderveen 4y agoGreat! Facebook no longer have to provide content moderation in all the various corners of the world where they could accidentally enable the dissemination of misinformation and hate speech in minority languages. They can simply transform it into English and run it back through the existing moderation tooling! Understanding foreign culture is about reading automated translations of online comments into your native language. It has nothing to do with putting the effort into learning a language and understanding the nuances and current events and issues of the culture it embeds. The ESL (English as a single language) speakers over at Facebook don't even need to understand foreign cultures, because they already know everyone in the world needs to spend their lives staring into the Metaverse. So grateful that they are working on the world's fattest pipeline for exporting Anglophone culture to every corner of the planet!
- LtWorf 4y agoFacebook translations are horrifying for the mainstream languages already. They go from completely wrong to kinda understandable but still wrong.
- rmbyrro 4y agoLooks like they're investing to get better. The model is also available and they called for contributions to improve it.
- LtWorf 4y agoWhy would I help them? If it was public data sure.
- ShamelessC 4y agoLook, I fucking hate Facebook to the point that I can't really be objective about their research. Whenever I see a section on ethical implications or impacts I just think about shit like Myanmar or the insurrection and laugh (cry). But this is a shallow dismissal that doesn't add anything valuable to the discussion. "Oh they made their _terrible_ (probably state of the art) machine translation _better_??! Those monsters!!"
- microtherion 4y agoAs a native Swiss German speaker, my native language is not only low resource in general, but has the additional difficulty of not having a standardized orthography (many native speakers will exclusively write in Standard German, and use Swiss German only for spoken communication). So you have a language with some economic opportunity (a few million speakers in a fairly wealthy country) but no clearly defined written interface, and an ambivalent attitude of many speakers towards the very idea of writing the language.
- rmbyrro 4y agoThis only makes the problem behind the NLLB project even more interesting to solve
- hello_im_angela 4y agosooo real. Many low-resource languages have many different natural variants, can be written in multiple scripts, don't have as much written standardization, or are mainly oral. As part of the creation of our benchmark, FLORES-200, we tried to support languages in multiple scripts (if they are naturally written like that) and explored translating regional variants (such as Moroccan Arabic, not just Arabic). As an aside, the question of how to think about language standardization is really complex. We wrote some thoughts in Appendix A of our paper: https://research.facebook.com/publications/no-language-left-behind/ https://research.facebook.com/publications/no-language-left-...
- visarga 4y agoAnother avenue for machine translation is to use audio instead of text. There is much more audio data available and being generated on a daily basis, especially for cases like yours it would be very useful.
- tsm 4y agoSimilar issue with Scots, which has many variant orthographies but is frequently written in mostly-English anyway.
- albertzeyer 4y agoNote that very recently Google has done something very similar: "Building Machine Translation Systems for the Next Thousand Languages": https://arxiv.org/abs/2205.03983 https://arxiv.org/abs/2205.03983 https://ai.googleblog.com/2022/05/24-new-languages-google-translate.html https://ai.googleblog.com/2022/05/24-new-languages-google-tr... The Facebook paper has some direct comparison to that work.
- jkw 4y agoEvaluation was important to us, and we really wanted to have a benchmark that covers all 200 languages
- enos_feedler 4y agoI was two sentences in before I realized the headline wasn’t “No Luggage Left Behind”
- onurcel 4y agothis is actually our recurring joke for our team meeting offsites!
- jkw 4y agoHey all, I work on this project. Full list of languages can be found here: https://github.com/facebookresearch/flores/tree/main/flores200 https://github.com/facebookresearch/flores/tree/main/flores2... As well as in the research paper: https://research.facebook.com/publications/no-language-left-behind/ https://research.facebook.com/publications/no-language-left-...
- labrador 4y agoI'll know AI translators are any good when the United Nations starts using them "Skills required: United Nations translators are required to have a perfect command of their main language and an excellent knowledge of, in most cases, two other official languages" https://www.un.org/dgacm/en/content/translation https://www.un.org/dgacm/en/content/translation
- samatman 4y agoAn organization built out of pure prestige, with no concept of monetary profit, has zero pressure to stop employing their classmates as translators, ever.
- deleted 4y ago[deleted]
- epolanski 4y agoMy ex is a translator at an embassy, and she always said that ai translators are a godsend. On one side they make their work easier as they can focus more on correcting the ai produced text and focus on author's meaning while eliminating lots of plumbing. On the other hand they increased the amount of business because much more text is translated than at any other point in history, which requires validation in most business, legal and even personal contexts. Without ai translators those translations would've not happened in the first place.
- labrador 4y agoI'm surprised this didn't occur to me until after I posted because it fits with my general feeling that AIs will be nothing more than collaborative tools for the foreseeable future.
- astrange 4y agoMost media translators consider MTL worse than nothing, because editing it is actually harder than just doing it yourself. Can especially be an issue for neural MTL because the output is both fluent (looks natural) and inaccurate.
- zzzeek 4y agoSo glad it's Facebook doing this and not some other weird company, when translating and delivering information to every culture on the planet it's good to have a trustworthy, ethical company without any past (or heck, even any current, ongoing) issues in spreading misinformation around the globe and contributing to the rise of fascism across the world while profiting massively off of it and denying any culpability, making sure it all goes smoothly.
- pdonis 4y agotl/dr: Now your words can be misconstrued by far more people than before, because AIs will translate the misunderstandings into as many languages as possible.
- NoInkling 4y agoI know DeepL doesn't do low-resource languages, but it would be interesting to see a translation quality comparison between the two.
- otreblatercero 4y agoNot a single mesoamerican language is present. Maya, Náhuatl, Otomí, Zapoteco, etc. And these languages are big, they are spoken by millions and even have literature. Náhuatl and Maya are spoken in Central America.
- bertil 4y agoAre there online corpora, like Wikipedia, that could be used to train the models? Are those under a permissive enough license to be used for model training? If there are spoken, with enough budget, a library of voices could be recorded. I think you’d prefer that collection to be gathered and maintained by a non-profit rather than Meta.
- otreblatercero 4y agoFor náhuatl, I found this: Wikipedia in nahuatl https://nah.wikipedia.org/wiki/Cal%C4%ABxatl https://nah.wikipedia.org/wiki/Cal%C4%ABxatl
- bertil 4y agoI’m wondering if 7065 articles is enough to train the model.
- yellowapple 4y agoHopefully the Scots language model wasn't trained on Wikipedia.
- schoen 4y agoI wonder if spy agencies have already developed, but not published, high-quality SMT methods for lots of minority and little-known languages. :-( (Edit: and speech-to-text models.)
- kgeist 4y agoI wonder how it differs from what Yandex.Translate did back in 2016: [0] >The affinity of languages allows one common model to be trained for their translation. That is, “under the hood” of the translator, the same neural network translates into Russian from Yakut, Tatar, Chuvash and other Turkic languages. This approach is called many-to-one, that is, "from many languages \u200b\u200binto one." This is a more versatile tool than the classic bilingual neural network. And most importantly, it is the many-to-one approach that makes it possible to use knowledge about the structure and vocabulary of the Turkic languages, learned on the rich material of Turkish or Tatar, to translate languages like Chuvash or Yakut, which are less “resource-rich”, but no less important for the cultural diversity of the planet. >In order to create a unified model for translating Turkic languages, Yandex developed a synthetic common script. Any Turkic language is translated into it, so that, for example, the Tatar “dүrt” (“four”) written in Cyrillic becomes similar to the Turkish dört (“four”), not only from the point of view of a person, but also at the level of similarity of lines for a computer. This way they added support for Turkic and Uralic languages which are very underrepresented on the Internet. But I don't know what the quality of their translation is: even though I live in a region where Mari is spoken (indigenous Uralic language) and my wife is Mari, none of us, sadly, speak the language. [0] https://techno-yandex-ru.translate.goog/machine-translation/on-small-languages?_x_tr_sl=ru&_x_tr_tl=en&_x_tr_hl=ru&_x_tr_pto=wapp https://techno-yandex-ru.translate.goog/machine-translation/...
- hello_im_angela 4y agoWe represent all languages in their natural script, rather than transliterating them into a common synthetic one. Regarding Mari: extremely interesting language, exciting to hear that you are from that region. We are interested in working on this one (likely in the "Hill Mari" variant), but currently do not support it.
- Groxx 4y ago>REAL-WORLD APPLICATION >Translating Wikipedia for everyone Hmmm. While there is very definitely utility in doing things like this, I do kinda fear "poisoning the well"-like effects of feeding (even partially-) AI-generated-data into extremely common AI-data-sources. There's some info on it in a blog post[1] and the MediaWiki "Content translation" page[2], but does anyone know of any studies on the quality of the translations produced? I can absolutely see it being a huge time-saver for people who are essentially fluent in both (there's a lot of semi-mechanical drudgery in translating stuff like this that could be mostly eliminated)... but people are pretty darn good at choosing the easy option of trusting whatever they're given rather than being as careful as they should be. It kinda feels like it runs the risk of passively encouraging people to trust the machine's choice over their own, as long as it isn't obviously nonsense, and the cumulative effect could be rather large after a while. [1]: https://diff.wikimedia.org/2021/11/16/content-translation-tool-helps-create-one-million-wikipedia-articles/ https://diff.wikimedia.org/2021/11/16/content-translation-to... [2]: https://www.mediawiki.org/wiki/Content_translation https://www.mediawiki.org/wiki/Content_translation
- jhugo 4y agoYeah, I really hope they don't do this. I live in a country where I don't speak the language well, so I am using Google Translate and DeepL [0] all day every day. The quality of translations of real-world text is so incredibly variable. There is literally no way to know when it will suddenly reverse the meaning of a sentence, or produce something that sounds like it makes sense, but in terms of meaning bears no relation to the input at all. A machine-translated Wikipedia would not be a trustworthy source of information at all, yet would look like one. I think that does significantly more harm than good. [0] Suggestions for better alternatives welcomed.
- debesyla 4y agoOn top of that - a lot of language specific content has to include sources in that same language. (As an example, it would be absurd for lithuanian wikipedia to include sources in japanese - that would be not usable AND not usefull for the wikipedia readers, editors...)
- thamer 4y agoDoes this mean that Facebook's advertising system will finally start rejecting ads calling for genocide in Myanmar, and that they will finally flag comments expressing the same intent? As recently as March of this year there were reports that Facebook accepted ads that said "The current killing of the Kalar is not enough, we need to kill more!" or "They are very dirty. The Bengali/Rohingya women have a very low standard of living and poor hygiene. They are not attractive". Full story: https://abcnews.go.com/Business/wireStory/kill-facebook-fails-detect-hate-rohingya-83576729 https://abcnews.go.com/Business/wireStory/kill-facebook-fail... These were submitted to test Facebook's systems, because there's a good reason not to trust their promises on this front. Facebook was used extensively to propagate hate speech in Myanmar during the crisis of 2017, with their moderation tools and hate speech detection system letting through a ton of hateful content with real-world consequences, in the course of an actual ethnic cleansing campaign. Other references: "Facebook Admits It Was Used to Incite Violence in Myanmar" https://www.nytimes.com/2018/11/06/technology/myanmar-facebook.html https://www.nytimes.com/2018/11/06/technology/myanmar-facebo... (2018) "Violent hate speech continues to thrive on Facebook in Myanmar, AP report finds" https://www.cbsnews.com/news/myanmar-facebook-violent-hate-speech-thrives-ap-report/ https://www.cbsnews.com/news/myanmar-facebook-violent-hate-s... (9 months ago)
- bertil 4y agoThe issue here wasn’t that Facebook didn’t have resources for a basic translation tool (able to translate open death threats) but that Burmese had inconsistent encoding. That delayed the translation effort. https://www.localizationlab.org/blog/2019/3/25/burmese-font-issues-have-real-world-consequences-for-at-risk-users https://www.localizationlab.org/blog/2019/3/25/burmese-font-...
- account42 4y ago> Essential cookies > These cookies are required to use Meta Products. They’re necessary for these sites to work as intended. What cookies does Facebook "need" to serve a simple article?
- _nalply 4y ago"No Language Left Behind" - really? Did the people at Meta think about the Signed Languages of the Deaf? I didn't find a mention. Even Ctrl-F deaf didn't yield anything.
- langsoul-com 4y agoSo so many words but not a hint of any demo. It's just magic according to Facebook. Plz couldn't they at least have a crappy demo to break?