4 ms·
The title is clickbait. The actual study[0] reads as follows: > Multi-way parallel, machine generated content not only dominates the translations in lower reso
by chacham15 3y ago
The title is clickbait. The actual study[0] reads as follows:
> Multi-way parallel, machine generated content not only dominates the translations in lower resource languages; it also constitutes a large fraction of the total web content in those languages
This is talking specifically about text being translated into less common languages being done mostly by machine translation and that in those languages most of the content found was as a result of machine translation. This has nothing to do with English or more common languages overall.
[0] https://arxiv.org/pdf/2401.05749.pdf https://arxiv.org/pdf/2401.05749.pdf
- hhs 3y agoTo be fair, the author does link to that study in the first paragraph of this piece, and then adds some context about languages near the end: “But while the English-language web is experiencing a steady — if palpable — AI creep, this new study suggests that the issue is far more pressing for many non-English speakers. What's worse, the prevalence of AI-spun gibberish might make effectively training AI models in lower-resource languages nearly impossible in the long run. To train an advanced LLM, AI scientists need large amounts of high-quality data, which they generally get by scraping the web. If a given area of the internet is already overrun by nonsensical AI translations, the possibility of training advanced models in rarer languages could be stunted before it even starts.”
- figassis 3y agoSo it is AI generated slime for people in those languages.