8 ms·
This is going to cause problems as many Mechanical Turk tasks are being used to train AI models...
by bmmayer1 3y ago
This is going to cause problems as many Mechanical Turk tasks are being used to train AI models...
- 1323portloo 3y agoI think this is wonderful. It's a human-bootstrapped AI centipede with predictable, inevitable results.
- logicprog 3y agoTwenty fingered monster hands and nonsense garbled word salad here we come! The AI oroborus begins
- boesboes 3y agoAnd even fine-tuning using AI causes the models to quickly degrade I've seen here previously. Perhaps we are already at peak of these LLMs now and the next gen won't come because all training data will be poisened with AI generate crap. That would be kinda funny to me tbh
- sillysaurusx 3y agoThe invention of nukes caused a similar problem. Scientists had to harvest metal from pre-ww2 sunken battleships, since that’s the only metal not contaminated with some measurable amount of radiation. I forget the details, but it’s very likely pre-chatgpt training data will be prioritized, with exceptions for news and current events (which largely doesn’t matter whether it’s AI generated anyway).
- TZubiri 3y agoAlso happens with wikipedia. Later edits with web references are much more prone to citogenesis, while an early edit with a web reference is probably clear, even if the link is dead. Even printed references can be based on wikipedia now
- deleted 3y ago[deleted]
- CharlesW 3y ago> …citogenesis… https://en.wikipedia.org/wiki/Circular_reporting#Circular_reporting_on_Wikipedia https://en.wikipedia.org/wiki/Circular_reporting#Circular_re...
- imtringued 3y agoMissed opportunity to have the article apply to itself. E.g. someone writes a made up citogenesis article on Wikipedia before the xkcd comic was drawn.
- cwillu 3y agohttps://en.wikipedia.org/wiki/Low-background_steel https://en.wikipedia.org/wiki/Low-background_steel Note also: “Since the end of atmospheric nuclear testing, background radiation has decreased to very near natural levels,[3] making special low-background steel no longer necessary for most radiation-sensitive applications, as brand-new steel now has a low enough radioactive signature that it can generally be used in such applications.”
- BoxOfRain 3y agoApparently a lot of it comes from what remains of the German fleet that scuttled itself at Scapa Flow rather than be divided among the allies. Most of the ships were salvaged between the world wars but some remain.
- vintermann 3y agoI have a feeling that some types of data will become a bit like "the DNA wealth of the rainforest", valuable because it's so rich, different and rare. In particular small and/or dying human languages. Languages carry with them an immense amount of information, from all the precious human experiences that have shaped them.
- fatherzine 3y agoLanguages at fine granularity and narrative complexes at high granularity. Some people explicitly refer to stories as the most efficient compression mechanisms there is, by a large margin. There is a reason the Hutter prize is what it is. Which, in a supreme irony for the positivist types, rationally entails that folk tales and the Bible (gasp) are the most valuable data there is.
- logicprog 3y agoWell, not necessarily, because while it may be highly compressed data shaped by a long history, we don't know if that data is actually any good. It could be the compounded misapprehensions of all the peoples that contributed to it. Just the fact that something holds a lot of highly efficiently compressed data doesn't mean it's actually good data. You might argue that if it's old and still common today then it must have guided the people who followed it well in order to survive so long, engaging in a kind of natural selection of ideas, but the thing is that ideas can also fall prey to suboptimal equilibria and all sorts of other evolutionary weirdness and badness. Like, the success of an idea is not directly related to how good the idea actually is for the people that follow it. As long as it isn't bad enough to kill off all its followers, an idea can spread by many other means. Maybe it's better at spreading, through appealing to our baser natures or cognitive biases, maybe it's good at locking itself in through thought stopping cliches and fear of reprisal and stuff like Pascal's Wager/Roko's Basilisk. Or maybe it's really good at setting up interlocking peer pressure and network effects and social enforcement of continuing to believe it. Or maybe it's better at suppressing other ideas. Or maybe it's better at creating a group geared toward conquest or evangelism, but actually living under it sucks otherwise. Maybe it's just really beneficial to a small group of people in power and so they work really hard to spread it. Not to mention, it only had to compete against the actual ideas that existed contemporary with ut in history and were strong enough to create evolutionary pressure, so if a new idea comes along about how to organize society or whatever, we don't actually know if it will automatically win out just because it won out against other different ideas in the past. What would we think an idea is unimprovable or optimal just because it is old? The long term existence and popularity of an idea or tradition doesn't necessarily make it good at all. Would you say that the Bible's injunctions against e.g. divorce or homosexuality or women being able to teach and hold positions of authority are encoded wisdom? The most valuable data there is? I'm by no means an actual positivist, which is an incoherent position, but I am a pragmatist, and I think the only way to actually figure out if a tradition is worth anything is to look at its reasoning, its assumptions, its history, and its effects in the world.
- the8472 3y agoThat seems unlikely to me. Models are multi-modal now, so they'll likely train on video and audio content too. Once that's tapped out they can learn better world models by not passively ingesting data but interacting with it, e.g. via simulated environments or capturing human-bot interactions that are flagged as interesting and incorporating those in the next run.
- causalmodels 3y agoHard disagree. Phi-1 seems to be a harbinger of what is to come. However i think there is a plausible argument to be made that their training process approximates a kind of distillation
- petesergeant 3y ago> even fine-tuning using AI causes the models to quickly degrade Really? My recent reading[0, 1] suggests AI training with AI-generated content provides excellent results, but happy to see some counter examples 0: https://arxiv.org/abs/2305.07759 https://arxiv.org/abs/2305.07759 1: https://arxiv.org/abs/2306.11644 https://arxiv.org/abs/2306.11644
- stefan_ 3y agoAlso of course for "science", as we learned from the latest fake scandal, lots of social science researchers base their studies on MTurk tasks..
- TZubiri 3y agoI liked the comparision to taking a jpeg of a jpeg.
- Salgat 3y agoYes and no. You can utilize ML to do 90% of the work in preparing an input, then have a human verify to ensure it's correct. It's similar to how ML models are starting to consume ML generated artwork, which people think is bad, but in reality those ML generated artworks are the ones humans found good enough to publish online, so they still retain value for training (as being appealing to humans).