3 ms·
Enjoyed this. More importantly its an emerging problem we need to urgently solve. Apart from my own desire to read the internet to find out what actual people
by jgord 26d ago
Enjoyed this. More importantly its an emerging problem we need to urgently solve.
Apart from my own desire to read the internet to find out what actual people think .. Im worried about a kind of bit-rot, where any decade now, gen-pop reader has no idea what content came from humans and what came from AI.
Im assuming the people who _train_ AI [ LLMs ] actually need to separate the two, and avoid the feedback loop of training the next LLM on the output of the previous LLM.
Presumably, this would lead to a kind of reversion to the mean, akin to making photocopies of photocopies back in the day, resulting in copy degradation.
We already have a kind of bit-rot from moving away from physical media - many documentaries on TV, and audio / video / movies / games are being lost as they are not moved to permanent archive storage.
Old out of print books should be scanned and made public and permanently available online, as they are part of our cultural heritage [ not least for the purposes of training current and future AI as the defacto archives / oracles ]
- Ifkaluva 26d agoOh they don’t have this data degradation problem at all. They have large data curation teams that scrutinize the “data mix”, ensures the model is always performing better.
- esseph 26d ago> Im assuming the people who _train_ AI [ LLMs ] actually need to separate the two, and avoid the feedback loop of training the next LLM on the output of the previous LLM. Check out Reinforcement Learning from AI Feedback (RLAIF). Then skim some of this maybe: https://arxiv.org/abs/2309.00267 https://arxiv.org/abs/2309.00267