3 ms·
Synthetic and generated data is also real and shouldn't be underestimated. LLM training can be very counterintuitive at times.
by WalterSear 2y ago
Synthetic and generated data is also real and shouldn't be underestimated. LLM training can be very counterintuitive at times.
- Out_of_Characte 2y agoSynthetic and generated data might not be garbage. Consider random chess positions. Those are generated,yet there's an engine out there that has objectively a higher chance of winning from that position. There was also a paper on persistent computer programs after running random data iteratively over millions of cycles. This doesn't even have a 'goal'. which by itself is very interesting because of how similar it feels to how organisms develop by the act of persisting after a few billion years. That is to say I largely agree with the premise that you can train and improve alot by absense of data. But if your original dataset is garbage, then training over and over might get limited by a hidden nash equilibrium that the algoritm can't escape due to lack of new information. Obligatory Anton petrov: https://www.youtube.com/watch?v=L_IWVZPmc-E https://www.youtube.com/watch?v=L_IWVZPmc-E Paper: https://arxiv.org/pdf/2406.19108 https://arxiv.org/pdf/2406.19108