3 ms·
They could release as much is as necessary to recreate it, the crawlers or list of links they used and configuration or scripts used to drive the training. Nobo
by guerrilla 2y ago
They could release as much is as necessary to recreate it, the crawlers or list of links they used and configuration or scripts used to drive the training. Nobody is asking for the entire web in their git repo, only the ability to retrain from scratch, possibly with modifications.
- echoangle 2y agoNot really, because there’s no guarantee that it will be available in the future. A script to download the data doesn’t mean I can reliably recreate the data in 5 years, I wouldn’t call that open source. To me, the data itself needs to be published.
- guerrilla 2y agoOh well, they did their best. You can't expect them to do better than what is possible. Enough Nirvana fallacy here.
- echoangle 2y agoI’m not expecting them to do the impossible, but they shouldn’t call it open source then. Either you provide all the data and call it open source, or you don’t provide the data because it is proprietary and don’t call the model open source.
- guerrilla 2y agoWell, I'm glad you exist to push the Overton window even further anyway. A lot of people are trying to claim that what's being pushed now (opaque data) is open source. I'll be satisfied if I can at least aplroximate the training witb whatever is online at the time I were to do it.
- tourmalinetaco 2y agoThis is incredibly unrealistic. Imagine hundreds to thousands of individuals scraping thousands of websites, DDoSing them because their mid-tier “open source” LLM project gave them a list of links to “recreate” their dataset. It’s far more sustainable to create a dataset filled GPL, public domain, and other permissively licensed data rather than nuke half the Internet’s bandwidth. And that’s ignoring the fact that scraping it yourself does not actually grant you a permissive license, that’s like saying that by watching a movie in theaters you have a legal right to sell that movie.