3 ms·
Agreed. I feel like the "training data must ALSO be open source" argument solely stems from trying to nitpick at Meta's efforts. The training data provides alm
by infotainment 2y ago
Agreed. I feel like the "training data must ALSO be open source" argument solely stems from trying to nitpick at Meta's efforts.
The training data provides almost no value, since it's usually just unstructured text dumped from the internet. However, by demanding it, one creates almost impossible-to-reach goalposts. If, somehow, Meta had also released their training data, I suspect the goalposts would immediately move to something else.
- blackeyeblitzar 2y agoIt’s absolutely valuable. It lets you determine the biases of the model, by examining what it was trained on. It also lets you alter the curation, pre, and post processing to achieve a different model that may be more accurate or truthful (if you can train it). But the transparency is necessary to audit what they do. > If, somehow, Meta had also released their training data, I suspect the goalposts would immediately move to something else. No it wouldn’t. There’s a finite list of what is needed for an LLM to be open source. See an example of this in AI2’s OLMo: https://blog.allenai.org/hello-olmo-a-truly-open-llm-43f7e7359222 https://blog.allenai.org/hello-olmo-a-truly-open-llm-43f7e73...
- papichulo2023 2y agoIt will never happened, no legal team will approve it, endless liability. Also I wonder what will normal users do with dozens of TB of synthetic data.