4 ms·
I'm always happy to see the proliferation of open-source resources for the next generative models. But I strongly suspect that OpenAI and friends are all using
by nyyp 2y ago
I'm always happy to see the proliferation of open-source resources for the next generative models. But I strongly suspect that OpenAI and friends are all using copywritten content from the wealth of shadow book repositories available online [1]. Unless open models are doing the same, I doubt they will ever get meaningfully close to the quality of closed-source models.
Related: I also suspect that this is one reason we get so little information about the exact data used to train Meta's Llama models ("open weights" vs "open source").
[1]: https://www.annas-archive.org/llm https://www.annas-archive.org/llm
- ilrwbwrkhv 2y agoOf course. This isn't even a suspicion. They are using any publicly accessible dataset especially the piracy hoards.