3 ms·
> There are massive proprietary datasets out there which people avoid using for training due to copyright concerns. The main legal concern is their unwillingne
by oldgradstudent 2y ago
> There are massive proprietary datasets out there which people avoid using for training due to copyright concerns.
The main legal concern is their unwillingness to pay to access these datasets.
- zozbot234 2y agoYup, there's also a huge amount of copyright-free, public domain content on the Internet which just has to be transcribed, and would provide plenty of valuable training to a LLM on all sorts of varied language use. (Then you could use RAG over some trusted set of data to provide the bare "facts" that the LLM is supposed to be talking about.) But guess what, writing down that content accurately from scans costs money (and no, existing OCR is nowhere near good enough), so the job is left to purely volunteer efforts.