4 ms·
I don’t think they’re using picture heavy book for LLM training, no?
by mateus1 2y ago
I don’t think they’re using picture heavy book for LLM training, no?
- moralestapia 2y agoYes they do, there's multimodal models.
- mnsu 2y agoFor multi-modal models, why not? They would be probably some of the best data.
- michaelt 2y agoSometimes the PDF of a book is big because the book's packed with important illustrations and charts - like a textbook or journal paper. Other times a PDF of a book is big because someone scanned it and didn't have trustworthy OCR, so they figured distributing images of text at 1.5 MB per page was better than risking OCR errors.
- WithinReason 2y agoPresumably they didn't create the torrent
- littlestymaar 2y agoEven if they didn't use the illustration(which isn't clear given multimodal models), they'd still make use the text in the books.
- deleted 2y ago[deleted]
- rbanffy 2y agoI don't think they need to be selective. It's not like Meta can run out of storage.
- RIMR 2y agoJust because the LLMs are trained on text doesn't mean that images we're a part of what they downloaded. You clean up the data after you acquire it, not before.
- hulitu 2y agoWhy not ? Do you think that AI doesn't enjoy porn ? /s