4 ms·
Imagine a single, 120 minute movie in a tar file. How much of this file's raw data would be encoding and metadata, vs the content of the movie?
by optimalsolver 4y ago
Imagine a single, 120 minute movie in a tar file.
How much of this file's raw data would be encoding and metadata, vs the content of the movie?
- decremental 4y agoThe thing is even the video data outside of the tar file is also encoded. Most likely the compressed video data will be basically random. You can't train it on random data, it's just noise. Rather it would make more sense to train on sequences of RGBA pixels. It's a seductive thought to be able to just throw raw bits at a model, regardless of what those bits represent, and have it just magically attain LLM qualities in reproducing the data you would want it to. Something to think about: GPT3/ChatGPT tokenize at the byte level. If they tokenized at the bit level the model would learn Utf8 encoding over time. Unicode characters that require more than one byte to represent, such as emojis, are not learned directly but the model can still reproduce them.