8 ms·
Would there be any benefit in modeling raw binary sequences rather than tokens? I think text prediction only gets you so far. But I guess you could use the sam
by optimalsolver 4y ago
Would there be any benefit in modeling raw binary sequences rather than tokens?
I think text prediction only gets you so far. But I guess you could use the same principles to predict the next symbol in a binary string. If this binary data represents something like videos of physical phenomena, you might get the AI to profound, novel insights about the Universe just with next-bit prediction.
Hmmm, maybe even I could code something like that.
- visarga 4y agoSo you are proposing a massive video model, on the likes of GPT-3? The architecture is simple, but making it train correctly and efficiently is really hard, especially for video.
- marmadukester39 4y agoIs it? Videos are just sequences of frames
- rdedev 4y agoEach frame of the image would have to be divided into many sequences. Atleast that's how transformer based image models work. Then you have to account for audio data too in the same way. It just blows up the compute required
- optimalsolver 4y agoNot quite. I meant something that models pure binary sequences, not higher level tokens. That way, it could learn from any source that can be represented as binary data. Could be video, text, audio, or all three at once. It wouldn't be "video model", it would be an "anything that can be expressed in binary" model.
- visarga 4y agoMaybe you are interested in this paper: > Perceiver: General Perception with Iterative Attention Biological systems perceive the world by simultaneously processing high dimensional inputs from modalities as diverse as vision, audition, touch, proprioception, etc. Perceiver is a deep learning model that can process multiple modalities, such as images, point clouds, audio, and video, simultaneously. It is based on the transformer architecture and uses an asymmetric attention mechanism to distill a large number of inputs into a smaller latent bottleneck. This allows it to scale to handle very large inputs and outperform specialized models on classification tasks across various modalities. https://arxiv.org/abs/2103.03206 https://arxiv.org/abs/2103.03206
- optimalsolver 4y agoThanks! This looks really interesting.
- deleted 4y ago[deleted]
- dwaltrip 4y agoI’m waiting for someone to make a GPT-style model trained for video and audio prediction (e.g. frame by frame, perhaps) in addition to the existing text prediction. Imagine using a significant percentage of YouTube content, for example. It would probably be insanely expensive. But I feel like it would be almost guaranteed to acquire a world model far richer and more robust than ChatGPT’s. Human babies learn by watching the world around them. Video frame prediction feels much closer to that than text prediction, and given the wildly impressive results we are seeing with large text prediction models alone, it seems like an obvious next step.
- boredemployee 4y ago>> Human babies learn by watching the world around them. While I understood what you meant, I'd just add that babies learn by a combination of multisensory triggers (so not only _watching_)
- deleted 4y ago[deleted]
- tintor 4y agoThere is VideoGPT paper. It uses small frame sizes, but that will improve with time.
- TomSwirly 4y ago> a significant percentage of YouTube content [...] a world model far richer and more robust than ChatGPT’s There are three objections to this. The first is the astonishingly large amount of CPU power that this would take, given how high bandwidth video is. The second is that it is hard to believe that some thing really coherent could emerge from this, and certainly, it has never been shown. The third is that the world might be "richer" in terms of information-rich but seeing the world through YouTube's eyes would likely be a degrading and incoherent experience.
- decremental 4y agoThat might be too low(?) resolution. It would be learning encodings instead of features of the thing that is being encoded. Like training it on terabytes of zip files and expecting it to reproduce from the files contained in the archives.
- optimalsolver 4y agoImagine a single, 120 minute movie in a tar file. How much of this file's raw data would be encoding and metadata, vs the content of the movie?
- decremental 4y agoThe thing is even the video data outside of the tar file is also encoded. Most likely the compressed video data will be basically random. You can't train it on random data, it's just noise. Rather it would make more sense to train on sequences of RGBA pixels. It's a seductive thought to be able to just throw raw bits at a model, regardless of what those bits represent, and have it just magically attain LLM qualities in reproducing the data you would want it to. Something to think about: GPT3/ChatGPT tokenize at the byte level. If they tokenized at the bit level the model would learn Utf8 encoding over time. Unicode characters that require more than one byte to represent, such as emojis, are not learned directly but the model can still reproduce them.
- ttul 4y agoGoogle Research has a character-based transformer that learns to tokenize text rather than relying on hand coded tokenizers. It demonstrates superior performance on a variety of LLM tasks. If you have the money, you can apply the transformer architecture to many different tasks and people are experimenting all the time. I think one of the big challenges is always to come up with methods for training such enormous models pragmatically without cost exploding. [1] https://huggingface.co/docs/transformers/model_doc/canine https://huggingface.co/docs/transformers/model_doc/canine
- nodemaker 4y agoThe reason this works for tokens is that tokens are put in a vector space where similar words are in a similar place. The same effect could not be achieved with characters or bits. If you think about it our brains also remember words and not characters
- eternalban 4y ago> our brains also remember words Word sounds. I can not read without hearing the word. (Now I wonder about those born deaf.) Based on that subjective experience which I presume is rather universal among the hearing, tokenized phones and phonemes seem promising.
- worik 4y agoI am not deaf. I do not hear words as I read. (As I write I do)
- com2kid 4y agoThere are a variety of reading styles, some involve recognizing whole words at a time, others involve sounding out words. I'm also a whole word reader, it is IMHO generally faster than needing to mentally hear words.
- eternalban 4y agoI just tried that. Not sure I like the feeling, seems to rob of the pleasure of reading. But it -is- an interesting effect. Do you get pleasure from it?
- com2kid 4y agoIt is how I learned to read, early on I could read lots of words I had no idea how to pronounce. Makes reading sci-fi with funky alien names much easier! :D
- simonh 4y agoIt’s dominant but not universal. My daughters and I do this but my wife doesn’t. She can’t even imagine what it’s like.
- stared 4y agoTokenization for models like GPT or BERT can be seen as compression. That is, frequent words are separate tokens. Frequent sequences are separate tokens. On the other hand, if a sequence is very uncommon, then it will contain many tokens. Sure, you encode bit-by-bit. But it is a fixed-length code, which is even worse than character-by-character. Maybe you only get worse training and inference time. But I wouldn't be surprised if the encoding also serves as a Bayesian prior, and with a different encoding, you get worse results (for given data).
- tbalsam 4y agoThere's a couple of massive intuition leaps here (around tokens and the ease of which predicting one modality extends to another), but if you're interested in diving into the field at the place where they're asking questions like this, you could start by looking at the transition from BPE to the tokenizer we have today for the tokenization front, and PercieverIO for the multimodal generalization front.
- tehsauce 4y agoThe closest well-known example of something like this is deepmind's "perceiver" architecture. https://www.deepmind.com/blog/building-architectures-that-can-handle-the-worlds-data https://www.deepmind.com/blog/building-architectures-that-ca...
- hooande 4y agoIt's important to remember that the "power" of gpt doesn't come from the model, but from the sheer scale of the dataset. It's trained on the entire internet, in text form. You can 100% use a transformer architecture to train on binary data. But what data do you have hundreds of tebibytes of? Language also follows very common and repeatable patterns. "Hello" is often followed by "How are you?", etc. Just like Zipf's Law dictates that some words are used exponentially more than others, there are linguistic and conceptual patterns that appear with predictable frequency. If your bits don't follow similar rules, the results might not be as clean. I'm pretty sure you could code a transformer to work on binary or video data. Sounds like a great github project. But it's unlikely you'll have the scale of data to do anything close to ChatGPT.
- Terretta 4y ago> "power" [comes from] the entire Internet Are you sure? See "the Pile": https://pile.eleuther.ai/ https://pile.eleuther.ai/ The paper cited there by contrast, argues for select training sets: Recent work has demonstrated that increased training dataset diversity improves general cross-domain knowledge and downstream generalization capability for large-scale language models. With this in mind, we present \textit{the Pile}: an 825 GiB English text corpus targeted at training large-scale language models. The Pile is constructed from 22 diverse high-quality subsets -- both existing and newly constructed -- many of which derive from academic or professional sources. Our evaluation of the untuned performance of GPT-2 and GPT-3 on the Pile shows that these models struggle on many of its components, such as academic writing. Conversely, models trained on the Pile improve significantly over both Raw CC and CC-100 on all components of the Pile, while improving performance on downstream evaluations. https://arxiv.org/abs/2101.00027 https://arxiv.org/abs/2101.00027