4 ms·
I am also not an AI researcher, but shouldn't there be some quantification of how efficient the learning is. I imagine a human has much more information than th
by _glass 3y ago
I am also not an AI researcher, but shouldn't there be some quantification of how efficient the learning is. I imagine a human has much more information than the pure audio, which would then level up the information density, i.e. the 500 million words are not an equivalent, but should be compared to all of the audio, visual, tactile, etc. information, plus the context, meaning the desire of the humans, behavior, response patterns.
- neurobama 3y agoFor sure. That's why I mentioned multi-modal training, where the language model is trained on text in tandem with visual and audio training data. GPT-4 was trained on both images and text and can even explain why visual jokes are funny: https://openai.com/research/gpt-4 https://openai.com/research/gpt-4 Training an LLM on video seems like the clear next step. I wonder if images and video might not encode visuospatial relations that are helpful to learning language that describes them.