4 ms·
Why is a constant stream of ultra high resolution input from ~5 senses for years not count as a huge volume of data? Not to mention all the pre-training data en
by nofum 3y ago
Why is a constant stream of ultra high resolution input from ~5 senses for years not count as a huge volume of data? Not to mention all the pre-training data encoded in the brains blueprint via DNA through natural selection? If anything, LLMs are more impressive than humans and animals in that they only have access to text to build a world model from a tabula rasa.
- chrisco255 3y agoWhile the input resolution is indeed high, it's not as if any animal on earth has access to photographic memory of all that data. It is so ephemeral you could scarcely tell me what the last 10 people you passed on the street were wearing, even if it was 10 minutes ago.
- jahewson 3y agoAt inference time, an LLM does not have access to all its training data either.
- chrisco255 3y agoAt inference time, a human being has access to 5 to 9 variables at most an average human can hold in their head at any given time. But I can feed pages and pages to an LLM and it has a perfect representation of every single word and letter in its working memory. It also has access to a look up table of billions of statistical correlations for what the next word or block of words in any sentence should be given the perfect priors. If I'm talking to you, especially in person, I won't be able to tell you what the second word of three sentences ago was. In fact, I might not have even been paying attention, because I'm already thinking at a higher level about what you're communicating to me and the word wasn't all that important. But a LLM certainly uses this information to guide its responses (along with its ginormous lookup table).
- augment002 3y ago> LLMs are more impressive than humans and animals in that they only have access to text to build a world model from a tabula rasa. They don’t build a world model at all. They make inferences from text. Considering they have been trained on 100,000 books, and all of Wikipedia, they are remarkably unintelligent, and really only able to produce text that is consistent with what they have been trained on.
- jiggawatts 3y agoPeople have pointed out that LLMs like GPT 3 and 4 can draw pictures, but they've been "blind since birth" and have never seen anything. They've just read descriptions of things. How capable would a human be, if they had grown up deaf, blind, and in all other ways insensate except for some sort of braille-like reading input? Another salient quote is from Andrej Karpathy (ex-Tesla AI team lead) who said that a camera is a high-bandwidth input "that puts many constraints on the world". Children learn from multi-modal inputs, and vision especially provides a large number of constraints that they can use to learn how the world works. I have a two-year old, and something I've noticed is that infants have a very strong instinctive urge to gain agency over the world. They try very hard to control things, to make things move, to make sounds, to be able to affect things in all sorts of ways. This must be a very critical part of learning, because they'll kick and scream if you remove their agency. It's as strong an instinct as wanting to eat or sleep. No current LLM has any kind of feedback loop, and has essentially zero agency.
- palad1n 3y agoThat might be a feature, not a bug, depending on what you mean by "agency". At the tippy top of that concept might be an AI system that decides by itself (as perhaps a weapons system) whether to kill human beings without direct instruction to do so. Yeah, Skynet. And every other movie where the AI "goes crazy".
- chrisco255 3y agoOn the contrary, an AI without agency could be instructed by a human to execute tasks that are evil and it will never push back either.
- chrisco255 3y agoWe have access to video training data for driving, absolutely tons of it, as we've been attempting to train AI cars for more than two decades now. If GPT4 is what you say it is, we should be able to train it on that video data and solve autonomous driving. There is nothing inherently about transformers that prevents them from taking in video data. They've already been used by some researchers (https://arxiv.org/abs/2104.09224 https://arxiv.org/abs/2104.09224). And yet, you can take a 16 year old who's never driven, and teach them within a week to be decent at it and maybe 50-100 hours of driving training and they're competent. You don't need to show them a billion man-hours of driving footage. Even the first people to buy cars in the late 1800s when they were first invented, were able to pick it up almost right away (there weren't even driving licenses back then). At any rate, driving is just one example. Despite being one of the oldest futuristic sci-fi examples, I don't see restaurants powered by AI. I don't see housekeeping powered by it. Ok, those are embodied examples, so you say they're unfair. Fine. What remote-friendly jobs are being swept away by AI? Can we even do customer service with AI right now? No, outside of some "front line" chat bot (which just replaces phone trees and terrible localized search engines), we can't. Even if a GPT is trained on a business's proprietary documentation, it's wrong or unresponsive enough that it would cost you more than it would save you by firing your customer support staff.
- zarzavat 3y agohttps://en.m.wikipedia.org/wiki/Kolmogorov_complexity https://en.m.wikipedia.org/wiki/Kolmogorov_complexity The raw volume of the data from the human senses is high but a lot of it is low entropy. The high resolution video from our retinas of reading the pages of a book over many hours, and the ASCII representation of said text are equivalent, modulo information about the physical object of the book.
- Robotbeat 3y agoI mean, Helen Keller did pretty well (with an appropriate tutor) with only very low bandwidth interfaces like touch and smell.