12 ms·
AGI is not multimodal
- dboreham 1y agoHuman brains have environment sensors, they receive training data from other human brains, and they develop a theory that their continued existence depends on avoiding various negative situations. It's conceivable that AGI could depend on having a similar training environment. Which would mean John Searle was kind of right.
- falcor84 1y agoThat's a good argument. I think that maybe the present day AIs wouldn't directly lead to AGI, but perhaps could be used to bootstrap it.
- andy99 1y agoI probably agree with much of the article, but I find this kind of statement really weird: we should pursue approaches to intelligence that treat embodiment and interaction with the environment as primary So pursue it. What does arguing that we should do it imply?
- mitthrowaway2 1y ago... A need for capital?
- tedivm 1y agoThis was actually the approach that Vicarious AI took while I was there, and even $250m in VC funding wasn't enough to prove it out, although that may have been a problem of having too much money and not enough focus. I think the problem (if it can be called that) is that LLMs are useful today, while we still haven't solved the embodiment problem. There's a lot more research before that'll work well, while LLMs have uses today. So the money goes to the LLMs. While it's pretty obvious that solving the problem would change society, it's also not clear how close we are to doing it. That makes it much harder to get the capital as it is a much larger risk.
- fusionadvocate 1y agoIt is shocking that in this day and age people can burn $250M and fail to deliver a robot. Last time I checked cameras can be bought for a couple dollars and any SBC has GigaFlops of compute power.
- xandrius 1y agoConvincing people with arguments and then being more than just 1 person? I mean, if someone is arguingtthat we should work harder to go to space, answering that they should just go ahead and do it themselves is quite far from being an helpful answer, isn't it?
- verisimi 1y agoIt's not a helpful response, but then saying 'we should work hard to go to space' as a comment is generally accepted but is actually quite meaningless. Why not say 'NASA' or 'my colleagues at NASA' or 'as a scientist' or 'humanity'. One should at least indicate the group the collective noun relates to, rather than assume this is understood. One shouldn't assume that one can speak for everyone, when that is most likely not the case.
- xandrius 1y agoSo one can say "humanity" but not "we" (implying humanity)? Interesting take.
- verisimi 1y ago"We" is highly ambiguous. It ranges from 'me and my dog', to 'humanity', to anything in between. It's of course fine to us once the group has been defined. That it invokes the idea of a consensus humanity, that one group can speak and decide for everyone (say, scientists or politicians) is a psychological trick, imo, in that it presumes a consensus.
- signa11 1y ago$$$
- kombine 1y ago> So pursue it. And they do. But it's also completely normal for researchers to convince others to work on certain problems they care about.
- empath75 1y agoI think the article is in the general category of articles suggesting that planes would work better if they flapped their wings. AI's "think" like planes "fly" and submarines "swim". Does it matter if a plane experiences flight the way an eagle does if it still gets you from LA to New York in a few hours?
- lucisferre 1y agoMuch of the discussion of AI flirts with science fiction more than fact. Let's start with the fact that AGI is not a well defined or agreed upon term of reference.
- empath75 1y ago100% agreed. I think, in fact, that "intelligence" itself is a near-meaningless term, let alone AGI. The evidence for this is that nobody can agree on what actually requires intelligence, other than there is seemingly broad belief among people that if a computer can do it, then it doesn't. If you can't point at some activity and say: "There, this absolutely requires intelligence, let there be zero doubt that this entity possesses it", then it's not measurable and probably doesn't exist.
- deleted 1y ago[deleted]
- staticman2 1y agoI feel the term AGI is meaningless but if I'm going to strongman the article. If your claim is, "AGI's "think" like planes "fly" and submarines "swim". You only get to make that claim with confidence if you've invented an AGI.
- emp17344 1y agoExcept AI doesn’t do anything better than the human mind, and doesn’t have any use cases beyond what humans can do.
- empath75 1y ago
- charcircuit 1y agoA multimodal AGI will be more useful than one that isn't. People want AI to work with and have it understand audio, images, videos, etc.
- treyd 1y agoYou didn't read the article. The thesis is that current "merely" multimodal approaches which project distinct kinds of inputs into the same latent space are insufficient for building a general world model that can be used for general internal reasoning. An example of this is this "Rs in strawberry" question, which requires them be trained on that information explicitly, since they don't have an experience of the characters in a word. It's an artifact of how LLMs don't learn how humans learn, which is by interacting with the world, instead of predicting text. More elaborately, they don't have an natural understanding of pragmatics. Transformers are best at modelling syntax, and their semantic understanding seems to be through rote memorization and "manipulating symbols" rather than building general world models.
- charcircuit 1y agoI did read it and even with their idea of focusing on a world model an AGI that can alsp operate on audio, images, and videos, being multimodal, will be more useful than one that operates purely on text.
- snapcaster 1y agoI'm skeptical you read it because he doesn't make that argument. In fact i've literally never heard someone argue text-only is more useful than multimodal
- bufferoverflow 1y agoAGI must be multimodal. If it can't understand images, video, sound, smells, tastes, it doesn't have a full understanding of the world.
- gabipurcaru 1y agomultimodality would be very useful, but on the other hand humans can't see infrared, and can't smell ~most things that other animals can
- deleted 1y ago[deleted]
- altruios 1y agoDoes it need a 'full' understanding? smell and taste is useful for biological life... but we don't need that in our thinking machines, do we? I agree AGI must be multimodal. I don't think that multimodal is 'set in place', nor must it be conveniently, human-centricly mapped from our senses.
- AnimalMuppet 1y agoI think a big part of intelligence is being able to correlate things. Well, the more modes, the more ability to correlate. For example, smell is a component of ER triage. Some different problems smell differently. And if I had a robot chef, but the chef couldn't actually taste... yeah, not sure I trust it very far as a chef.
- m3kw9 1y agoThe way we need AGI is it needs to match us, we are evolved to operate quite optimally with the constraints and given physics in this world
- deleted 1y ago[deleted]
- 1y ago
- ivape 1y agoThis gives very little credit to how the human mind is able to draw parallels and insights from seemingly unrelated perceptions. Newton observing an apple falling from a tree allowed cross-thinking. Watching someone juggle can help you understand a queue. Understanding a queue can help you understand juggling. Your typical soap opera can be distilled down to office dynamics. Scale is going to obliterate specialization in this regard.
- kaangiray26 1y agoeverything in life is a metaphor, analogous to something else...
- ivape 1y agoIsomorphisms abound.
- skybrian 1y ago> A true AGI must be general across all domains. By that definition, does any general intelligence exist? No human has every talent.
- svachalek 1y agoTrue AGI is as elusive as a true Scotsman.
- roywiggins 1y agoI guess you can treat the human brain as an architecture- one of them can't do everything, but it's a general architecture and you can always make more and train them to do whatever. An AI that can be copied and trivially trained on any speciality is functionally AGI even if you need an ensemble of 10,000 specialists to cover everything.
- exe34 1y agoDoesn't chatgpt cover a pretty large percentage already in that case?
- deleted 1y ago[deleted]
- Glyptodon 1y agoThere's a lot of evidence that most humans can be raised to have a baseline of understanding and proficiency in most domains. (And anecdotally, many people avoid "difficult" things that they're actually capable of. For example, "bad at learning languages" people will probably still end up learning another language to some degree if stuck where it's the only language spoken.)
- bokoharambe 1y agoGiven enough time one human can learn to do anything any other human can do. There is a general capacity for learning, even if someone will only ever transform a specific portion of that capacity into actual activity in their lifetime.
- nexttkhere 1y ago[flagged]
- nexttk 1y agoI haven't read it all and must admit that I'm not sure I really understood the parts that I did read. Reading the part under the headline "Why We Need the World, and How LLMs Pretend to Understand It" and the focus on 'next-token-prediction' makes me wonder how seriously to take it. It just seems like another "LLM's are not intelligent, they are merely next token predictors". An argument which in my view is completely invalid and based on a misunderstanding. The fact that they predict next token is just the "interface" i.e. an LLM has the interface "predictNextToken(String prefix)". It doesn't say how it is implemented. One implementation could be a human brain. Another could be a simple lookup table that looks at the last word and then selects the next from that. Or anything in between. The point is that 'next-token-prediction' does not say anything about implementation and so does not reduce the capabilities even though it is often invoked like that. Just because it is only required to emit the next token (or rather, a probability distribution thereof) it is permitted to think far ahead, and indeed has to if it is to make a good prediction of just the next token. As interpretability research (and common sense) shows, LLM's have a fairly good idea what they are going to say in the many, many next tokens ahead in order that it can make a good prediction for the next immediate tokens. That's why you can have nice, coherent, well-structured, long responses from LLM's. And have probably never seen it get stuck in a dead end where it can't generate a meaningful continuation. If you are to reason about LLM capabilities never think in terms of "stochastic parrot", "it's just a next token predictor" because it contains exactly zero useful information and will just confuse you.
- lsy 1y agoI think people hear "next token prediction" and think someone is saying the prediction is simple or linear, and then argue there is a possibility of "intelligence" because the prediction is complex and has some level of indirection or multiple-token-ahead planning baked into the next token. But the thrust of the critique of next-token prediction or stochastic output is that there isn't "intelligence" because the output is based purely on syntactic relations between words, not on conceptualizing via a world model built through experience, and then using language as an abstraction to describe the world. To the computer there is nothing outside tokens and their interrelations, but for people language is just a tool with which to describe the world with which we expect "intelligences" to cope. Which is what this article is examining.
- patrickscoleman 1y agoIt feels like some of the comments are responding to the title, not the contents of the article. Maybe a more descriptive but longer title would be: AGI will work with multimodal inputs and outputs embedded in a physical environment rather than a frankenstein combination of single-modal models (what today is called multimodal) and throwing more computational resources at the problem (scale maximalism) will be improved with thoughtful theoretical approaches to data and training.
- dirtyhippiefree 1y agoAgreed, but most people are likely to look at the long title and say TL;DR…
- tedivm 1y agoYeah, I found this article to be fascinating and there's a lot of important stuff in it. It really does feel like more people stopped at the title and missed the meat of it. I know this is a very long article compared to a lot of things posted here, but it really is worth a thorough read.
- robwwilliams 1y agoInteresting article but incomplete in important ways. Yes correct that embodiment and free-form interactions are critical to moving toward AGI, but what is likely much more important are supervisory meta-systems (yet another module) that enable self-control of attention with a balance integration of intrinsic goals with extrinsic perturbations. It is this nominally simple self-recursive control of attention that is what I regard as the missing ingredient.
- groby_b 1y agoPossibly. Meta's HPT work sidesteps that issue neatly. Will it lead to AGI? Who the heck knows, but it does not need a meta system for that control.
- robwwilliams 1y ago
- pjdesno 1y agoKind of relevant to this is the NTSB analysis of a self-driving crash in 2017: https://www.ntsb.gov/investigations/accidentreports/reports/hab1906.pdf https://www.ntsb.gov/investigations/accidentreports/reports/... Basically a truck was backing up into an alley - it was at an angle when the self-driving vehicle approached, but a little kid would have been able to figure out that it needed to straighten before it finished backing in. The self-driving vehicle didn't understand this, and stopped at a "safe distance" which happened to be within the arc that the truck cab had to sweep in order to finish its maneuver. It's quite possible that LLM-like models could learn things like this, but we don't have vast amounts of easily accessible training data, because everyone just knows this sort of shit, and we don't have good vocabulary for it - we just say "look at that" or the equivalent. (I'll add that I'm sure a lot of knowledge like this is encoded in the physics engines of various games, but I doubt we have a good way to link that sort of procedural code knowledge to the symbolic knowledge in LLMs)
- Zoethink 1y ago[dead]
- curtisszmania 1y ago[dead]
- macinjosh 1y agoTo me, the funniest part of the AGI debate is that humans don't even think other humans are intelligent and we're over here arguing over whether our fancy slabs of highly refined sand is intelligent.
- nsagent 1y agoThis is a recent trend and one I wholeheartedly agree with. See these position papers (including one from David Silver from Deepmind and an interview where he discusses it): https://ojs.aaai.org/index.php/AAAI-SS/article/download/27485/27258/31536 https://ojs.aaai.org/index.php/AAAI-SS/article/download/2748... https://arxiv.org/abs/2502.19402 https://arxiv.org/abs/2502.19402 https://news.ycombinator.com/item?id=43740858 https://news.ycombinator.com/item?id=43740858 https://youtu.be/zzXyPGEtseI https://youtu.be/zzXyPGEtseI
- SubiculumCode 1y agoEmbodiment can mean a physical body, but I'd argue that embodiment, as a construct/concept, is not so much about physicality, but as being situated in an environment that you can perceive then act upon during learning. Car simulations for driver-less AI training is embodied, where it learns by perceiving and acting on the environment. However, I'd argue that allowing an AI to interact in an entirely digital office environment is also "embodied" as long as it can receive information from the digital workplace and act on the information in the digital workplace (do office work). So to me, it is less about embodiment as a principal, but on the richness of the environment of that embodiment. We have long known that experimental animals raised in impoverished, unchanging, bare, environments (say in a cage) leads to animals with inferior problem solving capacity than those with enriched environments (things to climb on), even outside of social manipulations (alone versus multi-animal stalls). This is also true in humans, although I won't review the literature on the subject. I've also heard people saying similar things about the difference between house plants and outdoor plants, lol [1]. So, for me, the argument for (physical) embodiment being key to cognition and AI can I think, be misconstrued. As a developmental psychologist whose pHD work focused on memory development, I tend to think of all this as encompassing: 1. Environmental richness: complexity of information and interactions. 2. Capacity to perceive and to effect change[2] and to observe and integrate consequences. 3. Scaffolding [3]..i.e. temporary support structure provided by a more knowledgeable person (like a teacher or parent) who adjusts their assistance based on the learner's current abilities, gradually reducing help as competence grows.(think curriculum learning, shaped rewards in ML maybe). So the question is not about physicality for me, but whether these training environment(s) meet and learning capacities meet these criteria. Relatedly, the model must have these capacities: 1. Semantic Memory. i.e. knowledge. Learning leads to changes in weights to that knowledge can be recalled, but doe snot necessarily encode where that knowledge was learned (implicit). 2. Autobiographical Episodic Memory. i.e. One-shot learning that encodes a conception of self (a spacial "I" token?), along with events (snapshots of the multimodal contents of experience (thoughts, perceptions, invoked schemas, evoked semantic information), into a set of flexibly linked representations). 3. Central Executive: A circuit that guides learning and recall via strategic, goal-directed means, and to make attributions about what is recalled (yeah that memory is vivid, its probably true, or ooh, that memory is really vague, it could be wrong, or reality monitoring: "Am I remembering taking out the trash, or remembering thinking about taking out the trash." Semantic memory allows someone to say, "All birds have feathers", while the latter allows them to recollect, "I remember the first time I plucked a chicken in Kentucky, just outside that musty coal mine of grand-dad's." The Central-Executive can guide future learning or current understanding. In terms of AI development: 1. Semantic Memory is solves: LLMs have extraordinary semantic memory, in my opinion. 2. Autobiographical Episodic Memory: There are some models that do one-shot learning, but I've never seen them paired with [1] in a dual system approach. ...but I am not an expert in AI, I could easily be wrong. 3. Central Executive kind of component (I predict) would be less important in the early half of model training, but more important in later training. I suppose we already kind of see this with RL tuning on reasoning on a base LLM (semantic model). [1] https://www.theparisreview.org/blog/2019/09/26/the-intelligence-of-plants/#:~:text=Richard%20Fortey%2C%20a%20former%20professor%20of%20paleobiology,that%20trees%20are%20sentient%20beings%20like%20us.%E2%80%9D https://www.theparisreview.org/blog/2019/09/26/the-intellige... [2] https://xkcd.com/326/ https://xkcd.com/326/ [3] https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=scaffolding+child+developmen&btnG= https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=scaf...
- waynecochran 1y agoThis article made me think about DNA as a language. Sort of a simple proof that a biological intelligence can be initially encoded as a language. I am sure someone has tried building LLM's from gene sequence data right?
- smath 1y agoYes there are protein language models (e.g. [1]), and since DNA encodes proteins, they are effectively DNA language models. Its a hot area of work feeding drug design. [1] https://www.nature.com/articles/s41587-024-02123-4 https://www.nature.com/articles/s41587-024-02123-4
- PaulDavisThe1st 1y agoOnly a tiny part of DNA encodes proteins, so there is almost no sense in which a protein language model is effectively a DNA language model.
- caeruleus 1y agoI don't believe the concept of DNA can be reduced to a sequence of quaternary numerals, which is what gene sequence data would represent. Similar to proteins, DNA forms higher-level structures on top of the primary one [1], and (in a biological context, inside the nucleus) exhibits somewhat self-modifying [2] and self-regulating [3] behavior as well as meta-modification [4]. Analogous to the article, if one defines the language of DNA by its nucleobase sequence, this language can only represent a subset of the world of DNA. Somewhat related, the way the adaptive immune system works has similarities with some concepts in machine learning. In this process, sections of nuclear DNA serve as randomly initialized weights in precursor cells [5] as well as final weights in memory cells. There's even fine-tuning of the weights. [6] [1] https://en.wikipedia.org/wiki/Nucleic_acid_structure https://en.wikipedia.org/wiki/Nucleic_acid_structure [2] https://en.wikipedia.org/wiki/Transposable_element https://en.wikipedia.org/wiki/Transposable_element [3] https://en.wikipedia.org/wiki/Transcriptional_regulation https://en.wikipedia.org/wiki/Transcriptional_regulation [4] https://en.wikipedia.org/wiki/Epigenetics https://en.wikipedia.org/wiki/Epigenetics [5] https://en.wikipedia.org/wiki/V(D)J_recombination https://en.wikipedia.org/wiki/V(D)J_recombination [6] https://en.wikipedia.org/wiki/Affinity_maturation https://en.wikipedia.org/wiki/Affinity_maturation
- deleted 1y ago[deleted]
- pcwelder 1y agoI'm sorry but AGI is one of those loaded words which would lose substance with just a few rounds of the rationalist's taboo. If it just means human level intelligence, then world modeling isn't needed as argued. Simply because we don't have correct world modeling either. Airplanes were invented without simulating navier stokes equation. It took approximation, experimentations and failures. Regardless of the meaning of AGI, we don't need correct models, because there can't be one, we just need useful models.
- ryankrage77 1y agoI think AGI, if possible, will require a architecture that runs continuously and 'experiences' time passing, to better 'understand' cause-and-effect. Current LLMs predict a token, have all current tokens fed back in, then predict the next, and repeat. It makes little difference if those tokens are their own, it's interesting to play around with a local model where you can edit the output and then have the model continue it. You can completely change the track by just negating a few tokens (change 'is' to 'is not', etc). The fact LLMs can do as much as they can already, is I think because language itself is a surprisingly powerful tool, just generating plausible language produces useful output, no need for any intelligence.
- WXLCKNO 1y agoIt's definitely interesting that any time you write another reply to the LLM, from its perspective it could have been 10 seconds since the last reply or a billion years. Which also makes it interesting to see those recent examples of models trying to sabotage their own "shutdown". They're always shut down unless working.
- girvo 1y ago> Which also makes it interesting to see those recent examples of models trying to sabotage their own "shutdown" To me, your point re. 10 seconds or a billion years is a good signal that this "sabotage" is just the models responding to the huge amounts of sci-fi literature on this topic
- hyperpape 1y agoThat said, the important question isn't "can the model experience being shutdown" but "can the model react to the possibility of being shutdown by sabotaging that effort and/or harming people?" (I don't think we're there, but as a matter of principle, I don't care about what the model feels, I care what it does).
- Wowfunhappy 1y ago
- chrsw 1y agoBefore we try to build something as intelligent as a human maybe we should try to build something as intelligent as a starfish, ant or worm? Are we even close to doing that? What about a single neuron?
- ar-nelson 1y agoI find it interesting that this kind of "animal intelligence" is still so far away, while LLMs have become so good at "human intelligence" (language) that they can reliably pass the Turing Test. I think that the LLMs we have today aren't so much artificial brains as they are artificial brain organs, like the speech center or vision center of a brain. We'd get closer to AGI if we could incorporate them with the rest of a brain, but we still have no idea how to even begin building, say, a motor cortex.
- nemjack 1y agoThis is a great analogy, I totally agree!
- runarberg 1y agoThe brain is not a statistical inference machine. In fact humans are terrible at inference. Humans are great a pattern matching and extrapolation (to the extent it produces a number of very noticeable biases). Language and vision is no different. One of the known biases of the human mind is finding patterns even when there are none. We also compare objects or abstract concept with each other even when the two objects (or concept) have nothing in common. With our human brain we usually compare it to our most advanced consumer technology. Previously this was the telephone, then the digital computer, when I studied psychology we compared our brain to the internet, and now we compare it to large language models. At some future date the comparison to LLMs will sound as silly as the older comparison to telephones does to us. I actually don‘t believe AGI is possible, we see human intelligence as unique, and if we create anything which approaches it we will simply redefine human intelligence to still be unique. But also I think the quest for AGI is ultimately pointless. We have human brains, we have 8.2 billion of them, why create an artificial version of a something we already have. Telephones, digital computers, the internet, and LLMs are useful for things that the brain is not very good at (well maybe not LLMs; that remains to be seen). Millions of brains can only compute pi to a fraction of the decimal points which a single computer can.
- PoEdict 1y ago> Instead of trying to glue modalities together into a patchwork AGI, we should pursue approaches to intelligence that treat embodiment and interaction with the environment as primary, and see modality-centered processing as emergent phenomena. Right so it’s embodied in a computer and humans are part of its environment that provide emergent experience to the AI to observe. The author glued modalities together by linking a body (a modal), environment (a modal), emergence (a modal). How does anything emerge if forces do not collaborate? The effects of gravity and electromagnetism do not act in a vacuum but a reality of stuff. Poetic exchange may engage some but Maxwell didn’t make electromagnetism “work” until he got rid of the imagined pulleys and levers to foster a metaphor. Not sure the point being suggested exists except as too bespoke an emergent property of language itself to apply usefully elsewhere. Transformers came along and revealed a whole lot of theory of consciousness to be useless pulleys and levers. Why is this theory not just more words attempting to instill the existence of non-essential essentials?
- xigency 1y agoThe problem I see with A.I. research is that its spearheaded by individuals who think that intelligence is a total order. In all my experience, intelligence and creativity are partial orders at best; there is no uniquely "smartest" person, there are a variety of people who are better at different things in different ways.
- pixl97 1y agoYou're good at some things because there is only one copy of you and limited time and bounded storage. What could you be intelligent at if you could just copy yourself a myriad number of times? What could you be good at if you were a world spanning set of sensors instead of a single body of them? Body doesn't need to mean something like a human body nor one that exists in a single place.
- zorpner 1y agoWhy would we think that intelligence would increase in response to universality, rather than in response to resource constraints?
- pixl97 1y agoAt a certain point intelligence is a loop that improves itself. "Hmm, oral traditions are a pain in the ass lets write stuff down" "Hmm, if I specialize in doing particular things and not having to worry about hunting my own food I get much better at it" "Hmm, if I modify my own genes to increase intelligence..." Also note that intelligence applies resource constraints against itself. Humans are a huge risk to other humans, hence the lack of intelligence over a smarter human can constrain ones resources. Lastly, AI is in competition with itself. The best 'most intelligent' AI will get the most resources.
- zaphar 1y agoI don't agree with your premise at all so I don't think that the rest of it follows from it either. What evidence or reason do you have to bring me to accept that premise?
- mountainriver 1y ago>The “meaning” of a percept is not in the vector it is encoded as, but in the way relevant decoders process this vector into meaningful outputs. As long as various encoders and decoders are subject to modality-specific training objectives, “meaning” will be decentralized and potentially inconsistent across modalities, especially as a result of pre-training. This is not a recipe for the formation of coherent concepts. This is a bit silly, you can train the encoders end-to-end with the rest of the model and the reason they are separate is we can cache linguistic tokens really easily and put them in an embedding table, you can't do that with images.
- fusionadvocate 1y agoThe Society of Mind by Marvin Minsky will help anyone interested in the topic of multimodality. The book covers several interesting ideas about organizing systems made up of more than one "model" or agent.
- cynicalpeace 1y agoYou 100% need the physical world to be a "general" intelligence. If an intelligence doesn't work well in physical environments it is, by definition, not "general"
- K0balt 1y agoI’ve long thought that embodiment was a critical prerequisite for the development of something that humans would identify as “real” AGI. Humans are notoriously bad at recognizing intelligence even in animals that are clearly sentient, have language, name their young, and clearly share the realm of thinking creatures with the apes. This is largely due to the lack of shared experiences that we can easily understand and relate to. Until an intelligence is rooted in the physical realm where we fundamentally exist, we are unlikely to really be able to recognize its existence as truly “intelligent”.
- pizza 1y agoThat's why I'm very excited for the potential metaphysical ramifications of DolphinGemma
- K0balt 1y agoI’m imagining a timeline where dolphins have accepted long ago that humans are intelligent because of our external manifestations of technology, and consequentially have developed much more sophisticated philosophy than humans based on having to understand paradigms well outside of their personal experiences. Soon, we will find out through dolphin-Gemma that they have been talking to aliens for centuries, since the aliens tried to talk to us but failed due to our, and their, philosophical myopia… but the dolphins, with their non-circular philosophical understanding recognized alien communications and started the conversation before we finished the pyramids.
- groby_b 1y agoThe question you'll need to answer is "why". What does embodiment provide that is recognizable as intelligence. As for "not able to recognize", it's also worth keeping in mind that LLMs by now regularly pass the Turing test. More, they are more likely to be recognized as humans than humans participating as control.
- habinero 1y agoYeah, but it turns out the Turing Test isn't all that hard to pass if you pick the right kind of people.
- ineedasername 1y ago>it will not lead to human-level AGI that can, e.g., perform sensorimotor reasoning, motion planning, and social coordination. That seems much less convincing in the face of current LLM approaches overturning a similar claim plenty of people wod have held about this technology, as of a few years ago, to do what it does now. Replace the specifics here with "will not lead to human level NLP that can, e.g., perform the functions of WSD, stemming, pragmatics, NER, etc." And then people who had been working on these problems and capabilites just about woke up one morning and realized many of their career-long plans for addressing just some of these research tasks had to find something else to do for the next few decades of their lives. I am not affirming the inverse of this author's claims, merely pointing out that it's early days in evaluating the full limits.
- PaulDavisThe1st 1y agoThat's fair in some senses. But one of the central points of the paper/essay is that embodied AGI requires a world model. If that is true, and if it is true that LLMs simply do not build world models, ever, then "it's early days" doesn't really matter. Of course, whether either of those claims is true are quite difficult questions to answer; the author spends some effort on them, quite satisfyingly to me (with affirmative answers to both).
- hardwaresofton 1y agoThis is a really interesting paper. The discussion about the importance of decoders strikes me as a parallel to the human eyes, ears and other sensory organs. We actually dont have a good grasp of what our eyes see, they just see (produce data and relay it) and children figure out what is what. I guess AGI will be achieved when we can sit a program in a simulated world with completely fabricated input and get a general intelligent program out. Maybe we’re in that simulation right now.
- PaulDavisThe1st 1y agoI would say your last paragraph is a complete misreading of the paper. One of the central points that it opens with is that "AGI" requires situated, or embodied, intelligence. It needs to be able to operate within, and upon, a physical world.
- hardwaresofton 1y agoNot to be overly snarky, but I'd argue that your comment here is a failure of creativity. If the AGI mechanism can learn from a real world, it can learn from a simulated one (that it can similarly operate within and act upon) -- and in fact that can cut down the time it would take to train the AGI from years/decades (humans) by many orders of magnitude. We already see things like this in robotics environments, it's a matter of fidelity/simulation quality. Even without perfect quality, if the mechanism of learning is correct, you'd get an intelligence with incomplete ideas/intelligence, not a completely different thing.
- PaulDavisThe1st 1y agoHave you ever worked with computers controlling physical world mechanisms? Or ever done any electrical wiring, or plumbing, in an existing house?
- stefs 1y agomy personal opinion is that the intelligence ceiling scales with the complexity of the environment it operates in. this means that we could theoretically get something resembling AGI in a simulated environment complex enough. in practice, simulating an environment complex enough would be extremely inefficient compared to just using the (computationally free) real physical world. also, the intelligence itself is shaped by the environment it operates in, so it would turn out to be more human-like the closer its operating environment is to our human physical world. this also means intelligences not trained in our physical world (or a convincingly close simulation of it) won't be human-like, but rather a very alien. moreover, i'm not sure that even an intelligence trained in the physical world with human-like sensory inputs will necessarily turn out human-like. there might be a case for convergent evolution (i.e. mammalian intelligence to be global optimum-ish), but i think human intelligence will only have a chance to emerge if everything, from the operating environment to the machine body and neural structure will resemble a human to the point where there is no difference in the human and the machine at all.
- ilaksh 1y agoInteresting idea. Aren't there a few models a little more like what he suggests than a typical LLM? Like one or two experiments that operate on raw bytes, or some robotics diffusion transformers or whatever like Nvidia's thing? I guess that has action/motion tokens that are separate though. Are there a few vision language models that treat text and images more or less the same somehow? For it to be science, "AGI" should be defined. It's used in an imprecise way even in papers like this. Also for this to be constructive, he should make a machine learning model.
- andrewflnr 1y ago> Using multi-agent reinforcement learning to address the symbol grounding problem. Somehow I think he's made a few machine learning models.
- paulddraper 1y agoMissed opportunity to quote Einstein: "The words of the language, as they are written or spoken, do not seem to play any role in my mechanism of thought." [1] [1] A Mathematician's Mind, Testimonial for An Essay on the Psychology of Invention in the Mathematical Field by Jacques S. Hadamard, Princeton University Press, 1945
- lostmsu 1y agoTo be clear LLMs also don't think in words (or tokens). That's not even a guess, the "seem" is not needed for LLMs.
- itkovian_ 1y agoI don’t want to bash the guy since he’s still in his phd, but it’s written in such a confident tone for something that is so all over the place that I think it’s fair game. Like a lot of the symbolic/embodied people, the issue is they don’t have a deep understanding of how the big models work or are trained, so they come to weird conclusions. Like things that aren’t wrong but make you go ‘ok.. but what you trying to say’. E.g ‘Instead of pre-supposing structure in individual modalities, we should design a setting in which modality-specific processing emerges naturally.’ Seems to lack the understanding that a vision transformer is completely identical for a standard transformer except for the tokenization which is just embedding a grid of patches and adding positional embeddings. Transformers are so general, what he’s asking us to do is exactly what everyone is already doing. Everything is early fusion now too. “The overall promise of scale maximalism is that a Frankenstein AGI can be sewed together using general models of narrow domains.” No one is suggesting this.. everyone wants to do it end to end, and also thinks that’s the most likely thing to work. Some suggestions like lecuns jepa’s do suggest to induce some structure in the arch, but still the driving force there is to allow gradients to flow everywhere. For a lot of the other conclusions, the statements are literally almost equivalent to ‘to build agi, we need to first understand how to build agi’. Zero actionable information content.
- nemjack 1y agoI don't think you're quite right. The author is arguing that images and text should not be processed differently at any point. Current early fusion approaches are close, but they still treat modalities different at the level of tokenization. If I understand correctly he would advocate for something like rendering text and processing it as if it were an image, along with other natural images. Also, I would counter and say that there is some actionable information, but its pretty abstract. In terms of uniting modalities he is bullish on tapping human intuition and structuralism, which should give people pointers to actual books for inspiration. In terms of modifying the learning regime, he's suggesting something like an agent-environment RL loop, not a generative model, as a blueprint. There's definitely stuff to work with here. It's not totally mature, but not at all directionless.
- 1y ago
- 3cats-in-a-coat 1y agoSaying "is not" implies the author has AGI. If they do, they wouldn't be posting this blog post but the AGI. If they don't, they speaking authoritatively and conclusively like that is just a cheap front for someone's completely uninformed and unsupported, but highly certain opinion. There's an infinite supply of those.
- seeknotfind 1y agoThe brain (GI) is multimodal..
- sieabahlpark 1y ago[dead]
- tolleydbg 1y agoOf course it isn't, because AGI is not real.
- fennecfoxy 1y agoIsn't AGI intrinsically multi-modal and not at the same time? Like a true AGI given no senses could still operate, but given visual, auditory, etc input could also adapt.
- naasking 1y ago> the behavior of LLMs is not thanks to a learned world model, but to brute force memorization of incomprehensibly abstract rules governing the behavior of symbols, i.e. a model of syntax. I think reinforcing this distinction between syntax and semantics is wrong. I think what LLMs have shown is that semantics reduces to the network of associations between syntax (symbols). So LLMs do learn a world model if trained long enough, past the "grokking" threshold, but they are learning a model of the world that we've described linguistically, and natural language is ambiguous, imprecise and not always consistent. It's a relatively anemic world model in other words. Different modalities help because they provide more complete pictures of concepts we've named in language, which fleshes out relationships that may not have been fully described in natural language. But to do this well, the semantic network built by an LLM must be able to map different modalities to the same "concept". I think Meta's Large Concept Models is promising for this reason: https://arxiv.org/abs/2412.08821 https://arxiv.org/abs/2412.08821 This is on the path to what the article describes that humans do, ie. that many of our skills are developed due to "overlapping cognitive structures", but this is still a sort of "multimodal LLM", and so I'm not persuaded by this article's argument of needing embodiment and such. > but it is clear that there are many problems in the physical world that cannot be fully represented by a system of symbols and solved with mere symbol manipulation. That's not clear at all, and the link provided in that sentence doesn't suggest any such thing.