4 ms·
> Language and vision are just the beginning — the parts we were able to digitize first - not necessarily the most central to intelligence. I respectfully disa
by dinfinity 1y ago
> Language and vision are just the beginning — the parts we were able to digitize first - not necessarily the most central to intelligence.
I respectfully disagree. Touch gives pretty cool skills, but language, video and audio are all that are needed for all online interactions. We use touch for typing and pointing, but that is only because we don't have a more efficient and effective interface.
Now I'm not saying that all other senses are uninteresting. Integrating touch, extensive proprioception, and olfaction is going to unlock a lot of 'real world' behavior, but your comment was specifically about intelligence.
Compare humans to apes and other animals and the thing that sets us apart is definitely not in the 'remaining' senses, but firmly in the realm of audio, video and language.
- voxleone 1y ago> Language and vision are just the beginning — the parts we were able to digitize first - not necessarily the most central to intelligence. I probably made a mistake when i asserted that -- should have thought it over. Vision is evolutionarily older and more “primitive”, while language is uniquely human [or maybe, more broadly, primate, cetacean, cephalopod, avian...] symbolic, and abstract — arguably a different order of cognition altogether. But i maintain that each and every sense is important as far as human cognition -- and its replication -- is concerned.
- wizzwizz4 1y agoPeople who lack one of those senses, or even two of them, tend to do just fine.
- oasisaimlessly 1y agoMostly thanks to other humans helping them. If all humans lacked vision, the human race would definitely not do just fine.
- actionfromafar 1y agoI think we need to think about vision and world modelling somewhat separately. We could construct an artificial (tech enhanced) society where sight was not available. People would still "model the world in their minds" with the "abstract model" part of the vision/world system.
- dinfinity 1y agoVision is interesting in that it leverages the maximum speed with which it is easily possible to gather information about our surroundings in this universe. I believe that is what makes it special and very valuable. I also believe this aspect makes it a strong attractor for convergent evolution. Language allows encoding and compression of information about the world, which is of course incredibly powerful and increases communication bandwidth enormously (as well as tons of other stuff). I'd say that for high level cognitive processes, hearing and speaking were an important stepping stone because for some reason evolving organs that can generate relatively high bandwidth signals in audio seems to be easier than evolving something that does that for visuals (very few Teletubby screens on tummies in the natural world). Interesting games to think about in this sense: Pictionary/drawing games and charades.
- actionfromafar 1y agoRegarding visual communication, I think you downsell posturing, gesturing and facial expressions a little. They may not be as high bandwidth as talking but they are very low latency and pretty stealthy if necessary.
- dinfinity 1y agoWell, there is sign language, so I guess you're right. It would be interesting to see how high bandwidth gesturing can be compared to speaking. I thought about this some more and I think the prevalence of making sounds rather than gesturing etc. is due to sound being a broadcasting mechanism that works over long distances and without line of sight. Visually indicating that you've claimed some territory is pretty hard.
- actionfromafar 1y agoI was thinking more about "low level" communication which can be gleaned from body language, frowns, smiles, gaze, winks, pointing etc. Perhaps not very information dense, but very fast.
- computably 1y agoLanguage is literally an abstraction of sensory inputs and cognitive processes. One can make similar arguments about image generation. These abstractions might characterize the higher cognitive abilities of humans, but it makes no sense to ignore "lower level" cognition. Embodiment is the foundation of our rich internal world models, in particular spacetime, causality, etc. Current generative models merely mimic the output, with a fuzzy abstract linguistic mess in place of any physical/causal models. It's unsurprising that their capacity to "reason" is so brittle.
- dinfinity 1y ago> Language is literally an abstraction of sensory inputs and cognitive processes. Language can exist entirely independently from senses and cognition. It is an encoding of patterns in the world where the only thing that matters is if anybody or anything wielding it can map the encodings to and from the patterns they encode for (which is more of a sociological/synchronisation challenge). Does C, or Java, 'make no sense' because it 'ignores lower level cognition'? There are many parts of non-programming languages that similarly have nothing to do with embodiment. Some of them are even about incredibly abstract things impossible in our universe. One could argue that for many fields genius lies in being able to mentally model what is so foreign to the intuition our embodiment has imbued us with or to be able to find a mapping to facilitate that intuition. Said otherwise: the experience our embodiment has given us might limit how well we can understand the world (Quantum Mechanics anyone?). Again, embodiment is interesting and worth pursuing, but far from a requirement for far-reaching intelligence.
- pjmorris 1y ago> Does C, or Java, 'make no sense' because it 'ignores lower level cognition'? It makes sense in context, but that context includes the machine on which the compiled code runs. Without the underlying machine, there's no real purpose for C or Java. I'm open to the idea that 'lower level cognition' may be as relevant to language as the machine is to C or Java.
- motorest 1y ago
- azeirah 1y agoHumans are known for their exceptional sensitivity in their hands and fingers. There are only few animals that come close to our ability to manipulate objects. Only octopuses, elephants and apes are in a similar league with regards to dexterity and finesse.
- cma 1y agoYou can be born without hands and have zero cognitive deficits. Sensory info and action-feedback from hands, vision, hearing, isn't key to intelligence, but if you are born without vision and hearing it can cause developmental issues, but even if you lose vision and hearing before 2 years you can develop normally, like Hellen Keller.
- dbspin 1y agoActually this is wrong. There's a connection between bodily sensation and emotion so profound that quadriplegics can develop flat affect which in turn leads to decision paralysis and cognitive deficit. Emotions are regulated somatically, and inform decision making and other aspects of motivation and reasoning. Source: https://pmc.ncbi.nlm.nih.gov/articles/PMC2633768/ https://pmc.ncbi.nlm.nih.gov/articles/PMC2633768/
- drw85 1y agoSpinal cord injury implies that you were born with those senses and then abruptly lost them. In that case, all the processes and pathways in your brain are relying on and tied to those senses, so losing them might also disrupt or affect those pathways and how they function.
- ordu 1y ago> Touch gives pretty cool skills, but language, video and audio are all that are needed for all online interactions. We use touch for typing and pointing, but that is only because we don't have a more efficient and effective interface. It may be, that we are not using touch for anything important as adults. But babies rely on touch to explore their surroundings. They stick anything into their mouth, why? Because a tongue is the most touch-sensitive organ. They are exploring things by touching them with their tongues. I can only guess what people get from that, but my guess is they get understanding of geometry and of surface properties of objects, which you'll have problems to get by processing photos or texts. > your comment was specifically about intelligence. Talking about intelligence, I do not believe that LLMs can match humans without deep understanding or 3d-space and material science^W intuition. It needs touch and temperature sensitivity at least. Probably you can replace it with billions of words of texts describing these things, but I doubt it.
- dinfinity 1y agoIt is trivial to train AI on 3D representations. In fact, that already happens in cases where robot algorithms are trained in simulations. Another thing to remember is that the senses we have aren't the only ones in biology and far from the only ones possible. In fact, anything that gives you another type of information about the world (you're modeling) is a different sense. In that sense (ha), AI has access to an incredibly vast and varied array of senses that is inaccessible to humans. Lidar is a very simple example of that. I don't think touch and temperature sensitivity are needed to achieve it, but I do agree that training with senses specifically for understanding 3D space is very important. At the very least binocular video.
- ordu 1y ago> It is trivial to train AI on 3D representations. So AI developers understand limitations and trying to remove them. It will help, but it will not make AI vision to be on par with a human's. > In that sense (ha), AI has access to an incredibly vast and varied array of senses that is inaccessible to humans. Lidar is a very simple example of that. I don't think that current uses of lidars have anything to do with intelligence. Not every neuro-net is about intelligence. > I don't think touch and temperature sensitivity are needed to achieve it, I'm sure they are. To understand forms you need to explore them with touch. The ability to understand forms by just looking at them is an acquired skill. Maybe it is possible to train these abilities without the touch, but how? I believe it will take a shitload of training data, and I'm not sure it will be good enough. Temperature sensitivity is a big thing, because it allow you to guess thermal conductivity of a thing by just looking at it. It allows to guess wetness of a thing. It allows us to guess temperature of things by looking at them: like you see sun shining, fire burning, people touching things and yanking their hands from hot things. Or just how about a person that cautiously trying to learn a temperature of a thing, at first measuring infrared radiation, then a quick touch, then a touch for a longer time, and finally a long sustained contact: how could you understand all these proceedings without your own experience of grasping the hot thing, crying from a pain and dropping the thing on your feet? These are just obvious ideas from top of my mind. What else comes from temperature sensitivity I don't know and no one is, because no one really knows how people learn to use their senses and to think. There are theories about it, but they are more of descriptive nature: they describe what is known without having a lot of a predictive power. Because of this the optimism of AI crowd seems overinflated. They don't know what they are trying to do, and still they believe in their eventual success. Probably you can learn it by thinking, but can LLMs think, while training? You can learn it as a pattern of a behaviour, without understanding the meaning of it, but then you'll hallucinate this pattern all the time, just because some of the movements were close enough. > At the very least binocular video. I'm not sure that people can learn 3d by looking. At least they do not just rely on a binocular vision to learn it. They touch, they lick. They measure things in different ways (by sticking it in mouth, by grasping, by climbing on top of it or falling from it, by hugging it), they measure distances by crawling or walking along them. They are finding a spot where they can see what happens behind a pack of tree, or maybe behind something else. People not just using more senses, they are acting also, which allows them to learn causal relationships. Watching binocular video is not acting, so you can get correlation only without any hope to learn how to distinguish correlations from causations, and at the same time it is much more limited in a data available. Science says that 80 or 90% of information people get is coming from their vision? I'm skeptical about this, because I don't know how they measure "information", but in any case human vision was trained with support from other senses. I wouldn't be surprised, if at certain stages of a baby's development other senses are more advanced and are used to get labelled data to train vision.