4 ms·
Seems to me that the way to architect this is to have multiple quasi independent embedded controllers at different levels. For example, you might have 3 indepe
by empath-nirvana 3y ago
Seems to me that the way to architect this is to have multiple quasi independent embedded controllers at different levels. For example, you might have 3 independent finger controllers, managed by a hand controller, and so on, on up to the highest level LLM that drives everything. So you have an LLM that just says, pick up the green can, issues whatever structured data needs to go to the arm controller and on down the line, going to down to more real-time and less high level processing as you go.
As a human, I don't understand the detailed micro-second by microsecond movements my fingers have to do to pick something up, let alone how I'm touch-typing this sentence. It just sort of "happens" when I want it to happen. I don't think you need to design a robotic AI that understands how every part of it's mechanics work. The fingers don't need to know how the feet work, for example. There should be semi-autonomous "intelligence" embedded throughout the system, with only necessary feedback being fed back up.
- RecycledEle 3y agoI agree. We need feedback loops on each joint that keep doing what they are doing until a higher feedback loop decides to change things. For example, as I enter this text, most of my fingers are holding the phone and do not need to be told to keep holding it. The low level controllers should also be able to respond to simple things. It's like when I get hurt, my hand yanks back before my brain realizes what has happened. Or when I almost trip and my legs correct before I realized it happened. Low level tasks should be at a low level and not bother the higher professors.
- RecycledEle 3y agoLLMs have terrible spatial awareness. This probably comes from ONLY being trained on text. I wonder if it would help a LLM running a robot to have a separate controller calculating it's position and what is around it, and feed that into the LLM constantly. It would be like when a video game has a radar display or map to let you know where you are in relation to other things.
- moffkalast 3y agoI actually tried that a while back, giving 3.5-turbo a multishot prompt that consisted of distance readings for ahead, left, right and back in an array, as extracted from lidar data, then giving it movement instructions. It performed rather terribly. You've very much correct that their spatial awareness is terrible. Something as simple as drive forward, then back, turn left, etc. works just fine and they can generally translate it to a specified message format reasonably reliably, but give them something more complex to execute, like drive a robot in a square pattern (an example answer would be go forward, turn right, go forward, turn right, etc.) they start to generate nonsense. I also tested it with the 30B WizardLM at the time which performed almost as well in terms of message format but had even worse awareness. Part of the problem is that the training data contains next to no examples that would teach it how 3D space works. I considered making a dataset of driving a robot around with human movement commands and then logging the aggregated sensor data and commands for fine tuning so the prompt format would be pre-learned, but I'm not entirely sure how much it would help.
- adeelzaman 3y agoHi @moffkalast, I am interested in doing something similar (building a larger dataset of robot movements) and fine-tuning an LLM with it. Would love to have a quick call and chat further. You can reach me at "hi [at] adeelzaman [dot] me".
- StackOverlord 3y agoLLMs have a firm grip on common sense. It's because it allows them to deal with the utterly unexpected they are deemed useful in robotics. Not to perform delicate movements, but stop doing so when police enters the room.
- NalNezumi 3y agoThat is indeed what a lot of Machine Learning turned Robotics researchers/enthusiasts are banking on. A counter argument to that is what Dhruv Batra responded to the question "lol why not use LLM for everything" [1]. >As a human, I don't understand the detailed micro-second by microsecond movements my fingers have to do to pick something up, let alone how I'm touch-typing this sentence. It just sort of "happens" when I want it to happen. I don't think you need to design a robotic AI that understands how every part of it's mechanics work. This is true for us ofc, and is encompassed in what is called the "Moravec's paradox" [2]. You don't understand it because it's unconscious process, and It's harder to reverse-engineer an unconscious process (motor movement) than conscious ones (calculating math, playing game, writing text, reading). But the thing is that in the real world we do need to take in to account everything, including noise and time-delay. Evolution gave rise to complex language in the last 100k years compared to millions of year for motor movement. I do agree that there must be some "hierarchical" structure for complex motion, but we currently don't know where and how that hierarchy is. Boston dynamics uses Model Predictive Control for complex movement, which means that at least some model of the world is required, for dexterous motion. Now if that model is part of LLM or not is a hard guess. But if we don't know this it's hard to say what kind of data we need to collect to train a LLM-model applied to embodiment (robotics). Researchers in the past have already made the mistaken assumption of "oh cognition and language is the hard part of intelligence. Perception and motion is easy" [3] and then their work amounted to nothing because turns out the latter was way, wayyyyyy harder. There's an implicit bias in us that think Language, puzzles and logic are harder [2] and therefore models that accomplish this can just be rammed in to the "easier" issues. Edit: I too would like "LLM models will solve these" attitude because otherwise the research I'm doing right now is a dead-end, but the more I try (with my limited compute) the less I'm sure [1] https://imgur.com/eWsH5ui https://imgur.com/eWsH5ui originally https://twitter.com/DhruvBatraDB/status/1641871357020614656 https://twitter.com/DhruvBatraDB/status/1641871357020614656 [2] https://en.wikipedia.org/wiki/Moravec%27s_paradox https://en.wikipedia.org/wiki/Moravec%27s_paradox [3] https://youtu.be/x10964w00zk?list=PLSQhB89mdG7PsZsDz2_5hZL8CACYMUqAT&t=3358 https://youtu.be/x10964w00zk?list=PLSQhB89mdG7PsZsDz2_5hZL8C...
- whinenot 3y ago>It's harder to reverse-engineer an unconscious process Aside from some basic life support systems, don't almost all movements start with conscious effort? Whether you are deliberate about the exercise or not, you practice and practice until you develop 'muscle memory' where it becomes unconscious: walking, dribbling a basketball, holding a G chord on a guitar, etc.
- jimbokun 3y agoMaybe the octopus is a more tractable model for robotic control. I understand that octopus neurons are not as concentrated in a central brain, but spread throughout its limbs that are autonomous compared to humans.
- StackOverlord 3y agoThis is in fact what happens in the human body. For example, when you reach out to pick up a green can, your brain makes the decision to do the task but it's your spinal cord and peripheral nerves that carry out the detailed work – orienting the hand, managing grasp strength, controlling the arm movements etc. This process is mostly unconscious – you don't need to actively think about how to tense each muscle in the same way that an embedded controller wouldn't need to understanding the working of the entire robotic system to carry out its specific task. Much like the model suggested, the human body communicates feedback across layers — this process is crucial to maintaining balance, coordination and effectively reacting to the environment. For instance, if your fingers touch a hot stove, the sensory receptors in your skin will immediately send a signal to your spinal cord and a reflex action will make you pull your hand back even before you consciously perceive that the stove is hot.