5 ms·
How are you engaging with the fundamental problems explained in this blog post https://rodneybrooks.com/why-todays-humanoids-wont-learn-dexterity/ https://rodne
by anonymid 2mo ago
How are you engaging with the fundamental problems explained in this blog post https://rodneybrooks.com/why-todays-humanoids-wont-learn-dexterity/ https://rodneybrooks.com/why-todays-humanoids-wont-learn-dex...
- the relatively crude tactile and proprioceptive sensing apparatuses of robots when compared to humans
- the limited availability of multisensory, perception-action coupled training data
Genuinely curious!
- nl 2mo agoThis seems like a post by someone who hasn't really ingested the bitter lesson. Eg, even LeRobot (without proper fingers) can fold clothes now: https://www.youtube.com/watch?v=dPe9v4gqbdg https://www.youtube.com/watch?v=dPe9v4gqbdg The labs are spending huge money collecting "multisensory, perception-action coupled training data" (eg, there is the one in NY that gives you free cleaning in return for video data from the cleaner). Edit: The Gemini Robotics blog post has a video of it tying knots too. That's pretty good.
- jhanschoo 2mo agoYou can point to the bitter lesson to support your claim, but on the other hand I can point to the massive investment in capital and time to get self-driving cars viable to support my more bearish view.
- AndrewKemendo 2mo agoYesterday evening I rode 12 miles in a Waymo actively dodging pedestrians and obstacles dynamically in an open ended environment. Multiples better experience than the two Ubers I had later that evening. What are you bearish about precisely?
- qsera 2mo agoIt is amusing to observe that the tech marketing of today are milking the shit out of this trick. The trick being to tread continuously through some non-obvious happy path. And average people will be convinced that you really have some breakthrough tech. But hey, this is not something new. Magicians were taking advantage of such things for centuries ..
- AndrewKemendo 2mo agoPlease describe the trick you’re suggesting
- qsera 2mo ago1. Make some thing that work in very limited of amount of real world cases 2. Deploy it somewhere where it won't encounter things it won't handle. 3. Market the shit out of the above fact and how well it work there. 4. Let the naive population who have a tendency to take one look, and imagine how it will automatically progress to some arbitrary influx point. 5. Get a lot of funding from people in point 4 and feed it to point 3, and keep going.
- jryle70 2mo ago> 1. Make some thing that work in very limited of amount of real world cases > 2. Deploy it somewhere where it won't encounter things it won't handle Who are doing these? Waymo? how? You're talking BS if you can't elaborate.
- qsera 2mo agoOf course it will appear BS to you. That is why it works.
- AndrewKemendo 2mo agoWhy couldn’t this trick be pulled off two decades ago? After all autonomous vehicles has been well funded research since the 1980s the DARPA grand challenge being one of the previously most important benchmarks. I think you might just need a history lesson friend
- qsera 2mo agoInternet and the ad/marketing/propoganda/pr powers that comes with it..
- jhanschoo 2mo agoI'm not talking about the state of today's self-driving cars, I'm talking about how we got to here. Also don't forget the overarching claim regarding lack of sensor and actuator fidelity in the parent comment; self-driving cars in contrast have expensive LIDAR in them and we still don't think camera-based self-driving cars are safe enough.
- AndrewKemendo 2mo agoYea we got here through recognizing the bitter lesson is a description of reality and adopting techniques that exploit it The entire history of this field is precisely that problem and repeatedly demonstrated
- jhanschoo 2mo agoI don't know how you can read my comment, not respond to my comment saying that today's self-driving cars use LIDAR, and continue to reiterate your point. I don't think I was clear and explicit, it being tedious to write, and I apologize for that. I also apologize for shifting the goalposts as I had not written out my own position, which is not exactly in the "grandparent commenter"'s position (that I had not previously given enough attention understanding), but it is also not in agreement with yours. I don't mean to say that we are not presently in the "bitter lesson" (your idea of what the bitter lesson says) regime. I definitely think that a lot of progress can be done right now by emphasizing the humanoid robotics platform as a foundation. What I mean to say is that I don't know if that platform with the hardware we have today is sufficient for parity with human housekeeping tasks in the domains that we wish it to have parity. The bitter lesson itself (not your understanding of it) is in fact silent on this as it is in relation to feature engineering, where it is a clear point, but you seem to be adapting it uncritically wholesale to mean something more than what it is written about. My position is that, it is unclear whether today's sensor platform is sufficient for parity. It is less strong than the blog post author's "Why Today’s Humanoids Won’t Learn", it is a "We can't say whether or not today's humanoids will learn", but it is something that also contradicts a "the bitter lesson means today's humanoids will learn" thesis. The self-driving car supports my claim, because after so much investment in capital and time, we ended up with a car with comparatively expensive LIDAR sensors as our preferred platform.
- qsera 2mo agoBut how does LLMs help in making chat bots better, help with this "multi-sensory" data. Has this anything to do with the current AI surge with LLMs? Has it been demonstrated? Or is it just that such things now get a lot of funding now, for no good reason?
- nl 2mo ago>But how does LLMs help in making chat bots better, help with this "multi-sensory" data. All three offerings from the linked blog post are either Vision LLMs or Vision/Action LLMs.
- qsera 2mo agoAction LLMs work by generating text underneath. Just some higher level software interprets the text generated and do some action. So the immediate inference result is still text. That does not help a lot.
- ACCount37 2mo agoDepends entirely on VLA arch. Some have dedicated action diffusion heads that work in a standalone non-text action output space. Much like an LLM can either use an external TTS or have audio output heads attached to it directly for native S2S. But your entire premise is wrong regardless of that. Even if VLAs were forever bound to outputting text, you'd have to prove that they're fundamentally incapable of emitting text that maps to useful action sequences. No proof of that whatsoever - and plenty of empirical evidence suggests otherwise. Even non-specialist LLMs like ChatGPT are getting better at controlling robots and navigating 3D environments, if slowly.
- qsera 2mo ago>But your entire premise is wrong regardless of that. You don't understand what I am saying. The crux of your misunderstanding is here >emitting text that maps to useful action sequences If you have a static mapping from text to action, then you are throwing away all the advantage of using an AI. The whole point of AI is that you can get an output from an input without explicit mapping. So If you use explicit mapping anywhere in the chain, then you lose most of the advantage of using the AI. So if your hardware, physical vocabulary is limited, like move left/right/up/down then what you say could work. But something that have the dexterity of a human form, this vocabulary is nearly infinite. You won't be able to use explicit mapping there.
- YeGoblynQueenne 2mo agoAt this point the bitter lesson has become a meaningless shibboleth. It seems very few people have actually read the original article: The Bitter Lesson Rich Sutton March 13, 2019 http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html And even fewer are aware of the author's follow up on what his article says about the current trend in AI: Silicon Valley Doesn't Understand The Bitter Lesson – Richard Sutton https://youtu.be/QMGy6WY2hlM?si=0aOmgKiPfGEEXoK9 https://youtu.be/QMGy6WY2hlM?si=0aOmgKiPfGEEXoK9
- ACCount37 2mo agoIt's a load of bollocks, and always was. Modern robotics is, at its core, not a hardware problem. It's an AI problem. We have plenty of headroom in the hardware - what we don't have is an AI good enough to utilize it. We don't know the practical limits of current hardware because we can't make a robot AI that would make the hardware a meaningful bottleneck. Today's robots don't fail at tasks because they have poor fingers. They fail because they don't know how to perform those tasks. If you put an effort into solving that? You get demos like: Gemini Robotics 2 tying a garbage bag. Take one long look at that and think of manual dexterity. Human body is crude and suboptimal in a thousands different ways, and all of it is salvaged by advanced intelligence.
- YeGoblynQueenne 2mo agoYeah, look carefully at that demo of tying the garbage bag strings: it's done in a very peculiar style that suggests a very specific, very precise, "algorithm" taught in an imitation learning session, which has no chance to transfer to other tasks, or even other garbage bag strings. As usual with robot tech demos: WYSIWYG.
- ACCount37 2mo agoDid the past decades of AI research teach you absolutely nothing? Every time you see something that "suggests a very specific, very precise, "algorithm" taught in an imitation learning session"? Scale the imitation learning up x10, x100, x1000, and it suddenly generalizes! I'll be honest: I don't see what you see. I don't see anything that would suggest this algorithm is so brittle there's zero transfer to "even other garbage bag strings". AI robotics isn't innately brittle like conventional robotics is. But even if you are, somehow, completely right on that? Teach a hundred "very specific algorithms" like this - and watch them fuse into a manifold of algorithms that can be applied to different problems as needed. And that is what you need. If an algorithm for "tie a garbage bag with current generation robot hands" exists and can be learned by an AI, then the gains from getting better AI are far from exhausted. The limits of robotics are the limits of AI. This is why every AI robotics company is saying "we need more data". They understand what they're dealing with. They looked at the scaling laws and went "robotics isn't magic, that curve applies to us too". I don't get what makes people see robotics as a special magic thing, that makes them look at the advances in robot AI and say "this is intractable" and not "this is hard". It's hard. We're getting through it though.