6 ms·
This is awesome. How much effort does it take to go from this to a generalist robot: “Go to the kitchen and get me a beer. If there isn’t any I’ll take a seltze
by mbil 4y ago
This is awesome. How much effort does it take to go from this to a generalist robot: “Go to the kitchen and get me a beer. If there isn’t any I’ll take a seltzer”.
It seems like the pieces are there: ability to “reason” that kitchen is a room in the house, that to get to another room the agent has to go through a door, to get through a door it has to turn and pull the handle on the door, etc. Is the limiting factor robotic control?
- jah242 4y agoThis might be of interest to you (Google are getting there :))- https://palm-e.github.io https://palm-e.github.io
- cwillu 4y agoGPT-5 figures out that if it picks up the knife instead of the bag of chips, it can prevent the human with the stick from interfering with carrying out its instructions.
- airstrike 4y agoAnd ViperGPT will take said knife and make the muffin division fair when there there are an odd number of muffins by slicing either a muffin or a boy in half
- inawarminister 4y agoAh, the Solomon solution.
- jamilton 4y agoI wonder how much the hardware they're using costs.
- LeanderK 4y agoDisclaimer: I am not really into robotics. I think the limiting factors is the interface between ML models and robotics. We can not really train ML models end to end since since to train the interaction the model needs to interact, limiting the data size the model gets trained on. And simulations are not good enough for robust handling of the world. But I think we are getting closer.
- amelius 4y agoWhat I'd like to see is: "Take these pieces of LEGO and put them together given the assembly instructions in this booklet."
- spacebanana7 4y agoCould language models be able to avoid the need for labelled interaction data by developing a really good understanding of hardware documentation?
- alfalfasprout 4y agoTBH we're reaching a point where it's no longer about training a single model end-to-end. We now have computer vision models that can solve well-scoped vision tasks. Robots that can carry out higher level commands (going into rooms, opening doors, interacting with devices, etc.), and LLMs that can take a very high level prompt and decompose it into the "code" that needs to run. This all thus becomes an orchestration problem. It's just gluing together APIs admittedly at a higher level. And then you need to think about compute and latency (power consumption for these ML models is significant).
- westoncb 4y agoI suspect if an LLM were used to control a robot it would do so through a high level API that it's given access to; things like: stepForward(distance) or graspObject(matchId) The API's implementation may use AI tech too, but that fact would be abstracted.
- moffkalast 4y agoThat's definitely the interim solution until there's enough data to make it end-to-end. Right now there's more or less zero useful data on that.
- Bedon292 4y agoThe Boston Dynamics dog can open doors and things like that. It should be capable of performing all of the actions necessary to go get a beer. So I think it would be plausible to pull it all together, if you had enough money. It might take a bunch of setup first to program routes from room to room and things like that. Might look something like this: determine current room with an image from the 360 cam, select path from current room to target room, tell it to execute that path. Then use another image from the 360 cam and find the fridge. Tell it to move closer to the fridge, open the fridge, and take an image from the arm camera of the fridge content. Use that to find a beer or seltzer, grab it, and then determine the route to use and return with the drink. But, not so sure I would want to have it controlling 35+ kg of robot without an extreme amount of testing. And then there are things like: Go to the kitchen and get me a knife. Maybe not the best idea.
- hackerlight 4y agoThe point is to avoid the need to "program routes" or "determine current room". The LLM is supposed to have the world-understanding that removes the need to manually specify what to do.
- famouswaffles 4y agoIndeed an LLM doesn't need to be told what routes or actions to take to do that as has been demonstrated by palm-e and chatgpt for robotics.
- Bedon292 4y agoDetermine current room is a step GPT-4 would take care of by looking at the surroundings. The one thing I wasn't sure it could do, was figure out the layout of the house and determine a route for that. And I would rather provide it with some routes than have it wander around the house for an hour. I didn't figure real time video is what it was going to be best at. But it can certainly say the robot is in the living room, it needs to go down the hall to the kitchen. And if the robot knows how to get there already, it just tells the robot to go. I am sure there is another model out there that could be slotted in, but as far as just the robot plus GPT-4 goes, it might not quite be there. Just guessing at how they could fit together right now.
- cjohnson318 4y agoI think that even when systems are extremely accurate, the mistakes that they make are very un-human. A human might forget something, or misunderstand, but those errors are relatable and understandable. Automated systems might have the same success rate as human, but the errors can be very counterintuitive, like a Tesla coming to a stop on a freeway in the middle of traffic. There's things that humans would almost never do in certain situations. So yeah, I think that's the future, but I think the user experience will be wonky at times.
- ChatGTP 4y agoIt's also the kind of wonky that's like, a big problem wonky. "Plane taxis into fire truck" is especially not good wonky.
- lachlan_gray 4y agoI think we’re pretty much there. Like the other comment pointed out, palm-e is a glimpse of it. Eventually I think this kind of thing will work it’s way into autonomous cars and a lot of other mundane stuff (like roombas) as it becomes easier to do this kind of reasoning at the edge.
- maxwell 4y agoThe limiting factor may now mostly be cost. Notice where the funding is coming from on this though. Seems like the initial use case is more killer robots than robot butlers: situational awareness and target identification, under the guise of "common sense for robots." https://www.darpa.mil/program/machine-common-sense https://www.darpa.mil/program/machine-common-sense
- xapata 4y agoSometimes DARPA just funds basic-ish research (eg., the internet).
- maxwell 4y agoARPANET and TCP/IP were military tech first.
- eh9 4y agoI’m not advocating for killer robots, but wouldn’t we get the killer robots in our kitchens 10 years after the military gets them?