3 ms·
My best guess is that you want a supervisor GPT-4 like LLM planning the task and a lower level on-prem LLM doing the tasks like driving from one location to ano
by gitfan86 3y ago
My best guess is that you want a supervisor GPT-4 like LLM planning the task and a lower level on-prem LLM doing the tasks like driving from one location to another or grasping an item.
Sending every frame to GTP-4 right now is way too slow. But at Tesla FSD like model can drive from one location to another in a closed environment with perfection. All that is missing is training that style in a roomba/robot form and then having GTP-4 monitor and manage the tasks at a 10 or 20 second interval
- sho_hn 3y agoOh, I'm OK with going slow! It doesn't have to be all that practical, I'm more curious about playing with toy approaches. Trying to populate a world model with few captioned still frames plus basic IMU/dead reckoning seems like a fun challenge ... Reminds me of the recent HN post on jumping spider intelligence: They can do complex route planning, but need to stare for hours before they get going. This is probably more down to their tiny field of view on the front-facing good eyes, but still :-)