2 ms·
Vision-language-action models seem to be the broad category for the best approach, which basically combines a large vision-language model with robotic actions.
by Geee 1y ago
Vision-language-action models seem to be the broad category for the best approach, which basically combines a large vision-language model with robotic actions. For example, see https://www.physicalintelligence.company/blog/pi0 https://www.physicalintelligence.company/blog/pi0