4 ms·
How about Vision Language Models (VLMs)? They could easily take in a camera signal plus some basic hand sensor state and give structured output that would fit t
by nmstoker 2y ago
How about Vision Language Models (VLMs)? They could easily take in a camera signal plus some basic hand sensor state and give structured output that would fit the bill here.
https://huggingface.co/blog/vlms https://huggingface.co/blog/vlms