4 ms·
Interesting that it’s not vision based, I suspect you will get much better performance once vision is incorporated, using e.g LLaVa style models
by jerpint 3y ago
Interesting that it’s not vision based, I suspect you will get much better performance once vision is incorporated, using e.g LLaVa style models