3 ms·
One thing I find particularly interesting is that SOTA video understanding requires significantly less parameters than SOTA language understanding. Are we just
by jerpint 2y ago
One thing I find particularly interesting is that SOTA video understanding requires significantly less parameters than SOTA language understanding. Are we just overfitting way too much to language?
Also how long until SAM gets aligned to an LLM? Would be great to natively prompt it and not hackily chain through separate vision models
- random17 2y agoI wouldn’t call SAM video “understanding” though, it’s a model whose sole job is to segment frames into distinct objects, and has not demonstrated any innate understanding of the physics or logic of the videos themselves.