28 ms·
I'd be interested to see if using DiffusionGemma-as-Jev helps as you can feed the image directly into the model and it'll make decisions based on the image embe
by mmastrac 17d ago
I'd be interested to see if using DiffusionGemma-as-Jev helps as you can feed the image directly into the model and it'll make decisions based on the image embeddings.
- nowittyusername 17d agoI had a long talk with chat gpt about this today as well. I think its duable and prolly not too hard either, also you could do lotsa funky stuff with stitched frames of a video in one 4x4 grid for example and send that as one image for analysis. that way temporal understanding can be had for fractions of a second by jev... also because vlm works in pixel space you can get around the whole state machine issue as well, so many possibilities...
- zjy365 17d ago[flagged]