3 ms·
> It will happen for images and video. That's for sure. The main problem is how to put them together. There are several ways. First, tokenizing the media and
by two_in_one 3y ago
> It will happen for images and video.
That's for sure. The main problem is how to put them together. There are several ways. First, tokenizing the media and inserting it in the stream. Second, a media controller which accepts verbal instructions generated by LLM. They can be combined. Interesting variation is controller with (immediate) feedback. It can be anything, for example database access.
> I argue that it will even extend to things like time series analysis
AFAIK transformers have been used for low level robotic control. And LLM for high level have been reported by MS and Google.