4 ms·
According to this source, step 3 of the cascading model generates a 16 frame video at 24×48 resolution. So instead of sending a text prompt YouTube could almost
by fercircularbuf 4y ago
According to this source, step 3 of the cascading model generates a 16 frame video at 24×48 resolution. So instead of sending a text prompt YouTube could almost just as easily send 16 downsampled frames of the beginning of the video that your Neural Engine chip could work on instead.