3 ms·
So, it's D-Flash but at each transformer layer and share the KV cache of the original model? Very smart!
by littlestymaar 5mo ago
So, it's D-Flash but at each transformer layer and share the KV cache of the original model? Very smart!
- foobar10000 5mo agoKindof yeah - predictivity is a question though for larger layers - when trying to scale this up. But yeah, this is a "95% predictor in latent space is a 7x improvement in speed if done right" approach.