3 ms·
So do we think that some form of this is what they are using internally to get those long context lengths by already using sub-quadratic architectures [1] in th
by snthpy 1y ago
So do we think that some form of this is what they are using internally to get those long context lengths by already using sub-quadratic architectures [1] in their deployed models?
1: "Context Is The Next Frontier by Jacob Buckman, CEO of Manifest AI" (https://youtu.be/wJyl6kBCwmY?si=ruxMWdENjazu3rp6 https://youtu.be/wJyl6kBCwmY?si=ruxMWdENjazu3rp6)