3 ms·This is mostly about inference speed, while maintaining long context performance.by alekandreev 2y agoThis is mostly about inference speed, while maintaining long context performance.