2 ms·
For the first time, it introduced native sparse attention into the full training process, achieving up to 11× inference speedup while maintaining model performa
by CalmStorm 1y ago
For the first time, it introduced native sparse attention into the full training process, achieving up to 11× inference speedup while maintaining model performance.