3 ms·
It's still much worse than 1M context on 16GB VRAM with Reformer, but at the cost of inference speed. And you can use FlashAttention in your own models to get a
by bitL 4y ago
It's still much worse than 1M context on 16GB VRAM with Reformer, but at the cost of inference speed. And you can use FlashAttention in your own models to get a more efficient/sparse attention now as well.
- meghan_rain 4y agoHow could one apply the mentioned technologies to llama/alpaca?
- Tenoke 4y agoThe quality with reformer is much much worse, it's not really comparable.
- bitL 4y agoYeah, but it fits on a single GPU. Now imagine it scaled across 1000 GPUs.
- Tenoke 4y agoI finetuned one in 2020[0] to play around with and the results still seemed a bit worse than a gpt of comparable size. 0. https://svilentodorov.xyz/blog/reformer-99m/ https://svilentodorov.xyz/blog/reformer-99m/