3 ms·
Running LLMs with 3.3M Context Tokens on a Single GPU
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- charlie_xxx 2y agoTheir demo looks really cool: https://github.com/mit-han-lab/duo-attention https://github.com/mit-han-lab/duo-attention
3 ms·