4 ms·
Show HN: oLLM – LLM Inference for large-context tasks on consumer GPUs
- deleted 1y ago[deleted]
- anuarsh 1y agoHi everyone, any comments or questions are appreciated
- attogram 1y ago"~20 min for the first token" might turn off some people. But it is totally worth it to get such a large context size on puny systems!
- Haeuserschlucht 1y ago20 minutes is a huge turnoff, unless you have it run over night.... Just to get the hint that you should exercise self care in the morning when presenting a legal paper and have the ai check it for flaws.
- anuarsh 1y agoWe are talking about 100k context here. 20k would be much faster, but you won't need KVCache offloading for it
- Haeuserschlucht 1y agoIt's better to have software erase all private details from text and have it checked by cloud ai to then have all placeholders replaced back at your harddrive.