3 ms·
LLM inference is mostly read only, so high-bandwidth flash looks like it could provide huge cost savings over VRAM. It's not yet in commercial products but ther
by mrob 8mo ago
LLM inference is mostly read only, so high-bandwidth flash looks like it could provide huge cost savings over VRAM. It's not yet in commercial products but there are working prototypes already. Previous HN discussion:
https://news.ycombinator.com/item?id=46700384 https://news.ycombinator.com/item?id=46700384
- whosegotit 8mo agoAre you saying that intel’s optane product was just ahead of its time? Is optane the answer to LLM’s ever increasing appetite?