3 ms·
Even more important in a local context is the difference between token generation and prompt processing speed. We tend to focus on the former, but for multi-tur
by c7b 3mo ago
Even more important in a local context is the difference between token generation and prompt processing speed. We tend to focus on the former, but for multi-turn/agentic workflows the latter can dominate.
- kpw94 3mo agoYeah definitely. I've recently commented on that: https://news.ycombinator.com/item?id=48557890 https://news.ycombinator.com/item?id=48557890