3 ms·
Also see the demo from Jason Dai's post: https://www.linkedin.com/posts/jasondai_with-the-latest-ipex-llm-llamacpp-portable-activity-7303194182729244673-FcxL ht
by colorant 2y ago
Also see the demo from Jason Dai's post: https://www.linkedin.com/posts/jasondai_with-the-latest-ipex-llm-llamacpp-portable-activity-7303194182729244673-FcxL https://www.linkedin.com/posts/jasondai_with-the-latest-ipex...
- aurareturn 2y agoCPU inference is both bandwidth and compute constrained. If your prompt has 10 tokens, it’ll do ok, like in the LinkedIn demo. If you need to increase the context, compute bottleneck will kick in quickly.
- colorant 2y agoPrompt length mainly impacts prefill latency (FTFF), not the decoding speed (TPOT)
- moffkalast 2y agoDecoding speed won't matter one bit if you have to sit there for 5 minutes waiting for the model to ingest a prompt that's two sentences long.
- colorant 2y agoWith ~1000 input, the TTFT is ~10 seconds