3 ms·
llama3:70b using llama.cpp (used under the hood by Ollama) on a 11th Gen Intel i5-11400 @ 2.60GHz - no GPU, CPU inference only. "Write a haiku about Hacker New
by programd 2y ago
llama3:70b using llama.cpp (used under the hood by Ollama) on a 11th Gen Intel i5-11400 @ 2.60GHz - no GPU, CPU inference only.
"Write a haiku about Hacker News mentioning AI in the title"
Here is a haiku:
AI whispers secrets
HN threads weave tangled debate
Intelligence born
eval time = 30363.04 ms / 23 runs ( 1320.13 ms per token, 0.76 tokens per second)
total time = 34294.80 ms / 33 tokens
- bityard 2y agoThat really doesn't seem bad. When people talk about responses of self-hosted LLMs without a beefy GPU being unusably slow, I always assumed they meant 15 minutes to hours. I do not mind waiting a few minutes if it will summarize the answer a question that will take me many times longer to research.
- logicallee 2y agohow much disk space did it use?