5 ms·
I'm always wondering how people are affording kubernetes clusters and all that to run stuff like ollama.
by sweeter 2y ago
I'm always wondering how people are affording kubernetes clusters and all that to run stuff like ollama.
- skinkestek 2y agoThe author is using https://k3s.io/ https://k3s.io/ not the full k8s, so it doesn't have to be extremely expensive.
- tbrownaw 2y agoK3s is a proper full k8s and has instructions for running multiple-node clusters, it just ships with batteries included and with defaults that play nice with having a single node.
- skinkestek 2y agoThanks! Today I learned!
- suprjami 2y agoYou can run a local LLM on a $100 minipc with Vulkan GPU acceleration and get a usable token generation count.