6 ms·
Inferencing can be done in software entirely (e.g. INT8) but it's very slow compared to GPU or APU. nVidia cornered the market because everything (tensorflow an
by timnetworks 3y ago
Inferencing can be done in software entirely (e.g. INT8) but it's very slow compared to GPU or APU. nVidia cornered the market because everything (tensorflow and everything after) is optimized for it, but you can get good results on AMD now, and on ARC too in some cases. And slow results entirely in software (CPU-RAM), which for personal and non-constant use may be just fine too.
- user_7832 3y agoThanks! Do you have any guides/websites/github repos for running these models on CPUs?
- rcarmo 3y agoOllama will (nearly always) work provided you have enough RAM. I was actually pretty surprised that it didn't work on my N5105 (which has 16GB) because it relies on AVX instructions...
- user_7832 3y agoThanks! Someone else mentioned llama.cpp but it appears that ollama is just a gui frontend for llama (which is good because I find guis easier). I'll hopefully set it up soon!
- FergusArgyll 3y agoIt's not a GUI it's a cli but very easy to use "ollama run {model}". you can also `ollama serve` which serves an api, and then you can use or build a simple gui.
- user_7832 3y agoThanks, I’ll keep that in mind!
- jmorgan 3y agoNext upcoming Ollama version will support non-AVX CPUs