4 ms·
FWIW I have been using gptel with local inference by a deepcoder 14b q6 and llama.cpp and it’s not too bad. But it can’t handle big asks. 24GB of VRAM is not
by binary132 1y ago
FWIW I have been using gptel with local inference by a deepcoder 14b q6 and llama.cpp and it’s not too bad. But it can’t handle big asks. 24GB of VRAM is not enough to make local inference smart enough to be really useful. It’s okayish at things like writing commit messages or reformatting / refactoring local regions of code. 32B quants are slow and not that much better.