3 ms·
They're trying to corner the market and betting on the running costs of these models going down substantially in the future.
by swexbe 3y ago
They're trying to corner the market and betting on the running costs of these models going down substantially in the future.
- rgbrgb 3y agoI think it’s a good bet based on watching inference speed of llama.cpp consistently improving and model ability / size on a similar trajectory. I expect there’s similar room for optimization (probably more) with hosted models. If you don’t care about code/model privacy (I do but it seems like most don’t yet at least) there are even more batching/caching tricks to engineer.