5 ms·
The issue is that the field is still moving too fast - in 20 months, you might break even on costs, but the LLMs you are able to run might be 20 months behind "
by silversmith 1y ago
The issue is that the field is still moving too fast - in 20 months, you might break even on costs, but the LLMs you are able to run might be 20 months behind "state of the art". As long as providers keep selling cheap inference, I'm holding out.
- wmf 1y agoThe gap between local models and SOTA is around 6 months and it's either steady or dropping. (Obviously this depends on your benchmark and preferences.)
- criddell 1y agoSeriously? So I can run the best models from 2024 at home now? For example, what would I need to run Open AI's o1 model from 2024 at home? Are there good guides for setting this up?
- wmf 1y agoIt's not the same model, but for example GPT-OSS-120B is smarter than o1. The guide is buy 128 GB of VRAM then install LM Studio.
- criddell 1y agoAn NVIDIA 5090 with 128 GB of VRAM is $13k. It doesn’t make any sense to run that at home when you can pay OpenAI $20 / month to use it (it would take more than 50 years to spend $13k at OpenAI this way). So technically you might be able to run a six month old model at home, but it would be foolish to do so from a financial point of view. Or is there a way to get 128 GB of VRAM for a lot less than that?
- wmf 1y agoRyzen AI Max is $2,000, M4 Max is $3,500, and DGX Spark is $4,000. Still not really economically feasible but I see it as an insurance policy. And that's the most expensive local model; smaller models will run on any PC.
- tra3 1y agoThat's where I am at too. Also it's not clear what's going to happen with hardware prices. I think there's a huge demand for hardware right now, but it should fall off at some point hopefully.
- ants_everywhere 1y agoI agree, but also don't underestimate the value of developing a competency in self-hosting a model. Dan Luu has a relevant post on this that tracks with my experience https://danluu.com/in-house/ https://danluu.com/in-house/
- rz2k 1y agoFortunately the models are increasing in efficiency about as fast as they are increasing in performance, so your homelab surprisingly doesn’t become out of date as fast as you might expect. However, I expect there will also be very capable machines like 1TB or 2TB Mac Studio M5 or M6 Ultras within a year or two.
- qingcharles 1y agoAgreed. Right now this stuff is being massively subsidized in our favor. I'm taking advantage of that.