4 ms·
> It's free It isn't, the cost is included in your electricity bill, not even talking about the cost of your time to set it up. It's very possible that it cost
by usrnm 1mo ago
> It's free
It isn't, the cost is included in your electricity bill, not even talking about the cost of your time to set it up. It's very possible that it costs you more than a cloud mode would, you just don't want to calculate it properly.
- exe34 29d agoFree heating in winter.
- BoredomIsFun 1mo ago> It's very possible that it costs you more than a cloud mode would ...which is almost always true in a single request/reply mode and never true in batch mode. Single request usually 2x-3x more expensive than cloud and batch mode 2x-3x cheaper. Now, for narrow tasks, a finetuned tiny 8b model would dramatically outperform SOTA frontiers for a fraction of price, esp. on energy efficient hardware like Apple.
- visarga 1mo agoLocal is never cheaper than cloud because they can do batch inference, and that means you load model weights once to produce 128 tokens on 128 sessions in parallel not 1 token on 1 session like local models. Local models rarely get to high utilization factor, they spend most of their time waiting. If you had only batch inference and enough of it to fill the compute to 80% then you get cheaper local models.
- BoredomIsFun 1mo ago> not 1 token on 1 session like local models. Local models can absolutely run in batch, what are even talking about? > If you had only batch inference and enough of it to fill the compute to 80% then you get cheaper local models. Even if you ran sequentally, single session, a _finetuned_ tiny (8B) local model on narrow tasks would abolutely mog SOTAs, any of it - Fable, Opus, Sol you name it.
- helsinkiandrew 1mo ago> Local models can absolutely run in batch, what are even talking about? I think the point was that if you aren't running your local machine at 100% for 24 hours a day then a cloud - with multiple clients - that is, will be more efficient.
- visarga 1mo agoIf you buy the computer specifically for inference it is more expensive than cloud, but if you had it anyway it's free.
- trainingonme 1mo agoTrue, but how many people (realistically) buy a computer with 48GB+ of RAM?
- NamlchakKhandro 1mo ago48gb of vram. a machine like this is about a years rent for most people. a small car for most others.
- LeBit 1mo agoI think he’s talking about the Mac Mini unified memory. 48G RAM is pretty useful if you want to run k8s locally for tests / exploration
- bel8 1mo agotrue but if you're actually running k8s and similar workloads, chances are it might eat memory that LLM requires. you'll also notice these articles rarely specify their context window in tokens, because it is small, usually 30k to 70k tokens and it gets slower as it fills up.
- LeBit 1mo agoI actually have a Mac Mini M4 Pro with 48G. I gave the k8s example because this is what I was doing with it. Was because I am back to using Linux as my workstation. My Mac Mini is now a headless server for llama.cpp. So, you are right that for these workloads , I would not be using the Mac Mini for k8s AND llama. Another thing going against using a Mac for Linux containers is that there are no solutions that I know that properly manages memory : memory is given to the Linux vm , but never fluctuates if the needs in the vm are less than the initial request. I know Orb Stack does that but is it proprietary. I think UTM does it , but not sure I would use UTM instead of Lima, Colima , multipass , etc to run containers.
- thrw93747572007 1mo agoIt sounds like they are doing something similar to what I described in my other post below. Personal media station. That can be done on hardware that quite a lot of people basically just have and don't use 24/7 to the max - because it is their gaming machine or their programming and compiling workhorse, for example. Of course you are paying for additional electricity but even with napkin-math instead of a "proper" calculation, you are unlikely to pay more for running your own instead of something commercial (and that can be offset further with some of the "modern" electricity contracts and/or PV and battery storage). Especially if we are talking about a stack that runs most of/all the time when you are not using your machine and makes LLM calls regularly while running. The work in software/admin to get whatever you want set up is similiar no matter which infrastructure you use.
- albrewer 29d ago> the cost is included in your electricity bill Acting like an extra $20 on my electric bill is equivalent to a $200/mo subscription is... a take.