4 ms·
All those choices seem to have very different trade-offs? I hate $5,000 as a budget - not enough to launch you into higher-VRAM RTX Pro cards, too much (for me
by clusterhacks 10mo ago
All those choices seem to have very different trade-offs? I hate $5,000 as a budget - not enough to launch you into higher-VRAM RTX Pro cards, too much (for me personally) to just spend on a "learning/experimental" system.
I've personally decided to just rent systems with GPUs from a cloud provider and setup SSH tunnels to my local system. I mean, if I was doing some more HPC/numerical programming (say, similarity search on GPUs :-) ), I could see just taking the hit and spending $15,000 on a workstation with an RTX Pro 6000.
For grins:
Max t/s for this and smaller models? RTX 5090 system. Barely squeezing in for $5,000 today and given ram prices, maybe not actually possible tomorrow.
Max CUDA compatibility, slower t/s? DGX Spark.
Ok with slower t/s, don't care so much about CUDA, and want to run larger models? Strix Halo system with 128gb unified memory, order a framework desktop.
Prefer Macs, might run larger models? M3 Ultra with memory maxed out. Better memory bandwidth speed, mac users seem to be quite happy running locally for just messing around.
You'll probably find better answers heading off to https://www.reddit.com/r/LocalLLaMA/ https://www.reddit.com/r/LocalLLaMA/ for actual benchmarks.
- kpw94 10mo ago> I've personally decided to just rent systems with GPUs from a cloud provider and setup SSH tunnels to my local system. That's a good idea! Curious about this, if you don't mind sharing: - what's the stack ? (Do you run like llama.cpp on that rented machine?) - what model(s) do you run there? - what's your rough monthly cost? (Does it come up much cheaper than if you called the equivalent paid APIs)
- clusterhacks 10mo agoI ran ollama first because it was easy, but now download source and build llama.cpp on the machine. I don't bother saving a file system between runs on the rented machine, I build llama.cpp every time I start up. I am usually just running gpt-oss-120b or one of the qwen models. Sometimes gemma? These are mostly "medium" sized in terms of memory requirements - I'm usually trying unquantized models that will easily run on an single 80-ish gb gpu because those are cheap. I tend to spend $10-$20 a week. But I am almost always prototyping or testing an idea for a specific project that doesn't require me to run 8 hrs/day. I don't use the paid APIs for several reasons but cost-effectiveness is not one of those reasons.
- bigiain 10mo agoI don't suppose you have (or would be interested in writing) a blog post about how you set that up? Or maybe a list of links/resources/prompts you used to learn how to get there?
- clusterhacks 10mo agoNo, I don't blog. But I just followed the docs for starting an instance on lambda.ai and the llama.cpp build instructions. Both are pretty good resources. I had already setup an SSH key with lambda and the lambda OS images are linux pre-loaded with CUDA libraries on startup. Here are my lazy notes + a snippet of the history file from the remote instance for a recent setup where I used the web chat interface built into llama.cpp. I created an instance gpu_1x_gh200 (96 GB on ARM) at lambda.ai. connected from terminal on my box at home and setup the ssh tunnel. ssh -L 22434:127.0.0.1:11434 ubuntu@<ip address of rented machine - can see it on lambda.ai console or dashboard> Started building llama.cpp from source, history: 21 git clone https://github.com/ggml-org/llama.cpp 22 cd llama.cpp 23 which cmake 24 sudo apt list | grep libcurl 25 sudo apt-get install libcurl4-openssl-dev 26 cmake -B build -DGGML_CUDA=ON 27 cmake --build build --config Release MISTAKE on 27, SINGLE-THREADED and slow to build see -j 16 below for faster build 28 cmake --build build --config Release -j 16 29 ls 30 ls build 31 find . -name "llama.server" 32 find . -name "llama" 33 ls build/bin/ 34 cd build/bin/ 35 ls 36 ./llama-server -hf ggml-org/gpt-oss-120b-GGUF -c 0 --jinja MISTAKE, didn't specify the port number for the llama-server 37 clear;history 38 ./llama-server -hf Qwen/Qwen3-VL-30B-A3B-Thinking -c 0 --jinja --port 11434 39 ./llama-server -hf Qwen/Qwen3-VL-30B-A3B-Thinking.gguf -c 0 --jinja --port 11434 40 ./llama-server -hf Qwen/Qwen3-VL-30B-A3B-Thinking-GGUF -c 0 --jinja --port 11434 41 clear;history I switched to qwen3 vl because I need a multimodal model for that day's experiment. Lines 38 and 39 show me not using the right name for the model. I like how llama.cpp can download and run models directly off of huggingface. Then pointed my browser at http//:localhost:22434 on my local box and had the normal browser window where I could upload files and use the chat interface with the model. That also gives you an openai api-compatible endpoint. It was all I needed for what I was doing that day. I spent a grand total of $4 that day doing the setup and running some NLP-oriented prompts for a few hours.