8 ms·
Jamesob's guide to running SOTA LLMs locally
- beardsciences 3mo agoI am somewhere in the middle, where I want something with more than 48GB/$2k of VRAM, but less than 384GB/$40k. I'm curious if GMKtec's EVO-X2, with ~96GB of usable VRAM, is still a good solution for something like this for $3,399.
- sampullman 3mo agoI picked up the 128gb version when it was $2,199 and it runs Qwen 3.6 reasonably well with a 128kb context. Not very useful for complex tasks but it can handle some web stuff.
- mft_ 3mo agoIt has lower memory bandwidth than most comparable Macs.
- verdverm 3mo agoI've been happy with an OEM Spark (128G), enough so that I picked up a second one. Have 2x qwen and 1x gemma (both at 8bit and full context), plus embedding, Re-Ranker, and a 1.7B for little things. Running 6x models, probably going to add STT here soon, want to try talking more than typing. The caveat is that if you try to use multiple models on the same device at the same time, you thrash and destroy tok/s
- maxothex 3mo ago[flagged]
- datadrivenangel 3mo ago"A great way to go is 2x RTX 3090s for a total of 48GB VRAM total. You can then run Qwen3.6-27B, which is an awesome model." Just want to note that for $3k you can get an M5 macbook pro with 48gb of shared memory, and it will not be a giant box. Also, consider committing to spending that money on a cloud hosting provider, which will be at least somewhat cheaper if not significantly cheaper. It is awesome being able to run models locally though.
- jbellis 3mo agoThat's a reasonable option, just be aware that you get about 1/3 as much memory bandwidth with the M5 Pro, or 2/3 with the M5 Max [now you're at $4100 for the lowest-end]. So both your prefill (flops-bound, M5 has a lot less) and decode (bw-bound) will be slower.
- LeBit 3mo agoI’m an idiot who is unable to project itself in situations I’ve never experienced before. So, I always thought local LLMs were toys not worth pursuing. Only once have I tried something decent like Gemma 4 31B and Qwen 3.6 27B did I realize how incredibly useful they are. You stop fearing you are sharing sensitive information. You stop fearing you will run out of tokens. You stop fearing about the availability of the remote AI. Local LLMs are extremely valuable.
- bityard 3mo ago*for many tasks
- Aurornis 3mo agoI have an M5 MacBook Pro and I also have a separate GPU setup for running models. The difference in speed is significant. It's not just token generation speed, but time to first token (prompt processing). The M5 hardware is amazing for what it is, but GPUs are still so much faster. Running the models on the GPU box also means I can use the laptop on my lap instead of turning it into a hot plate.
- amelius 3mo agoWhat is your GPU setup?
- brcmthrowaway 3mo agotinygpu kernel driver
- 3mo ago
- xela79 3mo agodid he call Qwen a SOTA model?
- zackify 3mo agoYou can get amazing local STT using parakeet which can use as little as 600mb of vram. Better or as good as whisper v3 large
- subhobroto 3mo ago[flagged]
- mrgaro 3mo agoHave we reached the capability of a local STT+LLM system being constantly listening for normal speech in a room and being able to understand when the human is addressing the system instead of talking to another human?
- kgeist 3mo ago>$40k gets you almost-Opus GLM 5.2 is "almost Opus," and it needs at least 8xH200s for comfortable inference (so it's closer to $400k than $40k). They suggest using this modified model: >A REAP-pruned (≈22% of experts removed), Int8-mix NVFP4 quantized version of GLM-5.2, ≈594B parameters. I wonder how it behaves in practice outside of benchmarks. Qwen3.6, even at 6-bit quantization, often gets stuck in loops while reasoning. And here they've also removed some experts. I mean, sometimes an 8-bit or 16-bit small model can be smarter than a lobotomized large model. I heard the consensus is you shouldn't go below 8 bit for coding. Also, it's not clear what is left of the available context when you try to fit a lobotomized model into 4 RTX 6000s. Anything below 100k is barely usable because it often hits compaction before it's able to gather the necessary context P.S. found in the repos, 240k context
- amelius 3mo agoHow does this work with scaling? I assume you can then somehow run several hundreds of prompts concurrently?
- CamperBob2 3mo agoYou can get 1M context with the lukealonso NVFP4 quant on 8x RTX6000s, which remains coherent and useful through at least 400k. No real need to run 8x H200s unless you just want to. Or unless you need to serve many concurrent users or agents on a regular basis.
- rsync 3mo ago"GLM 5.2 is "almost Opus," and it needs at least 8xH200s for comfortable inference ..." What is the behavior if one were to run GLM 5.2 with only a single H200 ? Would it fail to run at all, or would it just run so slowly as to be unusable ? I would like to prove out the build, and concept, of a SOTA model locally, but then backfill the rest of the GPUs in 18-24 months when they cost significantly less ...
- BoorishBears 3mo ago> in 18-24 months when they cost significantly less ... going to need you to sit down for this one...
- api 3mo agoApple M series chips deserve a mention as another option, especially since you get a whole Mac laptop or desktop workstation too. They have unified memory and respectable inference performance, and for some variations can be cheaper than video cards, especially if you get an older-gen high-end M series with a lot of RAM used or refurbished. I've read that Apple has plans once the RAM bottleneck passes to offer more RAM in all their models, and that future M series GPUs and NPUs will be even better for local inference, so in the future I expect Apple to be a serious offering for local inference and AI research workstations. And what about AMD and Intel Arc GPUs? They don't get as much love but I've heard they can be compelling for certain shapes of a local LLM configuration. At this point though, I think we may be in a "renters market" for LLM compute. If you want privacy it might be better to rent GPU time in raw form or use spot pricing at various providers. It probably only makes sense to build if you have extreme privacy/security needs or just want to do it cause it's cool.
- mwcampbell 3mo ago> once the RAM bottleneck passes Do we have evidence that this will actually happen? Maybe the belief that it won't pass is what requires evidence, but I think there's a widespread feeling right now that things are just getting permanently worse and this is one example.
- justincormack 3mo agoMicron have sold RAM for the next 4 years at current prices, so there are buyers expecting this to stay the same.
- api 3mo agoThat means buyers have basically purchased options. If the price falls, they're underwater a little, but if the price spikes it protects them. People do that all the time, and sometimes it doesn't pay off.
- api 3mo agoIt'll probably take a few years. There's many fabs under construction. One thing holding back capacity expansion is that a lot of people are concerned this is a bubble. They're worried it'll pop and leave them with orphaned assets if they over-invest in production. Of course maybe they're right and that will happen. If the data center construction boom ends, RAM prices will fall.
- wxw 3mo agoI agree that local LLMs are the likely future and worth investing in… but at $40k for possible-SOTA right now, this isn’t worth it for the average consumer. I’m pretty bullish that Apple will deliver something very competitive for the average consumer in the next couple years.
- Aurornis 3mo agoI play with local LLMs a lot. I've spent more on hardware than I should. I'm friends with a local group of people who have spent a lot more than I have. The warning I would have for everyone is to temper your expectations and read the fine print carefully. The big build in article starts off with a $40K budget and then includes 4 GPUs that are $12K each. For those doing the math, this build is going to cost more like 50-55K. Local setups also often rely on quantization and techniques like REAP to fit the models on their hardware. You will read a lot of claims that 4-bit quantization is lossless, but those claims come from KL divergence measurements on a small corpus. Use one of these 4-bit models on long context coding tasks and the quality will be noticeably less. Even for non-coding tasks like dataset analysis, I can measure a substantial quality difference between 4-bit models, 8-bit quants, and even some times the full 16-bit source. This article is also encouraging the use of a REAP model, which means someone has cut out some of the weights to make it smaller. The idea is to remove weights that are less useful for certain tasks, but again this is going to reduce the overall quality of the output. The trap is that people say "I'm running GLM-5.2 locally!" and it sounds amazing when you look at the GLM-5.2 benchmarks. However they're not actually running GLM-5.2, they're running a model derived from GLM-5.2 that discards most of the bits and drops some of the experts. It does not perform the same as what you see in the benchmarks. In my experience, the divergence between a quantized/REAP model and the parent model is unnoticeable when you try it on very small tasks or chat, but becomes painful when you start trying to use it on long-horizon tasks where little errors start compounding. Then you get into the slippery slope of thinking you're $50K deep into this project, but what you really need is just one or two more of those $12K GPUs to use the next level of quantization that might improve the quality a little more and make your investment worthwhile...
- CamperBob2 3mo agoAll very true. Right now, running GLM 5.2 at its full BF16 quantization level needs 1.5 TB of VRAM. You can't run this locally at a usable speed for less than $250K or so, and frankly I'd be surprised if it could be done for less than $500K. The best NV4FP quant for 5.2 appears to be lukealonso's at https://huggingface.co/lukealonso/GLM-5.2-NVFP4 https://huggingface.co/lukealonso/GLM-5.2-NVFP4, and it is capable of good throughput (75-100 tps) without losing much reasoning performance. Allowing for overhead for the KV cache and other requirements, this quant will (barely) run in 8-way tensor-parallel mode on 8x RTX 6000 cards. Not too long ago it was possible to put an 8x machine together for less than $100K USD, but that's probably not true now, assuming you buy all-new components. It'll almost certainly be worth it, given the abusive behavior we've seen and will continue to see from the major closed-model providers. If I hadn't already put a similar rig together, I'd be kicking myself. But getting it running well is by no means as simple as buying a bunch of RTX6K cards and calling it a day, and people need to know what they're getting into. Local AI is in its Altair and IMSAI days. There's no turnkey Apple II or C64 on the market yet, much less an IBM PC. Hardware, yes -- you can buy a capable box off the shelf from various vendors -- but you have to be prepared to take up a whole new hobby when it comes to getting a complete system working well.
- turova 3mo agoFor qwen3.6-27b you can also run the q4 variant with full ~250K context on one 3090. It's fast enough to not be frustrating so the speed gains with 2x 3090s wouldn't be worth it to me. Running a q6 on 2x 3090s at half the speed with a smaller context is an option, but you're really not going to compete with SOTA models there anyway so unless you already have 2x 3090s, I would say 1 is the best investment given current prices. It's good enough to do a lot, especially with a well-configured harness.
- hypfer 3mo agoThat math (250k context, Q4 model, 24GB VRAM) only checks out at q4 quant for the K/V cache, which is probably not the best idea.
- nabakin 3mo agoAre you running qwen3.6-27b on one 3090 with your KV cache at q4? Ime there is significant long-context recall accuracy degradation at that precision. I prefer putting the KV cache at q8 and working with the 120k context
- Der_Einzige 3mo agoUse modern samplers and you don’t need to limit yourself to 8bit at half the context window. I could push it down to 1.58 bits and get decently good output easily by simply not using the garbage default top_p and top_k that vendors continue to wrongly recommend.
- anon373839 3mo agoWhere do you find optimal samplers and sampler settings for these models? Very interested in this as I, too, use Q8.
- chompychop 3mo agoIs Whisper still considered SOTA for STT? Since it came out years ago, I'd have assumed there are better models by now.
- randomblock1 3mo agoNo, there are quite a few models which are smaller, more accurate, and faster. For example Parakeet TDT v3 is half the size, way faster, and lower WER. There's also Voxstral, which is much larger but also even more accurate. But the ecosystem isn't as mature, so Whisper is still a valid option, even now. For example Parakeet uses Nemotron framework (made by Nvdia), normally you need CUDA, so you need to use an ONNX version instead on AMD. Meanwhile Whisper has VLLM and desktop apps like Buzz. There aren't many benchmarks and they often don't have all the models, since STT doesn't get nearly enough attention as normal LLMs, but this is one of the more complete ones: https://artificialanalysis.ai/speech-to-text/non-streaming https://artificialanalysis.ai/speech-to-text/non-streaming
- venusenvy47 3mo agoI don't have anything to compare against, since I have just started using it. But I was fairly happy with it on my personal recordings from my phone. Also, I ran it on my CPU (Core i7) and it was perfectly usable, as something to run when not using the machine for anything else.
- simonw 3mo agoI'm a big fan of Parakeet v3 - I run it using the MacWhisper app, it's a 494MB model and the quality is excellent.
- 3eb7988a1663 3mo agoRelated - what is the best isolation system available? Do I have to go full, fat VMs or can I get by with a Firecracker-like thing? Seemingly every available option has some subtle-gotchas about how easy it is to blow off your foot and effectively have no security at all. I use VMs because I actually trust that security is a foundational principle of the technology, not a well-if-you-use-these-20-flags-and-squint kind of deal.
- ZiiS 3mo agoFull fat VMs with GPU passthough I trust a lot less then CPU ones.
- elsombrero 3mo agofrom my understanding, you can run the inference server (llama.cpp/vllm/whatever) and the agent/harness in different contexts, event different machines. The risky part is in the agent/harness and what tools it has access to. You don't need to give GPU passthrough to the VM running the agent/harness. There is still a risk of a prompt messing with the inference server, but I think that's a much lower risk compared to an agent doing whatever on its own.
- dofm 3mo agoRight. All my experiments are naïve, I am sure, but I run the LLM on the host and expose it via OpenAI API to the VMs. This approach requires that you trust the llama.cpp codebase, essentially. It might be reasonable not to. I suppose in principle there is the risk of a prompt exploit corrupting the inference server.
- Catloafdev 3mo agoIt depends - for what? If your security model is sandboxing an agent to ensure they don't nuke your PC, then there are a lot of options, you can use something like bubblewrap[1] or a microVM like libkrun[2] if your goal is light-weight, up to full Docker if you want the tooling that comes with that. [1] https://github.com/containers/bubblewrap https://github.com/containers/bubblewrap [2] https://github.com/libkrun/libkrun https://github.com/libkrun/libkrun
- bobkb 3mo agoVery useful. The whisper setup is something similar to what we have been using. The LLM setup though is outstanding.
- vector_vibe 3mo agoi have the same Whisper Setup too!
- whalesalad 3mo agowhy in gods name is a RTX PRO 6000 $13,000? supply and command?
- pulse7 3mo agoexactly: high demand!
- gizajob 3mo agoAbout £4000 on eBay uk right now. Because if they were any lower we’d all be buying six each.
- petu 3mo agoNvidia reuses numbers for workstation cards, so there's multiple vastly different 'RTX 6000' cards: Quadro RTX 6000 (Turing / 2080 Ti / 24GB @ 672.0 GB/s) RTX A6000 (Ampere / 3090 / 48GB @ 768.0 GB/s) RTX 6000 Ada (4090 / 48GB @ 960.0 GB/s) RTX PRO 6000 Blackwell (5090 / 96GB @ 1.79 TB/s) For £4000 you were likely looking at RTX 6000 Ada listing.
- jacobgold 3mo ago> "~$40k At this price level, you get the next step up in model intelligence. Something pretty close to Claude Opus." That is equivalent to 16.8 years of Claude Opus 4.8 or Codex GPT 5.5 at $200/mo. I'm a huge fan of running local models, but they're still wildly expensive, lower quality, and possibly dangerous (if backdoored). I sincerely wish this wasn't the case.
- simonw 3mo agoThat $200/month is already more like $4,000/month if you have to pay full API pricing - "enterprise" companies for example. That drops the equivalent to 10 months. (I'd be surprised if that local rig really can drive the equivalent of $4,000/month of API spend though, given that a local rig can run prompts in parallel a lot less effectively than Anthropic's many data centers.)
- fweimer 3mo agoI think the decode phase of inference typically uses local compute resources poorly due to the very small batch size. If you can run many inference tasks in parallel, this will make local inference more competitive to centralized inference, not less.
- echelon 3mo agoStop trying to run them locally, folks. You don't own your fiber connection. So why try to own another rapidly depreciating, expensive, and annoying asset? Rent cloud GPUs! You get to participate in the ownership, data control, price control, and hacking culture without having to Frankenstein some hobbyist box that costs a ton, is distilled down to functional uselessness, and is a PITA to maintain.
- satvikpendem 3mo agoIf I'm gonna rent cloud GPUs I might as well just use a subsidized cloud agent like Claude or Codex. As for depreciation, that is true, but the bet is that models get better for a certain parameter count faster than your hardware becomes obsolete, such as Gemma models for example at the same 30 billion parameter count being much better than some years ago.
- Avicebron 3mo agoDoes anyone know any good data center to home conversion kits for gear?
- bcjdjsndon 3mo agoIf you can run sota on a 40k setup, why do openai etc spend maybe 100x that?
- dwroberts 3mo agoObvious one: Because they are serving it to millions of people at the same time, not just one local user
- c4pt0r 3mo agoLocal open weight models will definitely be a future trend. Imagine if an Opus-level model could run locally: many more latent use cases would likely emerge, since Opus is priced so high. Perhaps the future will be a multi-model architecture, where frontier models handle planning and local models carry out the concrete execution.
- maxxxml 3mo agoWhat harness is the best for local LLMs? I've been researching optimizing local LLM agent harness performance with context/ tools. Quite the endeavor and would love to learn what users prefer for this type of workflow.
- npodbielski 3mo agoI like vibe and pi. Vibe just looks nice and is good enough. But pi extensibility is just another level. There is also Dirac that is quite OK but seems like full of bugs. Zerostack is the simplest harness I saw. OpenCode is OK too. Rest I did not try.
- jzer0cool 3mo agoWhat's the technical reason we call call these a harness? Seems right but want to understand better.
- maxxxml 3mo agoThe model represents intelligence and the harness is toolset which allows the model to create more informed decisions with context. Specifically, loops, subagents, tools, connectors, prompts, skills, and much more. This is why Cursor performs so well.
- GTP 3mo agoThere also exists an in-between possibility, that is, if you get 128GB of vram (there are now multiple options in the market to get that amount with a unified memory architecture) you can run DeepSeek V4 flash at good speed via DwarfStar. I'm not going to spend money on this, but my gut feeling is that this would be the right compromise for a lot of people.
- jonaustin 3mo agoI just started using it on an m4 max 128 and it's the first time since buying the machine a year ago that it feels like local llm "just works" for reasonably decent coding. Use pi though; claude code has way too much bootstrap context; slows everything way down.
- LoganDark 3mo agoDefinitely seconding pi. Also avoid opencode, it doesn't support caching (mutates the system prompt constantly)
- rishabhaiover 3mo agoThis is a great guide. However, the economics just do not work in my favor at all. Even if I were to spend $2k, I get much more flexibility of model intelligence and choice from a provider for $20/month.
- QuantumNoodle 3mo ago$2k or $40k? One of those is not "self host."
- gizajob 3mo agoDepends how much money you made going long on AI stocks.
- maxignol 3mo agoDid not seem to find how much tokens per second he achieved with this setup ?
- aetherspawn 3mo ago80 tok/s which is kind of a lot for GLM. My experience running 80 tok/s on other LLM is that it ~seems faster than cloud inference, but that obviously depends what you use, in my case ChatGPT.
- SwellJoe 3mo agoI recently wrote up how I run local LLMs, because several folks had asked (https://swelljoe.com/post/how-i-run-local-llms/ https://swelljoe.com/post/how-i-run-local-llms/) and I think even my setup, which I spent maybe $4200 on, half on a Strix Halo and half on upgrades for my desktop, would be too expensive to justify today. I bought before prices went through the roof, and only did so because I like to tinker with hardware...not because I expected it to ever pay for itself vs. buying subsidized tokens from the big guys or the cheap tokens from efficient providers like DeepSeek. Buying four $13000 GPUs and several thousand dollars worth of supporting hardware seems crazy. This supply shortage has to end eventually, and I can buy billions of DeepSeek, MiMo, and GLM tokens, and use $100 or $200 a month subscriptions for the big guys in the meantime for the difference in price once that happens. And, you can't even run the full-sized GLM on that hardware, it is quantized and so is your KV cache; the degradation is small, but not non-existent. You're not running a model that's equal to what you get when you buy GLM tokens from Z.ai. My recommendation for self-hosting is this: If you already have a 24GB or 32GB GPU, or two, or a recent Mac with 32GB or more, run the appropriate quantization of Qwen 3.6 27B or Gemma 4 31B. If your hardware is older and too slow for that, use the MoE, but know it'll be dumber. Use the tiny model for the stuff that doesn't need deep smarts: Research (give it a Brave or Exa MCP for web search), summarization, simple Python scripts for basic tasks, simple websites or web apps, categorization of stuff (I used Gemma 4 to review my past writing for friendliness and helpfulness), etc. It can also be a sub-agent for bigger agents (for those same kinds of tasks). Gemma 4 12B is an incredibly good model for its size, particularly for vision tasks, and in the 4-bit quantization (7GB on disk) it runs on anything, even a modern tablet or phone. And, if you don't already have a big GPU or unified memory Mac, just wait. Use the cheap tokens every AI company wants to sell you, for now. A Claude or Codex or Gemini subscription is a good deal. Tokens from DeepSeek are a good deal, especially with Reasonix agent (which maximizes caching, which DeepSeek is uniquely good at, and cached tokens are uniquely cheap at DeepSeek). GLM is Good Enough and has a cheap coding plan. MiMo has the cheapest tokens for a 1T+ model in the game, though DeepSeek and GLM are better models, MiMo is fine. When prices come down, I'll be speccing out a beast to run the big models, too. But, I'm not paying 4x for RAM and GPU and storage, and y'all shouldn't either. That's crazy. Computer prices go down over time. It is the law.
- 3mo ago
- weystrom 3mo agoWhile I think that local LLMs are the future, i think these setups are insane. You shouldn't be trying to push the SOTA, most people underestimate how much you can get out of small LLMs. Why ask FABLE 5000 to "summarize this email thread" when a tiny model can do the job? Sure Codex3000 can oneshot your backlog, but why not use a subsidized subscription to do it for now? We're clearly not at the peak of these model's capabilities yet.
- saltamimi 3mo agoCould someone give me an actual guide for spending as little as possible to get as maximal gains with either SOTA or cheap models as a systems administrator and not someone like a full-stack developer? The models are so powerful and consequently so expensive and confusing to use, I don't get all of it.
- ineptech 3mo agoMight as well add my own experience since I just set up a local llm this week. I went with a 32GB card made by Intel called Arc B70, which is cheaper than a 3090 and more has ram, at the cost of a slower memory bus. edited to remove something incorrect, thanks diablod3 I went with this because a) the models I wanted to use are a little too big to fit comfortably in 24gb, plus I need room for a few additional small models for autocomplete and speech recognition, and b) I already had a cheap server to use and dual gpus would've required upgrading the mobo and power supply and probably the case as well. It was definitely a little tricky to set up. The Intel line requires a driver package called "level zero" to support something called SYCL (Intel's version of CUDA basically, AFAICT) that was tricky to get working. I am running llama.cpp in a docker container, which also required some fiddling to get the container to see the card. You also need a kernel from the last few months. Once I got it working though, the results are very impressive for a $1k investment. Qwen 3.6 35B at q4 quantization takes about 3/4 of the ram and delivers like 88 tokens/sec. So, if you want a decent-sized model for cheap, this is one way to go.
- DiabloD3 3mo agoThat is incorrect. They both have GDDR6. The B70 has 256 bit it bus at a clock speed of 2375mhz (608 GB/s), the 3090 has a 384 bit bus at a clock speed of 2438mhz (936 GB/s). It isn't slower, it just has less channels, ie, it is less wide.
- ineptech 3mo agoWhoops thanks, was going from memory. At any rate, the effect is that it's somewhat slower than the 3090, when using a model small enough to fit entirely in nvram, but can fit models the 3090 can't.
- DiabloD3 3mo agoYep, but the flip side is you can get two B70s for the price of a single 3090 (MSRP, obviously; 3090s used go for about the same as a new B70), and they're true 2 slot, so they can fit on x8/x8 consumer boards fine. The side effect is Intel has fired their entire Arc team, including the driver team, as well have canceled all Celestial products (only low end combined GPU + IO die products will have Celestial, as its too late in production to change). All future Intel products will have Nvidia graphics tiles. Llama.cpp is, of course, still trying to better support the existing Intel products, but its kinda hard when they might stop working any day due to driver breakage.
- charcircuit 3mo agoIf you want to host SotA models you need multiple machines. 384 GiB is nowhere near enough for SotA where models are terabytes big.
- misiti3780 3mo agoDoesnt an NVIDA Spark solve most of these problems? (at 5K)
- pulse7 3mo agoNVidia Spark is much slower (low memory bandwidth)!
- mateenah 3mo agoThis is extremely useful. Thank you so much!
- nullc 3mo agoThose cards would really prefer you use a pcie-5 switch, but I guess they're sold out.
- throwrioawfo 3mo agoIf you're going to fork out 40k, why not get an actual rack rather than fashioning one yourself out of plywood...
- gehsty 3mo agoAre they SOTA? I’m not sure
- tomnow 3mo ago[flagged]
- tomuow 3mo ago[flagged]
- gchamonlive 3mo agoThere's a sub 2k tier with a single 3090 that's also serviceable. Run https://github.com/noonghunna/club-3090 https://github.com/noonghunna/club-3090 with beellama, fast inference at the cost of a reduced 102k context window
- ursuscamp 3mo agoBitcoin is so dead that jamesob is posting about AI.
- brcmthrowaway 3mo agoBitcoin booster -> AI slopper pipeline
- luciana1u 3mo ago[flagged]
- joka88xj 3mo ago[flagged]
- nnevatie 3mo ago> SOTA LLMs locally Shouldn't the headline be about running SOTA _local_ LLMs, as GLM 5.2 is nowhere near a SOTA LLM?
- luciana1u 3mo ago[flagged]
- luciana1u 3mo ago[flagged]
- jkwang 3mo ago[dead]
- rldjbpin 3mo agono matter your luck with hardware or your sysadmin skills, doing local inference for just yourself and/or to emulate typical usage (e.g. your coding workflow and deep research, etc.) is just very inefficient in current model architecture. to me, this is a "truck" approach to city driving as a single person who does not do furniture hauling every weekend. the sense of privacy and freedom is nice but online inference is more "economical" as multi-user load is more effectively served than going solo. maybe new architectures would make it effective to do text inference locally [1], till then great on you if you can spend car money on your setup. hope it is a great learning experience as well. [1] https://deepmind.google/models/gemma/diffusiongemma/ https://deepmind.google/models/gemma/diffusiongemma/
- djx22 3mo agoin my experience running models that have been heavily quantized(q4) or altered to some extent has never made me say “wow, this is so amazing”. On the contrary, the model ended up in the thrash bin after a few prompts. I have an RTX 6000 PRO with 96GB, and what I can run comfortably is Qwen 3.6 27B or MoE, Gemma 4 31B. This is as far as it goes when you run the model at full precision and maximum context length. They perform well and you can use them for coding, doing research on the internet and what have you. So if you do the math and you see yourself spending more than the $2400/year to Anthropic, then it might make sense to get one of these cards but accept the quality drop. Otherwise, will humans even be coding in 5 years from now?
- broadsidepicnic 3mo agowhat you maybe forget here is the use case for people and businesses who can not send the data to 3rd party due to privacy/contractual reasons. This is what I'm looking at, we're bound by strict policies for data sharing outside of our premises.
- djx22 3mo agoyes, I understand the usecase. Where I'm coming from is quantized vs. unquantized. 4bit quants are lobotomizing the model heavily to the point that it's better to invest in some capable hardware than keep fighting the limitation. Refurbished server grade hardware is accessible. For the price of an RTX 6000 PRO you could probably get much more VRAM but 1-2 generations older.
- crymeth0t 3mo agoIt's nutty to me that anyone would go to such great lengths to use LLMs -- especially chasing the bleeding edge like this. If Claude and co. disappeared tomorrow, I wouldn't flinch. I don't understand why people are exchanging their brain wrinkles for access to a slop machine. I wonder if a good analogue would be a skilled carpenter being offered access to a machine which excretes furniture (one or two levels of quality beneath Ikea). Does it do the job? Most of the time. Does the carpenter enjoy the process? No.
- doobidoo 3mo ago[flagged]
- Fno44 3mo ago[flagged]