13 ms·
Can I run AI locally?
- bheadmaster 7mo agoMissing 5060 Ti 16GB
- nicklo 7mo agothe animation of the model name text when opening the detail view is so smooth and delightful
- John23832 7mo agoRTX Pro 6000 is a glaring omission.
- schaefer 7mo agoNo Nvidia Spark workstation is another omission.
- embedding-shape 7mo agoYeah, that's weird, seems it has later models, and earlier, but specifically not Pro 6000? Also, based on my experience, the given numbers seems to be at least one magnitude off, which seems like a lot, when I use the approx values for a Pro 6000 (96GB VRAM + 1792 GB/s)
- sxates 7mo agoCool thing! A couple suggestions: 1. I have an M3 Ultra with 256GB of memory, but the options list only goes up to 192GB. The M3 Ultra supports up to 512GB. 2. It'd be great if I could flip this around and choose a model, and then see the performance for all the different processors. Would help making buying decisions!
- utopcell 7mo agoUnfortunately, Apple retired the 512GiB models.
- ProllyInfamous 7mo agoSure, but those already sold still exist.
- gentleman11 7mo agoask apple to graciously allow you to install your own ram in the computer you "own"
- ActorNightly 7mo ago>. I have an M3 Ultra with 256GB of memory, Im sorry but spending this kind of money when you could have just built yourself a dual 3090 workstation that would have been better for pretty much everything including local models is just plain stupid. Hell, even one 3090 can now run Gemma 3 27b qat very fast.
- xiconfjs 7mo agoExcept if you are living in a region where electricity is quite expensive :/
- brulard 7mo agoAre you aware that your 3090s have nowhere close to 256GB of VRAM? Or maybe you are not aware that on macs you have unified memory (working both as RAM and VRAM).
- ActorNightly 7mo agoAre you aware that having ram doesn't matter when your tokens/second is slow as shit? You don't need to run large models, Gemma QAT 27B fits on one GPU and is quite good. Other models like Qwen3 are great for coding. 3090 gets 100+ tokens/second for QWEN, very close to what you would see with a cloud based model. M3 ultra gets ~30. Congrats, you played yourself.
- brulard 7mo agoDid I? Not only are you comparing apples to oranges, you even provide misleading numbers. 3090 gets 20-30 tokens a second for dense ~30B models (QwQ 32B, Gemma 3 27B Q4), similar to M3 ultra. If you are talking about Qwen3-Coder 30B (MoE), then both 3090 and M3 Ultra are around ~70 tok/s. But even if you were right about the speed - which you are not - speed is pointless if you need large model that wouldn't fit into your VRAM.
- nozzlegear 7mo ago> a dual 3090 workstation that would have been better for pretty much everything Doesn't run macOS
- GrayShade 7mo agoThis feels a bit pessimistic. Qwen 3.5 35B-A3B runs at 38 t/s tg with llama.cpp (mmap enabled) on my Radeon 6800 XT.
- Aurornis 7mo agoAt what quantization and with what size context window?
- GrayShade 7mo agoLooks like it's a bit slower today. Running llama.cpp b8192 Vulkan. $ ./llama-cli unsloth_Qwen3.5-35B-A3B-GGUF_Qwen3.5-35B-A3B-UD-Q4_K_XL.gguf -c 65536 -p "Hello" [snip 73 lines] [ Prompt: 86,6 t/s | Generation: 34,8 t/s ] $ ./llama-cli unsloth_Qwen3.5-35B-A3B-GGUF_Qwen3.5-35B-A3B-UD-Q4_K_XL.gguf -c 262144 -p "Hello" [snip 128 lines] [ Prompt: 78,3 t/s | Generation: 30,9 t/s ] I suspect the ROCm build will be faster, but it doesn't work out of the box for me.
- phelm 7mo agoThis is awesome, it would be great to cross reference some intelligence benchmarks so that I can understand the trade off between RAM consumption, token rate and how good the model is
- S4phyre 7mo agoOh how cool. Always wanted to have a tool like this.
- adithyassekhar 7mo agoThis just reminded me of this https://www.systemrequirementslab.com/cyri https://www.systemrequirementslab.com/cyri. Not sure if it still works.
- twampss 7mo agoIs this just llmfit but a web version of it? https://github.com/AlexsJones/llmfit https://github.com/AlexsJones/llmfit
- deanc 7mo agoYes. But llmfit is far more useful as it detects your system resources.
- dgrin91 7mo agoHonestly I was surprised about this. It accurately got my GPU and specs without asking for any permissions. I didnt realize I was exposing this info.
- dekhn 7mo agoHow could it not? That information is always available to userspace.
- bityard 7mo ago"Available to userspace" is a much different thing than "available to every website that wants it, even in private mode". I too was a little surprised by this. My browser (Vivladi) makes a big deal about how privacy-conscious they are, but apparently browser fingerprinting is not on their radar.
- swiftcoder 7mo agoIt's pretty hard to avoid GPU fingerprinting if you have webgl/webgpu enabled
- dekhn 7mo agoWe switched to talking about llmfit in this subthread, it runs as native code.
- rithdmc 7mo ago
- mrdependable 7mo agoThis is great, I've been trying to figure this stuff out recently. One thing I do wonder is what sort of solutions there are for running your own model, but using it from a different machine. I don't necessarily want to run the model on the machine I'm also working from.
- cortesoft 7mo agoOllama runs a web server that you use to interact with the models: https://docs.ollama.com/quickstart https://docs.ollama.com/quickstart You can also use the kubernetes operator to run them on a cluster: https://ollama-operator.ayaka.io/pages/en/ https://ollama-operator.ayaka.io/pages/en/
- rebolek 7mo agossh?
- g_br_l 7mo agocould you add raspi to the list to see which ridiculously small models it can run?
- vova_hn2 7mo agoIt says "RAM - unknown", but doesn't give me an option to specify how much RAM I have. Why?
- charcircuit 7mo agoOn mobile it does not show the name of the model in favor of the other stats.
- debatem1 7mo agoFor me the "can run" filter says "S/A/B" but lists S, A, B, and C and the "tight fit" filter says "C/D" but lists F. Just FYI.
- metalliqaz 7mo agoHugging Face can already do this for you (with much more up-to-date list of available models). Also LM Studio. However they don't attempt to estimate tok/sec, so that's a cool feature. However I don't really trust those numbers that much because it is not incorporating information about the CPU, etc. True GPU offload isn't often possible on consumer PC hardware. Also there are different quants available that make a big difference.
- havaloc 7mo agoMissing the A18 Neo! :)
- arjie 7mo agoCool website. The one that I'd really like to see there is the RTX 6000 Pro Blackwell 96 GB, though.
- ge96 7mo agoRaspberry pi? Say 4B with 4GB of ram. I also want to run vision like Yocto and basic LLM with TTS/STT
- boutell 7mo agoI've been trying to get speech to text to work with a reasonable vocabulary on pis for a while. It's tough. All the modern models just need more GPU than is available
- ge96 7mo agoWhispr? For wakewords I have used pico rhino voice I want to use these I2S breakout mics
- meatmanek 7mo agoFor ASR/STT on a budget, you want https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3 https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3 - it works great on CPU. I haven't tried on a raspberry pi, but on Intel it uses a little less than 1s of CPU time per second of audio. Using https://github.com/NVIDIA-NeMo/NeMo/blob/main/examples/asr/asr_streaming_inference/asr_streaming_infer.py https://github.com/NVIDIA-NeMo/NeMo/blob/main/examples/asr/a... for chunked streaming inference, it takes 6 cores to process audio ~5x faster than realtime. I expect with all cores on a Pi 4 or 5, you'd probably be able to at least keep up with realtime. (Batch inference, where you give it the whole audio file up front, is slightly more efficient, since chunked streaming inference is basically running batch inference on overlapping windows of audio.) EDIT: there are also the multitalker-parakeet-streaming-0.6b-v1 and nemotron-speech-streaming-en-0.6b models, which have similar resource requirements but are built for true streaming inference instead of chunked inference. In my tests, these are slightly less accurate. In particular, they seem to completely omit any sentence at the beginning or end of a stream that was partially cut off.
- LeifCarrotson 7mo agoThis lacks a whole lot of mobile GPUs. It also does not understand that you can share CPU memory with the GPU, or perform various KV cache offloading strategies to work around memory limits. It says I have an Arc 750 with 2 GB of shared RAM, because that's the GPU that renders my browser...but I actually have an RTX1000 Ada with 6 GB of GDDR6. It's kind of like an RTX 4050 (not listed in the dropdowns) with lower thermal limits. I also have 64 GB of LPDDR5 main memory. It works - Qwen3 Coder Next, Devstral Small, Qwen3.5 4B, and others can run locally on my laptop in near real-time. They're not quite as good as the latest models, and I've tried some bigger ones (up to 24GB, it produces tokens about half as fast as I can type...which is disappointingly slow) that are slower but smarter. But I don't run out of tokens.
- uncSoft 7mo ago[dead]
- Felixbot 7mo ago[flagged]
- itigges22 7mo ago[flagged]
- sshagent 7mo agoI don't see my beloved 5060ti. looks great though
- xylon 7mo agonor my plain 5060
- carra 7mo agoHaving the rating of how well the model will run for you is cool. I miss to also have some rating of the model capabilities (even if this is tricky). There are way too many to choose. And just looking at the parameter number or the used memory is not always a good indication of actual performance.
- jrmg 7mo agoIs there a reliable guide somewhere to setting up local AI for coding (please don’t say ‘just Google it’ - that just results in a morass of AI slop/SEO pages with out of date, non-self-consistent, incorrect or impossible instructions). I’d like to be able to use a local model (which one?) to power Copilot in vscode, and run coding agent(s) (not general purpose OpenClaw-like agents) on my M2 MacBook. I know it’ll be slow. I suspect this is actually fairly easy to set up - if you know how.
- AstroBen 7mo agoOllama or LM Studio are very simple to setup. You're probably not going to get anything working well as an agent on an M2 MacBook, but smaller models do surprisingly well for focused autocomplete. Maybe the Qwen3.5 9B model would run decently on your system?
- jrmg 7mo agoRight - setting up LM studio is not hard. But how do I connect LM Studio to Copilot, or set up an agent?
- brcmthrowaway 7mo agoBasically LM Studio has a server that serves models over HTTP (localhost). Configure/enable the server and connect OpenCode to it. Try this article https://advanced-stack.com/fields-notes/qwen35-opencode-lm-studio-agentic-coding-on-m1.html https://advanced-stack.com/fields-notes/qwen35-opencode-lm-s... I'm looking for an alternative to OpenCode though, I can barely see the UI.
- AstroBen 7mo agoCodex also supports configuring an alternative API for the model, you could try that: https://unsloth.ai/docs/basics/codex#openai-codex-cli-tutorial https://unsloth.ai/docs/basics/codex#openai-codex-cli-tutori...
- AstroBen 7mo ago
- AstroBen 7mo agoThis doesn't look accurate to me. I have an RX9070 and I've been messing around with Qwen 3.5 35B-A3B. According to this site I can't even run it, yet I'm getting 32tok/s ^.-
- misnome 7mo agoIt seems to be missing a whole load of the quantized Qwen models, Qwen3.5:122b works fine in the 96GB GH200 (a machine that is also missing here....)
- mongrelion 7mo agoWhich quantization are you running and what context size? 32tok/s for that model on that card sounds pretty good to me!
- unfirehose 7mo ago[flagged]
- varispeed 7mo agoDoes it make any sense? I tried few models at 128GB and it's all pretty much rubbish. Yes they do give coherent answers, sometimes they are even correct, but most of the time it is just plain wrong. I find it massive waste of time.
- boutell 7mo agoI'm not sure how long ago you tried it, but look at Qwen 3.5 32b on a fast machine. Usually best to shut off thinking if you're not doing tool use.
- mongrelion 7mo agoApparently there is a whole science behind running models. I have seen the instructions that unsloth publishes for their quants and depending on the model they'll tweak things like the temperature, top k, etc. The size of the quantization you chose also makes a difference. The GPU driver also plays an important role. What was your approach? What software did you use to run the models?
- orthoxerox 7mo agoFor some reason it doesn't react to changing the RAM amount in the combo box at the top. If I open this on my Ryzen AI Max 395+ with 32 GB of unified memory, it thinks nothing will fit because I've set it up to reserve 512MB of RAM for the GPU.
- bityard 7mo agoYeah, this site is iffy at best. I didn't even see Strix Halo on the list, but I selected 128GB and bumped up the memory bandwidth. It says gpt-oss-120b "barely runs" at ~2 t/s. In reality, gpt-oss-120b fits great on the machine with plenty of room to spare and easily runs inference north of 50 t/s depending on context.
- kylehotchkiss 7mo agoMy Mac mini rocks qwen2.5 14b at a lightning fast 11/tokens a second. Which is actually good enough for the long term data processing I make it spend all day doing. It doesn’t lock up the machine or prevent its primary purpose as webserver from being fulfilled.
- freediddy 7mo agoi think the perplexity is more important than tokens per second. tokens per second is relatively useless in my opinion. there is nothing worse than getting bad results returned to you very quickly and confidently. ive been working with quite a few open weight models for the last year and especially for things like images, models from 6 months would return garbage data quickly, but these days qwen 3.5 is incredible, even the 9b model.
- sroussey 7mo agoNo, getting bad results slowly is much worse. Bad results quickly and you can make adjustments. But yes, if there is a choice I want quality over speed. At same quality, I definitely want speed.
- meatmanek 7mo agoThis seems to be estimating based on memory bandwidth / size of model, which is a really good estimate for dense models, but MoE models like GPT-OSS-20b don't involve the entire model for every token, so they can produce more tokens/second on the same hardware. GPT-OSS-20B has 3.6B active parameters, so it should perform similarly to a 3-4B dense model, while requiring enough VRAM to fit the whole 20B model. (In terms of intelligence, they tend to score similarly to a dense model that's as big as the geometric mean of the full model size and the active parameters, i.e. for GPT-OSS-20B, it's roughly as smart as a sqrt(20b*3.6b) ≈ 8.5b dense model, but produces tokens 2x faster.)
- lambda 7mo agoYeah, I looked up some models I have actually run locally on my Strix Halo laptop, and its saying I should have much lower performance than I actually have on models I've tested. For MoE models, it should be using the active parameters in memory bandwidth computation, not the total parameters.
- littlestymaar 7mo agoWhile your remark is valid, there's two small inaccuracies here: > GPT-OSS-20B has 3.6B active parameters, so it should perform similarly to a 3-4B dense model, while requiring enough VRAM to fit the whole 20B model. First, the token generation speed is going to be comparable, but not the prefil speed (context processing is going to be much slower on a big MoE than on a small dense model). Second, without speculative decoding, it is correct to say that a small dense model and a bigger MoE with the same amount of active parameters are going to be roughly as fast. But if you use a small dense model you will see token generation performance improvements with speculative decoding (up to x3 the speed), whereas you probably won't gain much from speculative decoding on a MoE model (because two consecutive tokens won't trigger the same “experts”, so you'd need to load more weight to the compute units, using more bandwidth).
- lambda 7mo agoSo, this is all true, but this calculation isn't that nuanced. It's trying to get you into a ballpark range, and based on my usage on my real hardware (if I put in my specs, since it's not in their hardware list), the results are fairly close to my real experience if I compensate for the issue where it's calculating based on total params instead of active. So by doing so, this calculator is telling you that you should be running entirely dense models, and sparse MoE models that maybe both faster and perform better are not recommended.
- nilslindemann 7mo ago1. More title attributes please ("S 16 A 7 B 7 C 0 D 4 F 34", huh?) 2. Add a 150% size bonus to your site. Otherwise, cool site, bookmarked.
- amelius 7mo agoWhy isn't there some kind of benchmark score in the list?
- aplomb1026 7mo ago[dead]
- amelius 7mo agoWhat is this S/A/B/C/etc. ranking? Is anyone else using it?
- vikramkr 7mo agoJust a tier list I think
- relaxing 7mo agoApparently S being a level above A comes from Japanese grading. I’ve been confused by that, too.
- swiftcoder 7mo agoIt's very common in Japanese-developed video games as well
- bitexploder 7mo agoCommon in gaming culture. Kind of a meme template. S tier is the best tier of something. People make tier lists of all sorts of things with that grading.
- tcbrah 7mo agotbh i stopped caring about "can i run X locally" a while ago. for anything where quality matters (scripting, code, complex reasoning) the local models are just not there yet compared to API. where local shines is specific narrow tasks - TTS, embeddings, whisper for STT, stuff like that. trying to run a 70b model at 3 tok/s on your gaming GPU when you could just hit an API for like $0.002/req feels like a weird flex IMO
- itigges22 7mo ago[flagged]
- deleted 7mo ago[deleted]
- hatthew 7mo agoFor me and probably many other people, local has nothing to do with cost and everything to do with privacy
- tcbrah 7mo agogenuine question - what are you working on that needs that level of privacy? outside of NSFW stuff most API providers arent doing anything with your prompts
- hatthew 7mo agoI would answer that, but it's private :) I can think of several reasons: corporate policy, personal principles, NSFW stuff, illegal stuff
- sdingi 7mo agoWhen running models on my phone - either through the web browser or via an app - is there any chance it uses the phone's NPU, or will these be GPU only? I don't really understand how the interface to the NPU chip looks from the perspective of a non-system caller, if it exists at all. This is a Samsung device but I am wondering about the general principle.
- amelius 7mo agoIt would be great if something like this was built into ollama, so you could easily list available models based on your current hardware setup, from the CLI.
- rootusrootus 7mo agoSomeone linked to llmfit. That would be a great tool to integrate with ollama. Just highlight the one you want and tell it to install. Quick, someone go vibe code that.
- dugidugout 7mo agoThe latest level of abstraction! You just release your ideas half baked in some internet connected box and wake up with products! Yahoo! Onwards into the Gestell!
- rootusrootus 7mo agoOkay, now I’m tempted to set up a bluesky account that takes requests and spits out working software. I’m certain this has already been done. It’s too obvious, and too hilarious.
- amelius 7mo agoYou can literally mine the internet for ideas, use an LLM to determine their potential and implement them, and make $$$.
- am17an 7mo agoYou can still run larger MoE models using expert weight off-loading to the CPU for token generation. They are by and large useable, I get ~50 toks/second on a kimi linear 48B (3B active) model on a potato PC + a 3090
- brcmthrowaway 7mo agoIf anyone hasn't tried Qwen3.5 on Apple Silicon, I highly suggest you to! Claude level performance on local hardware. If the Qwen team didn't get fired, I would be bullish on Local LLM.
- golem14 7mo agoHas anyone actually built anything with this tool? The website says that code export is not working yet. That’s a very strange way to advertise yourself.
- cafed00d 7mo agoOpen with multiple browsers (safari vs chrome) to get more "accurate + glanceable" rankings. Its using WebGPU as a proxy to estimate system resource. Chrome tends to leverage as much resources (Compute + Memory) as the OS makes available. Safari tends to be more efficient. Maybe this was obvious to everyone else. But its worth re-iterating for those of us skimmers of HN :)
- ryandrake 7mo agoMissing RTX A4000 20GB from the GPU list.
- mark_l_watson 7mo agoI have spent a HUGE amount of time the last two years experimenting with local models. A few lessons learned: 1. small models like the new qwen3.5:9b can be fantastic for local tool use, information extraction, and many other embedded applications. 2. For coding tools, just use Google Antigravity and gemini-cli, or, Anthropic Claude, or... Now to be clear, I have spent perhaps 100 hours in the last year configuring local models for coding using Emacs, Claude Code (configured for local), etc. However, I am retired and this time was a lot of fun for me: lot's of efforts trying to maximize local only results. I don't recommend it for others. I do recommend getting very good at using embedded local models in small practical applications. Sweet spot.
- nine_k 7mo agoWhat kind of hardware did you use? I suppose that a 8GB gaming GPU and a Mac Pro with 512 GB unified RAM give quite different results, both formally being local.
- fzzzy 7mo agoA Mac Pro with 512 gb unified ram does not exist.
- nine_k 7mo agoMac Studio Ultra, my bad. The 512 GB option existed up until March 2026: https://macdailynews.com/2026/03/06/apple-drops-512gb-m3-ultra-mac-studio-option-ups-256gb-memory-upgrade-by-400/ https://macdailynews.com/2026/03/06/apple-drops-512gb-m3-ult...
- manmal 7mo agoWhat about running e.g. Qwen3.5 128B on a rented RTX Pro 6000?
- girvo 7mo agoIMO you’re better off using qwen3.5-plus through the model studio coding plan, but ymmv
- andy_ppp 7mo agoIs it correct that there's zero improvement in performance between M4 (+Pro/Max) and M5 (+Pro/Max) the data looks identical. Also the memory does not seem to improve performance on larger models when I thought it would have? Love the idea though! EDIT: Okay the whole thing is nonsense and just some rough guesswork or asking an LLM to estimate the values. You should have real data (I'm sure people here can help) and put ESTIMATE next to any of the combinations you are guessing.
- GeekyBear 7mo ago> Is it correct that there's zero improvement in performance between M4 (+Pro/Max) and M5 (+Pro/Max) Preliminary testing did not come to that conclusion. > Apple’s New M5 Max Changes the Local AI Story https://www.youtube.com/watch?v=XGe7ldwFLSE https://www.youtube.com/watch?v=XGe7ldwFLSE
- lostmsu 7mo agoFrom the video: 4.4k is "almost" 4x times 1.8k because 4.4k has "number 4" in the beginning, and the other one - number 1. For the lazy: that's less then 3x: 1.8 * 3 = 5.4
- andy_ppp 7mo agoIt’s not even the largest part, just prefill so I think maybe M5 Max is 30% faster overall. Still pretty good I think but the 4x nonsense is just marketing!
- mkagenius 7mo agoLiterally made the same app, 2 weeks back - https://news.ycombinator.com/item?id=47171499 https://news.ycombinator.com/item?id=47171499
- mongrelion 7mo agoWhat front-end framework did you use? I find the UI so visually appealing
- mkagenius 7mo agoThanks. I actually used Google AI Studio for this. Prompted with my color choices and let it do the rest, turned out pretty good.
- hatthew 7mo agoFWIW, while I find it appealing, I also strongly associate it with "vibe coded webapp of dubious quality," so personally I'm not gonna try to replicate it myself.
- prokajevo 7mo ago[dead]
- zitterbewegung 7mo agoThe M4 Ultra doesn't exist and there is more credible rumors for an M5 Ultra. I wouldn't put a projection like that without highlighting that this processor doesn't exist yet.
- rcarmo 7mo agoThis is kind of bogus since some of the S and A tier models are pretty useless for reasoning or tool calls and can’t run with any sizable system prompt… it seems to be solely based on tokens per second?
- polyterative 7mo agoawesome, needed this
- tristor 7mo agoThis does not seem accurate based on my recently received M5 Max 128GB MBP. I think there's some estimates/guesswork involved, and it's also discounting that you can move the memory divider on Unified Memory devices like Apple Silicon and AMD AI Max 395+.
- tencentshill 7mo agoMissing laptop versions of all these chips.
- mopierotti 7mo agoThis (+ llmfit) are great attempts, but I've been generally frustrated by how it feels so hard to find any sort of guidance about what I would expect to be the most straightforward/common question: "What is the highest-quality model that I can run on my hardware, with tok/s greater than <x>, and context limit greater than <y>" (My personal approach has just devolved into guess-and-check, which is time consuming.) When using TFA/llmfit, I am immediately skeptical because I already know that Qwen 3.5 27B Q6 @ 100k context works great on my machine, but it's buried behind relatively obsolete suggestions like the Qwen 2.5 series. I'm assuming this is because the tok/s is much higher, but I don't really get much marginal utility out of tok/s speeds beyond ~50 t/s, and there's no way to sort results by quality.
- J_Shelby_J 7mo agoIt’s a hard problem. I’ve been working on it for the better part of a year. Well, granted my project is trying to do this in a way that works across multiple devices and supports multiple models to find the best “quality” and the best allocation. And this puts an exponential over the project. But “quality” is the hard part. In this case I’m just choosing the largest quants.
- mopierotti 7mo agoSupporting all the various devices does sound quite challenging. I wouldn't expect a perfect single measurement of "quality" to exist, but it seems like it could be approximated enough to at least be directionally useful. (e.g. comparing subsequent releases of the same model family)
- downrightmike 7mo agoLLMs are just special purpose calculators, as opposed to normal calculators which just do numbers and MUST be accurate. There aren't very good ways of knowing what you want because the people making the models can't read your mind and have different goals
- comboy 7mo agoWhat is the $/Mtok that would make you choose your time vs savings of running stuff locally? Just to be clear, it may sound like a snarky comment but I'm really curious from you or others how do you see it. I mean there are some batches long running tasks where ignoring electricity it's kind of free but usually local generation is slower (and worse quality) and we all kind of want some stuff to get done. Or is it not about the cost at all, just about not pushing your data into the clouds.
- A7OM 7mo ago[flagged]
- JulianPembroke 7mo ago[flagged]
- A7OM 7mo ago[flagged]
- JulianPembroke 7mo ago[flagged]
- kennywinker 7mo agoAre you using 7/8b models for coding? I keep getting the impression from what i read that 8b is only good for autocomplete. Also, it seems like an 8b model will run on a $100 2nd hand gpu (e.g. an 8gb gtx 1050/1060/1070 kind of thing) - why would you need to quantize?
- SXX 7mo agoSorry if already been answered, but will there be a metric for latency aka time to first token? Since I considered buying M3 Ultra and feel like it the most often discussed regarding using Apple hardware for runninh local LLMs. Where speed might be okay, but prompt processing can take ages.
- teaearlgraycold 7mo agoWait for the M5 Ultra. It will get the 4x prompt processing speeds from the rest of the M5 product line. I hear rumors it will be released this year.
- tkfoss 7mo agoNice UI, but crap data, probably llm generated.
- anigbrowl 7mo agoUseful tool, although some of the dark grey text is dark that I had to squint to make it out against the background.
- lagrange77 7mo agoFinally! I've been waiting for something like this.
- mmaunder 7mo agoOP can you please make it not as dark and slightly larger. Super useful otherwise. Qwen 3.5 9B is going to get a lot of love out of this.
- ProllyInfamous 7mo agoI'm not usually one to whine, but agreed; additionally, add contrast to the modifiers (e.g. processor select). First thing I did when I visited was scale the website to 150% Super impressive comparisons, and correlates with my perception having three seperate generations of GPU (from your list pulldown). Thanks for including the "old AMD" Polaris chipsets, which are actually still much faster than lower-spec Apple silicon. I have Ollama3.1 on a VEGA64 and it really is twice as fast as an M2Pro... ---- For anybody that thinks installing a local LLM is complicated: it's not (so long as you have more than one computer, don't tinker on your primary workhorse). I am a blue collar electrician (admittedly: geeky); no more difficult than installing linux. I used an online LLM to help me install both =D
- ricardbejarano 7mo agoOP here, it's not mine though!
- aanet 7mo ago+1 The website is super useful. That theme though... low-contrast text on too-dark theme is, uh, barely readable for me.
- duskdozer 7mo agoHave to disagree in part at least. Text is pretty small which isn't good, but I'm glad to see it when sites don't succumb to the make-dark-mode-lighter trend.
- nozzlegear 7mo agoI can't see shit on this website lol. It'd be nice if they had a switch to toggle a light mode.
- reactordev 7mo agoThis shows no models work with my hardware but that’s furthest from the truth as I’m running Qwen3.5… This isn’t nearly complete.
- kennywinker 7mo agoWell… don’t keep us guessing -what hardware? And which size qwen3.5?
- azmenak 7mo agoFrom my personal testing, running various agentic tasks with a bunch of tool calls on an M4 Max 128GB, I've found that running quantized versions of larger models to produce the best results which this site completely ignores. Currently, Nemotron 3 Super using Unsloth's UD Q4_K_XL quant is running nearly everything I do locally (replacing Qwen3.5 122b)
- DrAwdeOccarim 7mo agoTotally doing this today! Have you tried OpenJarvis or NemoClaw (is it out yet?). I want to use my computer “through” the LLM.
- Yanko_11 7mo ago[dead]
- bearjaws 7mo agoSo many people have vibe coded these websites, they are posted to Reddit near daily.
- kuon 7mo agoI have amd 9700 and it is not listed while it is great llm hardware because it has 32Gb for a reasonable price. I tried doing "custom" but it didn't seem to work. The tool is very nice though.
- ipunchghosts 7mo agoWhat is S? Also, NVIDIA RTX 4500 Ada is missing.
- fraywing 7mo agoThis is amazing. Still waiting for the "Medusa" class AMD chips to build my own AI machine.
- kpw94 7mo agoPeople complaining about how hard to get simple answer is don't appreciate the complexity in figuring out optimal models... There's so many knobs to tweak, it's a non trivial problem - Average/median length of your Prompts - prompt eval speed (tok/s) - token generation speed (tok/s) - Image/media encoding speed for vision tasks - Total amount of RAM - Max bandwidth of ram (ddr4, ddr5, etc.?) - Total amount of VRAM - "-ngl" (amount of layers offloaded to GPU) - Context size needed (you may need sub 16k for OCR tasks for instance) - Size of billion parameters - Size of active billion parameters for MoE - Acceptable level of Perplexity for your use case(s) - How aggressive Quantization you're willing to accept (to maintain low enough perplexity) - even finer grain knobs: temperature, penalties etc. Also, Tok/s as a metric isn't enough then because there's: - thinking vs non-thinking: which mode do you need? - models that are much more "chatty" than others in the same area (i remember testing few models that max out my modest desktop specs, qwen 2.5 non-thinking was so much faster than equivalent ministral non-thinking even though they had equivalent tok/s... Qwen would respond to the point quickly) At the end, final questions are: are you satisfied with how long getting an answer took? and was the answer good enough? The same exercise with paid APIs exists too, obviously less knobs but depending on your use case, there's still differences between providers and models. You can abstract away a lot of the knobs , just add "are you satisfied with how much it cost" on top of the other 2 questions
- paxys 7mo agoI wish creators of local model inference tools (LM Studio, Ollama etc.) would release these numbers publicly, because you can be sure they are sitting on a large dataset of real-world performance.
- gopalv 7mo agoChrome runs Gemini Nano if you flip a few feature flags on [1]. The model is not great, but it was the "least amount of setup" LLM I could run on someone else's machine. Including structured output, but has a tiny context window I could use. [1] - https://notmysock.org/code/voice-gemini-prompt.html https://notmysock.org/code/voice-gemini-prompt.html
- vednig 7mo agoOur work at DoShare is a lot of this stuff we've been on it for 2 years
- sidchilling 7mo agoI have been trying to run Qwen Coder models (8B at 4bit) on my M3 Pro 18GB behind Ollama and connecting codex CLI to it. The tool usage seems practically zero, like it returns the tool call in text JSON and codex CLI doesn’t run the tool (just displays the tool call in text). Has anyone succeeded in doing something like this? What am I missing?
- MikeNotThePope 7mo agoI have the same hardware. Been curious about trying it with Opencode.
- mongrelion 7mo agoIt might be that the system prompt sent by codex is not optimal for that model. Try with open code and see if your results improve
- sand500 7mo agoHow does it have details for M4 ultra?
- starkeeper 7mo agoThis is awesome!!! Could you please add title="explanation" over each selected item at the top. For example, when I choose my video card the ram changes... I'm not sure if the RAM selection is GPU RAM? The GRAM was already listed with the graphics card. SO I choose 96GB which is my Main memory? And the GB/s I am assuming it's GPU -> CPU speed?
- pants2 7mo agoThis really highlights the impracticality of local models: My $3k Macbook can run `GPT-OSS 20B` at ~16 tok/s according to this guide. Or I can run `GPT-OSS 120B` (a 6X larger model) at 360 tok/s (30X faster) on Groq at $0.60/Mtok output tokens. To generate $3k worth of output tokens on my local Mac at that pricing it would have to run 10 years continuously without stopping. There's virtually no economic break-even to running local models, and no advantage in intelligence or speed. The only thing you really get is privacy and offline access.
- danny_codes 7mo agoA million tokens is like 5 minutes of inference for heavy coding use.
- girvo 7mo agoAt work I regularly hit my 7.5mil tokens per hour limit one of our tools has, and have to switch model of tool, and I’m not even really a remotely heavy user. I think people don’t realise how many tokens get burned with CoT and tool calls these days At 7.5mil per hour hard limit, 84 days to hit the grandparents $3k That said local models really are slow still, or fast enough and not that great
- reverius42 7mo agoThey already stated they can only generate 57,600 tokens per hour locally (expressed as 16 tokens per second). So that's the limiting factor here.
- xandrius 7mo agoYou're saying it as if privacy was worthless? Also not many people would consider the price of buying a macbook and put it strictly towards running a local model. Instead if you wanted to get a macbook anyway, you get to run local models for free on top. Very different story.
- pants2 7mo ago
- amdivia 7mo agoI found this to be inaccurate, I can run OSS GPT 120B (4 bit quant) on my 5090 and 64 ram system with around 40 t/s. Yet here the site claims it won't work
- Readerium 7mo agoQwen 3.5 4B is the goat then
- ThrowawayTestr 7mo agoFor image generation or even video generation, local models are totally feasible. I can generate a 5 second clip with wan 2.2 in about 30 minutes on my 3060 12G. Plus, I have full control on the loras used.
- dzink 7mo agoThis would be wonderful if it is accurate - instead of guesstimating, let people report their actual findings. I can confirm GLM 4.7 is possible on M1 Max and it can do nice comprehensive answers (albeit at 12 min an answer) locally. You can also easily do Mistral7B and OSS 20B and others. Structure it as a way to report accruals, similarly to Levels.xyz for salaries, instead of guestimating.
- torginus 7mo agoHuh, I never knew my browser just volunteers my exact hardware specs to any website without so much as even notifying me about it.
- Jaxan 7mo agoIt doesn’t really. The website thinks I’m on a iPhone 19 pro, although I’m actually on a iPhone SE 1st gen. So it’s off by roughly a decade.
- torginus 7mo agoMaybe that's one of Safari's numerous 'quirks' our frontend devs keep bitching about. Which in this case Im thankful that Apple isn't too keen on following standards like these.
- weikju 7mo ago> on a iPhone 19 pro I wish the website could tell us how life is like in 2027!
- tstrimple 7mo agoMine is radically off as well. Says I've got a GeForce 980 or equivalent with 4GB instead of a 5090. I'm guessing the detection only really works on Chromium based browsers.
- ebbi 7mo agoI thought that's how airlines do the whole trickery around having different pricing if you access the site from Windows or Mac...
- DanielHB 7mo agoThis stuff is used a lot in browser fingerprinting for tracking purposes. More privacy-focused browsers usually feed randomized info.
- hotsalad 7mo agoThe latest Librewolf prompted me to allow the site permission to make a WebGL context. That's what it used for hardware detection.
- comrade1234 7mo agoI can't tell at a glance what this page is showing, but I am curious about the licenses on the various models that let me run it locally and make money off it. Awhile ago only deepseek let you do that - not sure now.
- mind_heist 7mo agonice, this is an interesting idea. Can you elaborate on the licensing issue ? how do you get blocked for using the models commercially ?
- comrade1234 7mo agoJust read the license agreement. Last time I looked into this the only model I could run locally and do what I want was deepseek. I think it was the MIT license. The others had various restrictions that just didn't make it worth it. I stopped researching this because buying the hardware to run deepseek full model just isn't practical right now. Our customers will have to be happy with us sending data to OpenAI/deepseek/etc if they want to use those features.
- singpolyma3 7mo agoqwen3.5 is just apache
- urba_ 7mo agoMan, I wonder when there will be AI server farms made from iCloud locked jailbroken iPhone 16s with backported MacOS
- Akuehne 7mo agoCan we get some of the ancient Nvidia Teslas, like the p40 added?
- dirk94018 7mo agoWe wrote the linuxtoaster inference engine, toasted, and are getting 400 prefill, 100 gen on a M4 Max w 128GB RAM on Qwen3-next-coder 6bit, 8bit runs too. KV caching means it feels snappy in chat mode. Local can work. For pro work, programming, I'd still prefer SOTA models, or GLM 4.7 via Cerebras.
- 0xbadcafebee 7mo agoCouple thoughts: - The t/s estimation per machine is off. Some of these models run generation at twice the speed listed (I just checked on a couple macs & an AMD laptop). I guess there's no way around that, but some sort of sliding scale might be better. - Ollama vs Llama.cpp vs others produce different results. I can run gpt-oss 20b with Ollama on a 16GB Mac, but it fails with "out of memory" with the latest llama.cpp (regardless of param tuning, using their mxfp4). Otoh, when llama.cpp does work, you can usually tweak it to be faster, if you learn the secret arts (like offloading only specific MoE tensors). So the t/s rating is even more subjective than just the hardware. - It's great that they list speed and size per-quant, but that needs to be a filter for the main list. It might be "16 t/s" at Q4, but if it's a small model you need higher quant (Q5/6/8) to not lose quality, so the advertised t/s should be one of those - Why is there an initial section which is all "performs poorly", and then "all models" below it shows a ton of models that perform well?
- TheCapn 7mo ago@OP are you the creator? Could you add my GPU to the list? Radeon VII https://www.amd.com/en/support/downloads/drivers.html/graphics/radeon-rx/radeon-rx-vega-series/amd-radeon-vii.html#amd_support_product_spec https://www.amd.com/en/support/downloads/drivers.html/graphi...
- ementally 7mo agoIn mobile section it is missing Tensor chips (used by Google Pixel devices).
- johneth 7mo agoRe: the design of the site. Please use higher contrast colours, especially the barely visible grey text on black background. It's annoying to try to read.
- nazbasho 7mo agoits perfect
- adamhsn 7mo agoCool project!! It would be useful to filter which model to use based on the objective or usage (i.e., for data extraction vs. coding). Also, just looking at VRAM kind of misses that a lot of CPU memory can be shared with the GPU via layer offloading. I think there is ultimately a need for a native client, like a CPU/GPU benchmark, to figure out how the model will actually perform more precisely.
- zahirbmirza 7mo agoThis was depressing. But, also, I can't figure why AI companies are valued so high. The models will reach a limit (ie for what most people want to use a model for), and compute will increase over time.
- zahirbmirza 7mo agoAlso, I have to add, this project is an excellent piece of work.
- hotsalad 7mo agoThis says I can't run anything, because it's missing some of the smallest models. I know that I can run Qwen3.5 up to 4B, Ministral 3B, Qwen3VL up to 4B, and I know there are some Gemmas and Llamas in my size range.
- winterismute 7mo agoOddly, the website lists "M4 Ultra" which however does not exist... Also, it does not account for Apple Silicon chips to have up to 512GB of memory in some cases, but that might be only a limitation of the gathered data.
- intrasight 7mo agoYour LLM visited the future
- remote3body 7mo agoThe 'spent 100 hours configuring' part hits home. That fragmentation is exactly why we started building Olares (https://github.com/beclab/Olares https://github.com/beclab/Olares). It’s basically an open-source OS layer that standardizes the local AI stack—Kubernetes (K3s) for orchestration, standardized model serving, and GPU scheduling. The goal is to stop fiddling with Python environments/drivers and just treat local agents like standardized containers. It runs on Mac Minis or dedicated hardware.
- storus 7mo agoMissing latest Nvidia cards like RTX Pro 6000; M3 Ultra can have at most 192GB selected etc.
- aplomb1026 7mo ago[dead]
- butILoveLife 7mo agoThis is borderline irresponsible. Conflating first tokens with all tokens is terrible. Apple looks far better than it actually is. Just ask any Apple user, they don't actually use local models.
- RagnarD 7mo agoI have an RTX 6000 Pro Max-Q, which has 96GB VRAM. It identified the hardware correctly but incorrectly thought it had 4GB, at least if I interpret the RAM dropdown correctly. Then it shows the full resolution models, which are completely unnecessary to run quality inference. Quantized models are routine for local inference and it should realize that. Needs work.
- 3Sophons 7mo agoa lighter-weight alternative of docker and python is the Rust+Wasm stack https://github.com/LlamaEdge/LlamaEdge https://github.com/LlamaEdge/LlamaEdge
- deleted 7mo ago[deleted]
- dxxvi 7mo agoNot sure if there's anybody like me. I use AI for only 2 purposes: to replace Google Search to learn something and to generate images. I wonder where there are not many models that do only 1 thing and do it well. For example, there's this one https://huggingface.co/Fortytwo-Network/Strand-Rust-Coder-14B-v1 https://huggingface.co/Fortytwo-Network/Strand-Rust-Coder-14... for Rust coding. I haven't used it yet, so don't know how it's compared to the free models that Kilo Code provides.
- dyauspitr 7mo agoFor learning and general searching I find ChatGPT to be the best. Nano Banana Pro for anything image and video related. Grok Imagine for pretty decent porn generation.
- never_inline 7mo agoDidn't want to hear the grok thing from handle named "dyauspitr". My day is ruined.
- dyauspitr 7mo agoI would imagine the Dyaus was pretty randy if Zeus is anything to go by.
- deleted 7mo ago[deleted]
- casey2 7mo agoSomething notable is that Qwen3.5:0.8B does better on benchmarks than GPT3.5. Runs much faster on local hardware than GPT3.5 on release. However Qwen3.5:0.8B dumber and slower than GPT3.5. It's dumber: it can do 3*3, but if asked to explain it in terms of the definition (i.e. 3+3+3=9) it fails. It's slower: It's a thinking model so your 900T/S are mainly spent "thinking" most of the time it will just repeat until it hangs. It pretty obvious that this reasoning scaling is a mirage, parameters are all you need. Everything else is mostly just wasting time while hardware get better.
- d0100 7mo agoWhy is there no RTX 5060ti?
- scorpioxy 7mo agoBesides trying to run on your own hardware, anybody have recommendations for running some decent models on one of the many "AI clouds" providers? This is for sporadic use and so maybe one of the "serverless" providers that bill by the hour or minute or similar as opposed to monthly renting GPUs. There are quite a few of them but their marketing is just confusing and full of buzz words. I've been tinkering with OpenRouter that acts as a middleman.
- metrix 7mo agouse openrouter, and call it a day. auto switching between providers, connectivity to all clouds and even works with free models
- scorpioxy 7mo agoYeah, that's what I've been doing. But in terms of privacy policies, I have to review(and trust) 2 providers instead of 1. OpenRouter and whatever provider is used for any particular model. I agree with you that it is more convenient though.
- ActorNightly 7mo agoI mean AWS bedrock fits your use case pretty much. They have a bunch of models that are serverless that you can use on a per token pricing cost. Gemini api use also comes with a free tier.
- scorpioxy 7mo agoThanks, I'll check out Bedrock. I was under the impression they only provide "enterprise" access as OpenRouter uses them as one of the providers but I didn't actually check. Looking at their docs now, seems I was wrong.
- raiph_ai 7mo agoGreat site, I have an M2 and M3pro and was thinking about getting and Ultra M4 and wanted to know if it was going to be worth it. Now I can see exactly what models I can run locally.
- starkparker 7mo agoEvery time I refresh the page, I get a higher tokens/second value, presumably because of the keying off memory bandwidth.
- rahimnathwani 7mo agoThis site presents models in an incomplete and misleading way. When I visit the site with an Apple M1 Max with 32GB RAM, the first model that's listed is Llama 3.1 8B, which is listed as needing 4.1GB RAM. But the weights for Llama 3.1 8B are over 16GB. You can see that here in the official HF repo: https://huggingface.co/meta-llama/Llama-3.1-8B/tree/main https://huggingface.co/meta-llama/Llama-3.1-8B/tree/main The model this site calls 'Llama 3.1 8B' is actually a 4-bit quantized version ( Q4_K_M) available on ollama.com/library: https://ollama.com/library/llama3.1:8b https://ollama.com/library/llama3.1:8b If you're going to recommend a model to someone based on their hardware, you have to recommend not only a specific model, but a specific version of that model (either the original, or some specific quantized version). This matters because different quantized versions of the model will have different RAM requirements and different performance characteristics. Another thing I don't like is that the model names are sometimes misleading. For example, there's a model with the name 'DeepSeek R1 1.5B'. There's only one architecture for DeepSeek R1, and it has 671B parameters. The model they call 'DeepSeek R1 1.5B' does not use that architecture. It's a qwen2 1.5B model that's been finetuned on DeepSeek R1's outputs. (And it's a Q4_K_M quantized version.)
- zargon 7mo agoThey appear to be using Ollama as a data source. Ollama does that sort of thing regularly.
- Decabytes 7mo agoDoes anyone use the super tiny models for anything ? Like in the 2billion or lower parameter level?
- genpfault 7mo agoSpeculative decoding[1]? [1]: https://github.com/ggml-org/llama.cpp/blob/master/docs/speculative.md#draft-model-draft https://github.com/ggml-org/llama.cpp/blob/master/docs/specu...
- sriramgonella 7mo ago[dead]
- manlymuppet 7mo agoWould be useful if comparable scores for performance are added, perhaps from arena.ai or ARC. I know scores can be imperfect, but it would be nice to be able to easily see what the best model your machine can handle is.
- eichin 7mo agoI'm surprised that this shows anything running usefully on my 2021-era thinkpad (with "Iris Xe"'TigerLake graphics) which inspires me to ask - are external GPUs useful for this sort of thing?
- markdown 7mo agoProtip: Website requires CMD+ a few times to increase font size to 200%.
- spidrahedron 7mo agoThis is super useful for people not having access to GPUs and Servers
- suheilaaita 7mo agoThe simplest way to really start, use anything like claude code, vs code, cursor, antigratvity, (or any other IDE) ask them to install ollama and pull the latest solid local model that was released that you can run based on your computer specs. Wait 5-10 minutes, and should be done. It genuinely is that simple. You can even use local models using claude code or codex infrastrucutre (MASSIVE UNLOCK), but you need solid GPU(s) to run decent models. So that's the downside.
- cloogshicer 7mo agoGenuine question, will this actually give you the latest solid local model? I would've thought no, because of the knowledge cutoff in whatever model you use to download it.
- suheilaaita 7mo agoI think it will give you a good "starter model". But then, it ultimately depends on what you want to do with the model exactly and your computer's specs. For example, I needed a local model to review some transactions and output structured output in .json format. Not all local models are necesserily good at structured outputs, so I asked grok (becuase it has solid web search and is up to date), what are the best recommended models given this use case and my laptop's specs. It suggested a few models, I chose one and went for it and now it's working. To summarise, - find model given use case and specs. - trial and error - test other models (if needed) - rinse repeat - because models are always coming out and getting better
- jiggunjer 7mo agoWhat knowledge cutoff? They all have web agents to Google it.
- suheilaaita 7mo agoThey all do, true. But some are better than the others in how they retrieve, digest and present you with the information. Boils down to personal preferences and experimenting.
- Stronz 7mo agoOne thing I noticed with local models is that conversation behavior tends to drift over time as well. Even when running locally, the model often starts structured but gradually becomes more verbose or explanatory in longer threads. Curious if others have seen similar behavior when using local setups.
- stared 7mo agoI would love to see on this list some (any) benchmark. "I can run a model" is mildly interesting. I can run OSS-20B on my M1 Pro. It works, I tried it, just I don't find any application.
- dale_glass 7mo agoIt's missing Ryzen AI MAX+, which is sort of the Apple Silicon equivalent.
- modernerd 7mo agoWould love it more if it could help me to answer: - Which models in the list are the best for my selected task? (If you don't track these things regularly, the list is a little overwhelming.) Sorting by various benchmark scores might be useful? - How much more system resources do I need to run the models currently listed at F, D or C at B, A, or S-tier levels? (Perhaps if you hover over the score, it could tell you?)
- rando1234 7mo agoOn a related question, I'm in the market to buy a new laptop for development and want to get something with good support for local models. What is a good recommendation in terms of GPU support etc? I currently have a Dell XPS 13. Should I just get a MacBook? Or are there good non-Mac options?
- olivercoleai 7mo ago[flagged]
- StefanoC 7mo agoCan anybody share their setup using 64GB macs? I have an M2 Ultra studio and I'm trying Qwen 3.5 MLX models hosting them from the CLI, but I'm a bit stuck picking bigger models, more context, 4/8 bits, Opus-Reasoning-Distilled, coder... There are a bit too many permutations between mlx CLI flags, env variables, and models. At the moment I'm exploring: - nightmedia/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-qx64-hi-mlx - BeastCode/Qwen3.5-27B-Claude-4.6-Opus-Distilled-MLX-4bit - mlx-community/Qwen3-Coder-Next-4bit
- 5o1ecist 7mo ago[dead]
- lrpe 7mo agoThis site desperately needs a light mode.
- GTP 7mo agoIf I'm sorting models by score (the default), which kind of score is it?
- dexterlagan 7mo agoSame as top comment, have spent a lot of time on local models. IMHO, qwen3.5 is the very first model that is actually usable for serious work, ever - and I've tried them all. The 35B 3B is very smart. It understands things no other local model I've ever used does, it's that good. The 9B runs on my slow Mac, and it's also very 'smart'. I can say with confidence that 2026 is the year of the local model, at last.
- Western0 7mo agoI tried using agents on a small Orange Pi and other small machines, and I’ve come to the conclusion that, unfortunately, it’s not feasible to do this in a practical way. Of course, I started writing an agent that could run in such an environment (long timeouts, retries, etc.), but it’s a real pain. For this to make sense, you need much more powerful hardware (a 5-year-old Mac mini is fine); the other issue is power consumption. Unfortunately, until you can run Mavericks 2 on your own hardware, it’s pretty expensive.
- Western0 7mo agoI like a old gemma 3n 4EB for small device
- rurban 7mo agoNo, not yet. Up to a single H100 there is no single local model, which doesn't make your code worse. (excluding trivial stuff like ruby, python, typescript). Implement features, fix bugs. Right now we started experimenting with 2 H100's, 160GB models. But even a single one is wide out of anyone others league.
- Bydgoszczo 7mo agook, what buy a hardware for local ai for agent (coding, and other) in 2026
- opengrass 7mo agoIt would be more broad if you entered a Passmark score. Us CPU/integrated suckers can only see if we beat a RPI4, and same goes for GPUs.
- coinexpert 7mo agoMost of the friction around local AI comes from juggling different runtimes for different providers. We built Milady specifically to solve that — one unified runtime that works with Ollama, OpenAI, Anthropic, and others. Switch providers without rewriting a line of code. Fully offline capable, zero telemetry. Happy to answer questions if anyone's curious: milady.ai
- vincentbusch 7mo agoInteresting point about #2. I've been doing something similar but from a different angle — running the same question through Claude, GPT-4o and Gemini to see where they disagree. Turns out they give completely different root causes about 30% of the time, which honestly surprised me. What's your experience with qwen3.5 for debugging tasks? I've mostly stuck with the big models so far.
- aimemobe 6mo ago[flagged]
- NathanielLucas 6mo ago[flagged]