7 ms·
It's awesome that stuff like this is open source, but even if you have a basement rig with 4 NVIDIA GeForce RTX 5090 graphic cards ($15-20k machine), can it eve
by TIPSIO 10mo ago
It's awesome that stuff like this is open source, but even if you have a basement rig with 4 NVIDIA GeForce RTX 5090 graphic cards ($15-20k machine), can it even run with any reasonable context window that isn't like a crawling 10/tps?
Frontier models are far exceeding even the most hardcore consumer hobbyist requirements. This is even further
- bigyabai 10mo agoPeople with basement rigs generally aren't the target audience for these gigantic models. You'd get much better results out of an MoE model like Qwen3's A3B/A22B weights, if you're running a homelab setup.
- Spivak 10mo agoYeah I think the advantage of OSS models is that you can get your pick of providers and aren't locked into just Anthropic or just OpenAI.
- hnfong 10mo agoReproducibility of results are also important in some cases. There are consumer-ish hardware that can run large models like DeepSeek 3.x slowly. If you're using LLMs for a specific purpose that is well-served by a particular model, you don't want to risk AI companies deprecating it in a couple months and push you to a newer model (that may or may not work better in your situation). And even if the AI service providers nominally use the same model, you might have cases where reproducibility requires you use the same inference software or even hardware to maintain high reproducibility of the results. If you're just using OpenAI or Anthropic you just don't get that level of control.
- Aachen 10mo agoWho is the target audience of these free releases? I don't mind free and open information sharing but I have wondered what's in it for the people that spent unholy amounts of energy on scraping, developing, and training
- noosphr 10mo agoHome rigs like that are no longer cost effective. You're better off buying an rtx pro 6000 outright. This holds both for the sticker price, the supporting hardware price, the electricity cost to run it and cooling the room that you use it in.
- torginus 10mo agoI was just watching this video about a Chinese piece of industrial equipment, designed for replacing BGA chips such as flash or RAM with a good deal of precision: https://www.youtube.com/watch?v=zwHqO1mnMsA https://www.youtube.com/watch?v=zwHqO1mnMsA I wonder how well the aftermarket memory surgery business on consumer GPUs is doing.
- ThrowawayTestr 10mo agoLTT recently did a video on upgrading a 5090 to 96gb of ram
- dotancohen 10mo agoI wonder how well the opthalmologist is doing. These guys are going to be paying him a visit playing around with those lasers and no PPE.
- CamperBob2 10mo agoEh, I don't see the risk, no pun intended. It's not collimated, and it's not going to be in focus anywhere but on-target. It's also probably in the long-wave range >>1000 nm that's not focused by the eye. At the end of the day it's no different from any other source of spot heating. I get more nervous around some of the LED flashlights you can buy these days. I want one. Hot air blows.
- noosphr 10mo agoIt's 45w of lasing power. I have a scar on my hand that's 15 years old from running one of those at 10% power and getting a reflection from a bare metal sheet. This will absolutely scar, if not char, your cornea faster than you can blink.
- halyconWays 10mo agoAs someone with a basement rig of 6x 3090s, not really. It's quite slow, as with that many params (685B) it's offloading basically all of it into system RAM. I limit myself to models with <144B params, then it's quite an enjoyable experience. GLM 4.5 Air has been great in particular
- lostmsu 10mo agoDid you find it better than GPT-OSS 120B? The public rankings are contradictory.
- halyconWays 10mo agoI haven't used GPT-OSS 120B, or other GPT-OSS models, and I mostly go on personal recommendations rather than benchmarks directly.
- tarruda 10mo agoYou can run at ~20 tokens/second on a 512GB Mac Studio M3 Ultra: https://youtu.be/ufXZI6aqOU8?si=YGowQ3cSzHDpgv4z&t=197 https://youtu.be/ufXZI6aqOU8?si=YGowQ3cSzHDpgv4z&t=197 IIRC the 512GB mac studio is about $10k
- hasperdi 10mo agoand can be faster if you can get an MOE model of that
- dormento 10mo ago"Mixture-of-experts", AKA "running several small models and activating only a few at a time". Thanks for introducing me to that concept. Fascinating. (commentary: things are really moving too fast for the layperson to keep up)
- whimsicalism 10mo agothat's not really a good summary of what MoEs are. you can more consider it like sublayers that get routed through (like how the brain only lights up certain pathways) rather than actual separate models.
- _lyxd 10mo agoThe gains from MoE is that you can have a large model that's efficient, it lets you decouple #params and computation cost. I don't see how anthropomorphizing MoE <-> brain affords insight deeper than 'less activity means less energy used'. These are totally different systems, IMO this shallow comparison muddies the water and does a disservice to each field of study. There's been loads of research showing there's redundancy in MoE models, ie cerebras has a paper[1] where they selectively prune half the experts with minimal loss across domains -- I'm not sure you could disable half the brain and notice a stupefying difference. [1] https://www.cerebras.ai/blog/reap https://www.cerebras.ai/blog/reap
- 10mo ago
- reilly3000 10mo agoThere are plenty of 3rd party and big cloud options to run these models by the hour or token. Big models really only work in that context, and that’s ok. Or you can get yourself an H100 rack and go nuts, but there is little downside to using a cloud provider on a per-token basis.
- cubefox 10mo ago> There are plenty of 3rd party and big cloud options to run these models by the hour or token. Which ones? I wanted to try a large base model for automated literature (fine-tuned models are a lot worse at it) but I couldn't find a provider which makes this easy.
- big_man_ting 10mo agohave you checked OpenRouter if they offer any providers who serve the model you need?
- cubefox 10mo agoI searched for "base" and the best available base model seems to be indeed Llama 3.1 405B Base at Hyperbolic.ai, as mentioned in the comment above.
- reilly3000 10mo agoIf you’re already using GCP, Vertex AI is pretty good. You can run lots of models on it: https://docs.cloud.google.com/vertex-ai/generative-ai/docs/model-garden/available-models https://docs.cloud.google.com/vertex-ai/generative-ai/docs/m... Lambda.ai used to offer per-token pricing but they have moved up market. You can still rent a B200 instance for sub $5/hr which is reasonable for experimenting with models. https://app.hyperbolic.ai/models https://app.hyperbolic.ai/models Hyperbolic offers both GPU hosting and token pricing for popular OSS models. It’s easy with token based options because usually are a drop-in replacement for OpenAI API endpoints. You have you rent a GPU instance if you want to run the latest or custom stuff, but if you just want to play around for a few hours it’s not unreasonable.
- potsandpans 10mo agoI run a bunch of smaller models on a 12gb vram 3060 and it's quite good. For larger open models ill use open router. I'm looking into on- demand instances with cloud/vps providers, but haven't explored the space too much. I feel like private cloud instances that run on demand is still in the spirit of consumer hobbyist. It's not as good as having it all local, but the bootstrapping cost plus electricity to run seems prohibitive. I'm really interested to see if there's a space for consumer TPUs that satisfy usecases like this.
- wickedsight 10mo agoWhich ones are your favorites that fit on the 3060?
- seanw265 10mo agoFWIW it looks like OpenRouter's two providers for this model (one of whom being Deepseek itself) are only running the model around 28tps at the moment. https://openrouter.ai/deepseek/deepseek-v3.2 https://openrouter.ai/deepseek/deepseek-v3.2 This only bolsters your point. Will be interesting to see if this changes as the model is adopted more widely.
- deleted 10mo ago[deleted]