3 ms·
But most people don't have an RTX 5090 lying around, so the story doesn't apply to them, right?
by enraged_camel 2mo ago
But most people don't have an RTX 5090 lying around, so the story doesn't apply to them, right?
- OtomotO 2mo agoCorrect. If the premise doesn't hold, it's ex falso quodlibet for anyone.
- deleted 2mo ago[deleted]
- AbsurdCensor 2mo agoYou don't need a 5090 to run local AI. A whole lot of people out there are doing it with Macs. Unified ram is the biggest thing.
- irishcoffee 2mo agoBack in “the day” nerds just bought the hardware to fuck with. Some of us still do. Claiming that compute is the barrier to entry just means you’re not a nerd. That’s ok.
- Forgeties79 2mo agoI'm running Qwen 27B no problem with an AMD 9070XT + 24gb DDR5 ram. Does basic web search for me (tool call with tavily, costs nothing I get 1000 searches a month) and is great for creative writing (primarily breaking writer's block). Until the recent surge in ram costs, that wouldn't be hard to do. I built the computer for ~$1600 a year ago.
- NekkoDroid 2mo agoI am getting ~13-15 tps with my 9070XT for the 27B (~35tps for the 35B-A3B), but I think for me the main bottleneck is the 64gb of DDR4 3600 memory. What kinda speeds are you getting with what speed of DDR5?
- Forgeties79 2mo agoI’m a little more novice than a lot of the people on this site so take my response with a grain of salt. The wall I keep hitting is I can run models like I described (Q3-4 usually), but it’s very sensitive to context. Once I start getting past 7.5k or so it can really fall apart. Sometimes before that. It just really depends. If I run smaller ones that offload less to ram, they stay somewhat coherent but don’t quite do what I want them to do. Your token speeds are not that much slower than mine. I imagine part of it is I’m not fine-tuning it very well. On a good day I’ll get like…15tps.