9 ms·
Pretty interesting watching their tech explainers on YouTube about the changes in their AI solutions. Apparently they switched from CNNs to transformers for ups
by blixt 2y ago
Pretty interesting watching their tech explainers on YouTube about the changes in their AI solutions. Apparently they switched from CNNs to transformers for upscaling (with ray tracing support) if I understood correctly though for frame generation makes even more sense to me.
32 GB VRAM on the highest end GPU seems almost small after running LLMs with 128 GB RAM on the M3 Max, but the speed will most likely more than make up for it. I do wonder when we’ll see bigger jumps in VRAM though, now that the need for running multiple AI models at once seems like a realistic use case (their tech explainers also mentions they already do this for games).
- bick_nyers 2y agoCheck out their project digits announcement, 128GB unified memory with infiniband capabilities for $3k. For more of the fast VRAM you would be in Quadro territory.
- terhechte 2y agoIf you have 128gb ram, try running MoE models, they're a far better fit for Apple's hardware because they trade memory for inference performance. using something like Wizard2 8x22b requires a huge amount of memory to host the 176b model, but only one 22b slice has to be active at a time so you get the token speed of a 22b model.
- FuriouslyAdrift 2y agoProject Digits... https://www.nvidia.com/en-us/project-digits/ https://www.nvidia.com/en-us/project-digits/
- throwaway48476 2y agoI guess they're tired of people buying macs for AI.
- cma 2y agoYou can also run the experts on separate machines with low bandwidth networking or even the internet (token rate limited by RTT)
- logankeenan 2y agoDo you have any recommendations on models to try?
- stkdump 2y agoMixtral and Deepseek use MOE. Most others don't.
- Terretta 2y agoMixtral 8x22b https://mistral.ai/news/mixtral-8x22b/ https://mistral.ai/news/mixtral-8x22b/
- terhechte 2y agoIn addition to the ones listed by others, WizardLM2 8x22b (was never officially released by Microsoft but is available).
- memhole 2y agoI planted garlic this year. Thanks for documenting! I can’t wait to see what I get harvest time. I like the Llama models personally. Meta aside. Qwen is fairly popular too. There’s a number of flavors you can try out. Ollama is a good starting point to try things quickly. You’re def going to have to tolerate things crashing or not working imo before you understand what your hardware can handle.
- memhole 2y agoI haven’t had great luck with the wizard as a counter point. The token generation is unbearably slow. I might have been using too large of a context window, though. It’s an interesting model for sure. I remember the output being decent. I think it’s already surpassed by other models like Qwen.
- terhechte 2y agoLong context windows are a problem. I gave Qwen 2.5 70b a ~115k context and it took ~20min for the answer to finish. The upside of MoE models vs 70b+ models is that they have much more world knowledge.
- ActionHank 2y agoThey are intentionally keeping the VRAM small on these cards to force people to buy their larger, more expensive offerings.
- Havoc 2y agoSaw someone else point out that potentially the culprit here isn’t nvidia but memory makers. It’s still 2gb per chip and has been since forever
- tharmas 2y agoGDDR7 apparently has the capability of 3gb per chip. As it becomes more available their could be more VRAM configurations. Some speculate maybe an RTX 5080 Super 24gb release next year. Wishful thinking perhaps.
- tbolt 2y agoMaybe, but if they strapped these with 64gb+ wouldn’t that be wasted on folks buying it for its intended purpose? Gaming. Though the “intended use” is changing and has been for a bit now.
- whywhywhywhy 2y agoXX90 is only half a gaming card it's also the one the entire creative professional 3D CGI, AI, game dev industry runs on.
- knowitnone 2y agohmmm, maybe they can had different offerings like 16GB, 32GB, 64GB, etc. Maybe we can even have 4 wheels on a car.
- mox1 2y agoNot really, the more textures you can put into memory the faster they can do their thing. PC gamers would say that a modern mid-range card (1440p card) should really have 16GB of vram. So a 5060 or even a 5070 with less than that amount is kind of silly.
- resource_waste 2y ago> after running LLMs with 128 GB RAM on the M3 Max, These are monumentally different. You cannot use your computer as an LLM. Its more novelty. I'm not even sure why people mention these things. Its possible, but no one actually does this out of testing purposes. It falsely equates Nivida GPUs with Apple CPUs. The winner is Apple.
- vonneumannstan 2y agoIf you want to run LLMs buy their H100/GB100/etc grade cards. There should be no expectation that consumer grade gaming cards will be optimal for ML use.
- throwaway314155 2y ago> There should be no expectation that consumer grade gaming cards will be optimal for ML use. And yet it just so happens they work effectively the same. I've done research on an RTX 2070 with just 8 GB VRAM. That card consistently met or got close to the performance of a V100 albeit with less vram. Why indicate people shouldn't use consumer cards? It's dramatically (like 10x-50x) cheaper. Is machine learning only for those who can afford 10k-50k USD workstation GPU's? That's lame and frankly comes across as gate keeping. Honestly I can't really imagine how a person could reasonably have this stance. Just let folks buy hardware and use it however they want. Sure if may be less than optimal but it's important to remember that not everyone in the world has the money to afford an H100. Perhaps you can explain some other better reason for why people shouldn't use consumer cards for ML? It's frankly kind of a rude suggestion in the absence of a better explanation.
- vonneumannstan 2y agoIf you can do research on a mid tier consumer card then more power to you. I'm specifically referencing the people who are complaining that the specs on consumer video game GPUs are not good for ML work. Like theres just no reasonable expectation that they will be.
- throwaway314155 2y agoAh, I see what you mean. Yeah I think it comes from a place of viewing increase in VRAM as relatively low cost and therefore an artificial limitation of sorts used to differentiate between consumer and workstation products (and the respective price disparities). Which may be true although there are more differences than just VRAM and I assume those market segments have different perceptions of the real value Gamers want it cheaper/faster, institutions want it closer to state of the art, more robust to lengthy workloads (as in year long training sessions), and better support from nvidia. Among other things.
- quadrature 2y agoWhy are transformers a better fit for frame generation. Is it because they can better utilize context from the previous history of frames ?