5 ms·
> My day job is primarily ML as well, so I might just go for the 3090. 24 GB of memory is a game changer for what I can do locally. I really just wish Nvidia wo
by throwaway9d0291 6y ago
> My day job is primarily ML as well, so I might just go for the 3090. 24 GB of memory is a game changer for what I can do locally. I really just wish Nvidia would get its shit together with Linux drivers.
You might be better off waiting for Quadro.
GeForce, from NVIDIA's perspective, is the consumer line. The drivers are focused on getting the best possible performance for gamers (on Windows you'll see a big deal made of "Game-ready" drivers) and the hardware is intended for desktop usage, i.e. a couple of hours of gaming. GeForce hardware isn't intended for long-running jobs like ML and the drivers aren't stability-focused.
That's not to say that you can't train ML models on a GeForce card or that the GeForce drivers will lead to constant crashing or anything, just that NVIDIA isn't focused on this for GeForce.
Quadro and Tesla on the other hand are all about reliability. They're intended for use in servers and workstations where they're subject to all-day (Quadro) or 24/7 (Tesla) load. The drivers are focused on making sure things don't break.
Though I think the 3090 is intended to be the successor to the previous-generation Titan, so it could be that NVIDIA has different ideas about what goes where these days.
- fock 6y agoahem yes. It seems like your overlord won't pay you for this into the future.
- dannyw 6y agoThere are indeed driver differences but I have been running long-lived, sometimes month-lived ML workloads on GTX and RTX cards for years. I've never had any stability issues. You are more or less buying into their marketing, and paying 100-250% price premiums for a "Quadro" or "Tesla".
- formerly_proven 6y ago> the hardware is intended for desktop usage, i.e. a couple of hours of gaming. Tons of people have been flooring previous generation of cards 24/7 for months with compute workloads, and not nearly all of them were using watercooling. Those cards are still fine. ECC memory might be an argument though, for CAD/CAE type of work. Doubt ML cares about a few bitflips.
- ericd 6y agoTons of researchers and companies use the GeForce line for ML. It’s the reason for the infamous “no use in data centers” addition to the EULA, and potentially part of the reason for the switch from blowers, which made it feasible to stuff 4-16 1080Ti’s in a single case.