4 ms·
In a laptop?
by dman 5y ago
In a laptop?
- forgetfulness 5y agoMaybe the commenter has a very interesting use for it, but why would you buy a 256 GB RAM machine (which isn't a lot of memory for this arena either) to develop the model, instead of using something smaller to work on it and leasing a big cluster for the minutes to hours of the day that you'll need to train it on the actual dataset?
- aldanor 5y agoKicking learns off on a cluster is surely a thing as well. And in some fields, as you correctly mentioned, memory requirements may be measured in terabytes. It's more of a 'production use case' though - what I meant is the 'dev use case'. For instance, playing with mid/high frequency market data, plotting various things, just quickly looking at the data and various metrics, trying and testing things out often requires up too 100gb of memory at disposal at any time. It's definitely not only about model training. And 'something smaller to work on' principle doesn't always work in these cases. If the whole thing fits in my local ram, I would of course prefer to work on it locally until I actually need cluster resources. (But seriously though... what is 16gb these days? Looking at process monitor, I think my firefox takes over 5gb now since I have over a thousand tabs in treestyletab, clion + pycharm take another 5gb, parallels vm for some work stuff is another 10gb; if doing any local data science work that's usually at least a dozen gb or a few dozen or more)
- nonameiguess 5y agoI don't do this kind of thing any more, but back when I did, the one thing that consistently bit me was exploratory analysis requiring one-hot encoding of categorical data where you might have thousands upon thousands of categories. Take something like the Walmart shopper segmentation challenge on Kaggle that a hobbyist might want to take a shot at. That's just exploratory analysis, not model training. Having to do that in the cloud would be quite annoying when your feedback loop involves updating plots that you would really like to have in-memory on a machine connected to your monitor. Granted, you can forward a Jupyter server from an EC2, but also the high-memory EC2s are extremely expensive for hobbyists, way more than just buying your own RAM if you're going to do it often.
- rubatuga 5y agoI think there are studies showing that one hot encodings are not as efficient as an embedding, so maybe you would want to reduce the dimensions before attempting the analysis.
- JonathanFly 5y ago>Maybe the commenter has a very interesting use for it, but why would you buy a 256 GB RAM machine (which isn't a lot of memory for this arena either) to develop the model, instead of using something smaller to work on it and leasing a big cluster for the minutes to hours of the day that you'll need to train it on the actual dataset? A 128GB of a ram in a consumer PC is about $750 dollars (desktop anyway, laptop may be more?). That's less than a single high end consumer gaming GPU. Or a fraction of a Quadro GPU. So to the extent that developers ever run things on their local hardware (CPU, GPU, whatever) 128GB of RAM is not much of a leap. Or 256GB for Threadripper. It's in the ballpark of having a high-end consumer GPU.
- efxhoy 5y agoI have 128GB in my desktop, which is the same amount of ram that our cluster compute nodes have. I use it for Python/pandas forecasting work. The extra ram in my machine means I can work on the same datasets we use in prod locally, which is a huge productivity boost. Most of our developer time is actually spent building datasets. loading a random sample doesn’t really work when doing time series or spatial transforms and using a time or space limited subset makes description (graphs, maps) a chore. More ram is a huge productivity enabler when working with in memory tools like pandas.
- llampx 5y agoThe same M chip is used in 4 product lines so I'm going to assume a Pro version of the iMac or Mac Mini is what the parent means, but if you need that much memory, setting up a VM should be worth it. Same if you need a GPU.