4 ms·
also rules out e.g. the fairly common practice of using R or pandas for some ad-hoc processing. This is essentially the reason why all organizations are adopti
by aub3bhat 9y ago
also rules out e.g. the fairly common practice of using R or pandas for some ad-hoc processing.
This is essentially the reason why all organizations are adopting Spark, since it allows you to write imperative code on dataframes and build ML models.
- rspeer 9y agoNo, "all organizations" are not adopting Spark. You do not need Spark to use dataframes. And distributed computing is terrible for machine learning. Maybe you've worked at a job or two where nobody can comprehend not using distributed computing, as you describe, but it's nonsense to claim that "all organizations" work that way.
- aub3bhat 9y agoNo, "all organizations" are not adopting Spark. All organizations which already have a Hadoop cluster. You do not need Spark to use dataframes. Never claimed this. To clarify Spark allows you to directly port Pandas code while leveraging existing Hadoop cluster infrastructure. And distributed computing is terrible for machine learning. Distributed computing (Both traditional hadoop/spark and latest TF/PyTorch with parameter server) are essential for scaling ML beyond a certain point. Maybe you've worked at a job or two where nobody can comprehend not using distributed computing, as you describe, but it's nonsense to claim that "all organizations" work that way. If you have experience routinely training models on Terabytes of data intended for production deployment. I am happy to hear. There is a vast difference between training a model on your machine for research and building a reliable ML system that scales across large datasets and teams while taking infrastructure costs into account.
- rspeer 9y ago"All organizations which already have a Hadoop cluster are adopting Spark" is not a very interesting claim. Here's how I would leverage Hadoop infrastructure to use Pandas: delete Hadoop so I've got more disk space to run Pandas. I don't get to train ML on terabytes of data very often. I do NLP, so "terabytes" means training a background model on the entire Common Crawl. Usually I'm doing something more specific and interesting than learning about random web pages. But when I do deal with the Common Crawl, I deal with it on one computer. Terabytes are not scary. How does distributed computing even help? ML models need memory locality, sometimes to the extreme of being localized within a GPU's memory. And the limiting factor is the ability to iterate over the data. Sending the data over a network during training would be the worst thing you can do there.
- aub3bhat 9y agodelete Hadoop so I've got more disk space to run Pandas. Except when you have ~2000 node cluster that runs 10,000 ETL tasks daily all of which are IO bound, And you cannot "just" uninstall. However during certain periods the same cluster has significant underutilization this opens up possibilities of doing lots of cool stuff for almost zero cost. I can understand your confusion. Training embedding model on Common Crawl is a toy problem. I recommend you thinl from perspective of a Tech company ideally in a production ML setting to understand the cost tradeoffs that go into making these decisions. Regarding memory locality if the problem is small enough its possible to tune Spark to use fewer or even just one worker with enough amount of memory allocated. Sending the data over a network during training would be the worst thing you can do there. When your data itself is in 100s of Terabytes and sharded across multiple racks, often a properly tuned spark pipeline is more reliable and performs well. Again there is a difference between using Common Crawl subset that you manually download filter train etc and in ensuring that models are updated/trained automatically daily across 100s of TB data.
- deleted 9y ago[deleted]
- rspeer 9y agoWell, thanks for telling me what I do, that was very informative, except for the fact that it's all nonsense. You are posturing.