3 ms·
delete Hadoop so I've got more disk space to run Pandas. Except when you have ~2000 node cluster that runs 10,000 ETL tasks daily all of which are IO bound, An
by aub3bhat 9y ago
delete Hadoop so I've got more disk space to run Pandas.
Except when you have ~2000 node cluster that runs 10,000 ETL tasks daily all of which are IO bound, And you cannot "just" uninstall. However during certain periods the same cluster has significant underutilization this opens up possibilities of doing lots of cool stuff for almost zero cost.
I can understand your confusion. Training embedding model on Common Crawl is a toy problem. I recommend you thinl from perspective of a Tech company ideally in a production ML setting to understand the cost tradeoffs that go into making these decisions. Regarding memory locality if the problem is small enough its possible to tune Spark to use fewer or even just one worker with enough amount of memory allocated.
Sending the data over a network during training would be the worst thing you can do there.
When your data itself is in 100s of Terabytes and sharded across multiple racks, often a properly tuned spark pipeline is more reliable and performs well. Again there is a difference between using Common Crawl subset that you manually download filter train etc and in ensuring that models are updated/trained automatically daily across 100s of TB data.
- deleted 9y ago[deleted]
- rspeer 9y agoWell, thanks for telling me what I do, that was very informative, except for the fact that it's all nonsense. You are posturing.