3 ms·
Why bring Hadoop into it? MR is a paradigm for parallel programming, and having it compiled into OpenCL for you is hugely convenient. But the best Hadoop use ca
by alanctgardner2 13y ago
Why bring Hadoop into it? MR is a paradigm for parallel programming, and having it compiled into OpenCL for you is hugely convenient. But the best Hadoop use cases always leverage massive storage capacity: people who do ML on Hadoop are doing it because the data is already in Hadoop, not because it seemed like a nifty thing to do. Do you actually leverage HDFS and the Hadoop scheduling infrastructure to run ParallelX jobs? Or can you safely jettison them and just sell this as: "Write your ML/graph analysis/whatever code in Java, have it execute on a GPU"
Also, what are the limits like for input data sizes? I've done a little OpenCL, but I've never gone past the GPU RAM size.
- tonydiv 13y ago"Do you actually leverage HDFS and the Hadoop scheduling infrastructure to run ParallelX jobs?" -- Yes. We could also provide what you're suggesting outside of Hadoop, no problem! "What are the limits like for input data sizes?" --- There are no limits. You can store your data in AWS, and we will crunch it for you. That said, initially, for IO-bound and disk-bound jobs, ParallelX might not be ideal. This is a problem we are solving as we scale. Thanks for the feedback! We appreciate it!
- jbooth 13y agoAre you porting string processing routines to OpenCL? The grid model of GPU computing seems to be a bit of a mismatch with the jagged-edge model of a bunch of variably-sized records in 2 1TB files that I want to join on in my typical hadoop use-case. I guess I've only delved into GPU stuff for dense matrix math, though, which is a pretty bad fit with Hadoop. Maybe you guys can come up with some other use-cases for them.
- Hydraulix989 13y agoYes, we are porting string processing routines to OpenCL. We are going to have many of the Java libraries accessible within the compiled GPU code, including file system I/O.
- alanctgardner2 13y agoWhat's a typical input size, if you don't mind me asking?
- tonydiv 13y agoWhat do you mean 'typical?'
- espeed 13y agoYeah, it's unclear how the Hadoop cluster comes into play here. However, in general, a Spark cluster (http://spark.incubator.apache.org/ http://spark.incubator.apache.org/) will have better IO performance than Hadoop, and Spark code is much simpler than equivalent Hadoop code. Spark is a new computing framework out of Berkeley's AMPLab (https://amplab.cs.berkeley.edu/software/ https://amplab.cs.berkeley.edu/software/), and it might be an interesting platform to target. It's being adopted by Twitter, Yahoo, Amazon (http://www.wired.com/wiredenterprise/2013/06/yahoo-amazon-amplab-spark/all/ http://www.wired.com/wiredenterprise/2013/06/yahoo-amazon-am...), and it's now commercially backed by Databricks (http://databricks.com/ http://databricks.com/), which just recieved funding from Andreessen Horowitz. See https://news.ycombinator.com/item?id=6466222 https://news.ycombinator.com/item?id=6466222