5 ms·
Deep Learning with Spark and TensorFlow
- hcrisp 11y agoImpressive, but it seems an inversion of paradigms. Small data to compute ratios is usually associated with high performance computing (HPC). Why use Spark when the data is small and is broadcast to each worker? You have to pay the serialization-deserialization penalties of moving the data from Python to JVM and back again. In fact the JVM isn't really needed here at all since all the computation is done in the pure-Python workers in an embarrassingly-parallel way. Seems to me that you would just move onto an HPC and use TensorFlow within a IPython.parallel paradigm and be done much sooner.
- rxin 11y agoThe "broadcast" is pretty cheap because often you already have the data in some distributed file system, or if on a single node the network bandwidth is pretty high. The problem with a lot of the deep learning workloads is that it is very compute intensive and as a result takes a long time to run. For example, it is not uncommon to take a week to train some models.
- gcr 11y agoDeep learning workloads are typically compute-intensive, but they also tend to be extremely I/O intensive, and convergence may depend on a synchronous step where all the nodes must finish making their contribution to the model before any of them can continue. (This may not be quite true though -- see Google's DistBelief paper--but most frameworks work this way). Often times, adding more machines to a cluster may make training proportionally slower.
- rxin 11y agoDid you actually read the article? It was using Spark to parallelize hyperparameter tuning, which is embarrassingly parallel.
- doobwa 11y agoWhy not just use GNU Parallel (or something similar) instead of Spark?
- elyase 11y agoI think this could have been done with GNU parallel. One advantage I see with Spark is that is that it is easier to interact with Python, for example these two lines are all is needed to call the relevant Python function: urls = sc.parallelize(batched_data) labelled_images = urls.flatMap(apply_batch) So if you already have a cluster with Spark installed (like Databrick does) then it takes less work to just call your Python code than setting up a GNU Parallel cluster and a writing a small wrapper script. Additionally a Python script would have to load/init the models on every call from Parallel. I agree that this is not a great demonstration of Spark main strengths.
- orm 11y agoI think one reason would fault tolerance. Is there a fault tolerance layer in GNU parallel? last time I checked their homepage ( a few minutes ago), there was no reference to fault tolerance. Another reason is, perhaps, scheduling.
- chimtim 11y agowhat fault tolerance does spark give you in this scheme? It cannot look into TF progress and checkpoint all state. Using Spark with TF, seems like an overkill -- you need to manage and install two framework what should ideally be a 200 line python wrapper or small mesos framework at most.
- ole_tange 11y agoDoes --retries count as fault tolerance?
- gcr 11y agoOh dear. You're right, sorry. Shouldn't have commented before actually reading the article...
- elcct 11y agoThat article reminded me of this: http://i.imgur.com/XQJ3ACO.jpg http://i.imgur.com/XQJ3ACO.jpg
- rxin 11y agoThe blog post actually provides code to reproduce all the steps and the chart. See http://go.databricks.com/hubfs/notebooks/TensorFlow/Distributed_processing_of_images_using_TensorFlow.html http://go.databricks.com/hubfs/notebooks/TensorFlow/Distribu... http://go.databricks.com/hubfs/notebooks/TensorFlow/Test_distributed_processing_of_images_using_TensorFlow.html http://go.databricks.com/hubfs/notebooks/TensorFlow/Test_dis... You might've missed the section "How do I use it?" Maybe we should've made that section more obvious.
- hellofunk 11y agoI am still laughing from that graphic. So simple, but, you know what? So true, too.
- obituary_latte 11y agohttp://i.imgur.com/boZRjbB.png http://i.imgur.com/boZRjbB.png
- mindcrime 11y agoThere's actually a little bit more info out there for would-be "Watson builders". https://www.ibm.com/developerworks/community/blogs/InsideSystemStorage/entry/ibm_watson_how_to_build_your_own_watson_jr_in_your_basement7 https://www.ibm.com/developerworks/community/blogs/InsideSys... http://www.theregister.co.uk/2011/02/21/ibm_watson_qa_system/ http://www.theregister.co.uk/2011/02/21/ibm_watson_qa_system... http://learning.acm.org/webinar/lally.cfm http://learning.acm.org/webinar/lally.cfm http://www.cs.nmsu.edu/ALP/2011/03/natural-language-processing-with-prolog-in-the-ibm-watson-system/ http://www.cs.nmsu.edu/ALP/2011/03/natural-language-processi... Of course, there's still a big gap between "Download some stuff" and "Build Watson", but at least there's a trickle of details on what happens in the "a miracle happens here" step. :-)
- vonnik 11y agoSo the cool thing here is that you can use Spark and TF to find the best model like Microsoft Research did with Resnets. http://www.wired.com/2016/01/microsoft-neural-net-shows-deep-learning-can-get-way-deeper/ http://www.wired.com/2016/01/microsoft-neural-net-shows-deep... They're showing you how to train different architectures simultaneously, and then compare their results in order to select the best one. That's great as far as it goes. The drawback is that with this schema, you can't actually train a given network faster, which is what you want to do with Spark. What is the role of a distributed run-time in training artificial neural networks? It's easy. NNs are computationally intensive, so you want to share the work over many machines. Spark can help you orchestrate that through data parallelism, parameter averaging and iterative reduce, which we do with Deeplearning4j. http://deeplearning4j.org/spark http://deeplearning4j.org/spark https://github.com/deeplearning4j/dl4j-spark-cdh5-examples https://github.com/deeplearning4j/dl4j-spark-cdh5-examples Data parallelism is an approach Google uses to train neural networks on tons of data quickly. The idea is that you shard your data to a lot of equivalent models, have each of the models train on a separate machine, and then average their parameters. That works, it's fast, and it's how Spark can help you do deep learning better.
- amelius 11y agoI have a question about neural networks. Say, you are training a NN to recognize handwritten characters 0 and 1, and you have 1000 training images for each character (so 2000 images in total). All images are bitmaps with 0 for black and 1 for white. Now, by accident, all the "0" training-images have an even number of black pixels, and all the "1" training-images have an odd number of black pixels. How do you know that the NN really learns to recognize 0's and 1's, as opposed to recognizing whether the number of pixels in an image is even or odd?
- muizelaar 11y agoYou don't necessarily know, but the unlikeliness of this occurring in the training images and the structure of the neural net making it difficult for the net to learn to even or oddness of the number of pixels makes one reasonably confident.
- mindcrime 11y agoI would say that if you're using a single layer NN, the answer is "you don't really know". And that gets to a point about how we still don't entirely understand how neural networks work, even when they do work. If you were using a deep network though, and if the current theory is correct, it would be a slightly different story. The current thinking, as I understand it, is that with deep networks, each layer learns representations of certain features (say "slashes", "edges", "right slanted lines", "left slanted lines", etc.) and the progressively higher layers learn representations composed from those more primitive features. So if a deep net were recognizing your handwritten characters, you could probably reason that it isn't just considering whether the number of black pixels is even or odd. Now in reality this is a pretty contrived, and probably unlikely scenario. But it's a valid question, because there's a deeper point to all of this, which involves transference of learning. That is, how do you take the learning done by a neural network - trained to do one thing - and then leverage that learning in another application. We still don't exactly know how to do that, and that's in part because we don't entirely understand the nature of the representations the networks build up. So a very good answer to your question would arguably help understand how to do transference, which would make NN's even more useful.
- tachim 11y ago0.1% accuracy increments correspond to 10 images in the testing set; they should be reporting standard error bars with those numbers.