6 ms·
Pro tip: justify why dividing data set size by time to solution and saying number is small is a good analysis. Let's try imagenet training. Intel's best time i
by twtw 8y ago
Pro tip: justify why dividing data set size by time to solution and saying number is small is a good analysis.
Let's try imagenet training. Intel's best time is 3h25m on 128 nodes. Imagenet is ~150 gb.
(150 GB)/(3 hr * 3600 sec/hr * 128 nodes) = less than one megabyte per second per node! Caffe is slow! CPUs are slow!
Or bs metrics are bad?
https://dawn.cs.stanford.edu/benchmark/ImageNet/train.html https://dawn.cs.stanford.edu/benchmark/ImageNet/train.html
- jmalicki 8y agoIntel's run wasn't 3:25 per epoch, which is the difference. A simple ETL should be one linear scan over the data, with maybe some merges, nothing I saw here seemed more complex than that?
- paulsutter 8y agoThis is ETL processing, not training a neural network. If their (undisclosed) actual task is not parallelizable in Spark, the comparison is even sillier so let's be generous and assume it is.
- felipe_aramburu 8y agoThe GPU workload is used in feeding XGBOOST to perform classification. We are omitting this timing on purpose because the training time on purpose becuase SPARK took 1000's of times longer for this part of the workflow. You can see the workflow for yourself and see how the task is quite parallelizable.
- scarejunba 8y agoSo just the ETL portion is that slow? That's really odd. I think I see faster performance than that ETLing on Hadoop. Unless there's something complex going on here.
- felipe_aramburu 8y agoDid you see the steps being performed in the workflow? How would you perform these steps on hadoop and what kind of times would you expect for the different steps in the workflow?
- scarejunba 8y agoI went back to see and only just realized you had a “TLDR here’s the full details” bit. Missed that and saw only the 5x faster ETL which I figured was what you were crowing about. I’ll have to look at the rest later. I still don’t get why the ETL phase has to be so slow but maybe it’ll be obvious when I look, like you’re doing extensive transformations or something. But no one would just call that “just ETL”. Anyway, I’ll see for sure later.
- felipe_aramburu 8y agoCacaw! I am glad the crowing reached you :). You probably did not get far if you are still on the old numbers that was in the first paragraph. ETL can be very extensive. For example, we first built this when we had to take data from 15 different database systems that represented individuals and their pension contributions and join across these systems. It was a largish join about ~4 tables from each system so around a 60 table join. That was "JUST ETL". The job was preparing the data for training. ETL is often times a large part of people's workflow. Looking for needle situations in a haystack. That can be JUST ETL. If you believe there is a more apt word for extracting data from a system, performing unspeakable transformations on it, then making that information available to another process then please tell me. Being of Peruvian stock myself I take great license with my language and grammar.