4 ms·
More than performance, the primary concern these days is where the data lives and how it moves: Do we need to copy it to a different storage for processing? Cha
by Radim 8y ago
More than performance, the primary concern these days is where the data lives and how it moves: Do we need to copy it to a different storage for processing? Change our company processes around how and where we store data? How many moving pieces to juggle?
It's a bit of a chicken and egg problem, since when designing new systems, management asks "How fast can we run this?", to which engineers reply "How much data is there and what speed are we aiming at?".
At this point, a failure to specify the desired volume and latencies leads to a sad cycle where scales are overblown ("just in case") and a complex solution chosen. Distributed solutions necessarily come with a massive overhead, so it is later discovered that even larger systems are needed…
Many such anecdotes circulate the industry, and I can add our own: In 2012 we implemented SVD, a math algorithm at the core of many machine learning techniques like PCA, LSI etc. This was a fast streamed "local" SVD implementation in Gensim (and now even faster in ScaleText). Suddenly, use-cases that needed a cluster of 12 beefy Hadoop (Mahout) machines could be performed on a single laptop, faster, and after a single `pip install`.
Our perf testing of word2vec (yet another streamed ML algo) showed a similar ROI pattern of local-vs-generic: https://rare-technologies.com/machine-learning-hardware-benchmarks/ https://rare-technologies.com/machine-learning-hardware-benc... (note the Spark graph at the end).