7 ms·
Movidius launches a $79 deep-learning USB stick
- sillysaurus3 9y agoSo what can you do with a deep-learning stick of truth? EDIT: Looks like the explanation is in a linked article: https://techcrunch.com/2016/04/28/plug-the-fathom-neural-compute-stick-into-any-usb-device-to-make-it-smarter/ https://techcrunch.com/2016/04/28/plug-the-fathom-neural-com... How the Fathom Neural Compute Stick figures into this is that the algorithmic computing power of the learning system can be optimized and output (using the Fathom software framework) into a binary that can run on the Fathom stick itself. In this way, any device that the Fathom is plugged into can have instant access to complete neural network because a version of that network is running locally on the Fathom and thus the device. This reminds me of Physics co-processors. Anyone remember AGEIA? They were touting "physics cards" similar to video cards. Had they not been acquired by Nvidia, they would've been steamrolled by consumer GPUs / CPUs since they were essentially designing their own. The $79 price point is attractive. I wonder how much power can be packed into such a small form factor? It's surprising that a lot of power isn't necessary for deep learning applications.
- legolassexyman 9y ago> The $79 price point is attractive. I wonder how much power can be packed into such a small form factor? It's surprising that a lot of power isn't necessary for deep learning applications. It runs pretrained NN, which is the cheap part. So this is a chip optimized to preform floating point multiplication and that's it.
- shams93 9y agoYeah I can run pretrained models on my pi3, that's not that exciting, its more exciting that 2nd handle graphics cards are dumping onto the market.
- shams93 9y agoIf you can use a set of say 3 of these running in paralell on a pi3 with tensorflow to train models from scratch then this is more interesting.
- RBerenguel 9y agoNo training, only evaluation: https://ncsforum.movidius.com/discussion/104/can-i-use-this-for-training-of-neural-network-with-tensorflow https://ncsforum.movidius.com/discussion/104/can-i-use-this-...
- wyldfire 9y agoI wonder how it stacks up against the snapdragon 410e. You can buy one on a dragonboard for roughly the same price ~$80 [1]. The dragonboard has four ARM cores, a GPU, plus a DSP. You could run OpenCV/FastCV on any or all three. [1] https://www.arrow.com/en/products/dragonboard410c/arrow-development-tools https://www.arrow.com/en/products/dragonboard410c/arrow-deve...
- zitterbewegung 9y agoWhy not both? Plug in the USB deep-learning USB stick. Use the snapdragon to do ETL and or download models to be run on the Movidlus so that it can perform inference. I am waiting for the day that we can do some nontrivial training on mobile hardware.
- legolassexyman 9y ago> Movidius's NCS is powered by their Myriad 2 vision processing unit (VPU), and, according to the company, can reach over 100 GFLOPs of performance within an nominal 1W of power consumption. Under the hood, the Movidius NCS works by translating a standard, trained Caffe-based convolutional neural network (CNN) into an embedded neural network that then runs on the VPU. This is sure to save me money on my power bill after marathon sessions of "Not Hotdog."
- joshvm 9y agoIgnoring the price tag this is about half the performance of the Jetson TX2 which can manage around 1.5TFLOPS on 7.5W. Interesting that you could use this to accelerate systems like the Raspberry Pi. The Jetson is a pain in the backside to deploy (at a production level) because you need to make your own breakout board, or buy an overpriced carrier. EDIT: I use the Pi as an example because it's readily available and cheap. There are lots of other embedded platforms, but the Pi wins on ecosystem.
- MacsHeadroom 9y ago1.5TFLOPS would have made the supercomputer top500 12 years ago. That's amazing.
- Dylan16807 9y agoKeep in mind that supercomputers are a lot less specialized than circuits for running neural nets. 12 years ago you could have gotten a stack of 5-8 7800 GTX cards and had 1.5TFLOPS of single precision. 11 years ago you could have had a stack of 5 cards with unified shaders. It's not fair to compare against the significantly more complicated route of getting 100 CPU cores working together with only 1-4 per chip.
- amelius 9y agoBut can't you configure the device to do e.g. fast matrix-vector multiplications instead of inference? I can be wrong, but I suspect that's what people do mostly on supercomputers anyway.
- j_s 9y agoCurrently out of stock as best I can tell.
- tuxracer 9y agoReally disappointing there doesn't appear to be a USB-C option
- skrebbel 9y agoOr a blue bike shed option, for that matter.
- make3 9y agoI know it may be surprising, but bandwidth is a really important factor for speed in deep learning, and USB-C would help with that.
- Quequau 9y agoMy assumption is that they never would build more capacity into a device, whose only interface was USB 2.0, than USB 2.0 can actually handle.
- alexchamberlain 9y agoI thought the stick was USB 3.0?
- Quequau 9y agoO.K. yes. I went back, looked, and yes it does support USB 3.0. Actually given the chip itself also apparently supports GigE it's a shame there isn't the option with that brought out.
- anoother 9y agoUSB-C is a connector, and has no effect on speed. I think you're referring to USB 3.1 gen2, which would double the theoretical bandwidth to 10Gbps.
- Dylan16807 9y ago
- visarga 9y agoInteresting applications for drones and robots. The small form factor and low energy requirements are the key.
- oelmekki 9y agoIt took me a while to find how it interfaces with the system (driver? dedicated application? just drop model and data in a directory which appeared on mounted key?), so I'll post it here. To access the device, you need to install a sdk which contains python scripts that allow to manipulate it (so, it seems like it's a driver embedded in utilities programs). Source: https://developer.movidius.com/getting-started https://developer.movidius.com/getting-started
- nl 9y agoIt's surprising how much attention this has had over the last few days, without any discussion of the downside: it's slow. It's true that it is fast for the power it consumes, but it is way (way!) to slow to use for any form of training, which seems to be what many people think they can use it for. According to Anandtech[1], it will do 10 GoogLeNet inferences per second. By very rough comparison, Inception in TensorFlow on a Raspberry Pi does about 2 inferences per second[2], and I think I saw AlexNet on an i7 doing about 60/second. Any desktop GPU will do orders of magnitude more. [1] http://www.anandtech.com/show/11649/intel-launches-movidius-neural-compute-stick http://www.anandtech.com/show/11649/intel-launches-movidius-... [2] https://github.com/samjabrahams/tensorflow-on-raspberry-pi/tree/master/benchmarks/inceptionv3 https://github.com/samjabrahams/tensorflow-on-raspberry-pi/t... ("Running the TensorFlow benchmark tool shows sub-second (~500-600ms) average run times for the Raspberry Pi")
- olegkikin 9y agoBut all those other solutions will consume orders of magnitude more power, especially the GPU. It's actually impressive what can be achieved on 1W of power.