5 ms·
So ist works just like i thought it would. Why are CNN so hyped? Wasnt all this already known decades ago? Or is it just because we can afford the computing pow
by MrFeynmannsJoke 10y ago
So ist works just like i thought it would. Why are CNN so hyped? Wasnt all this already known decades ago? Or is it just because we can afford the computing power?
- bbctol 10y agoThe underlying math was figured out a long time ago, but it's only been in recent years that we've had the computing power to test these out on lots of complicated, real-world classification problems, and had some incredible success.
- gcp 10y agoComputing power, but also some implementation tricks that turn out to make things a lot better. For example, activation via ReLU instead of sigmoid/tanh significantly improves the performance of deep neural networks. Then there's stuff like BatchNorm, Pooling, Dropout etc...
- nkozyra 10y ago1. Computing power 2. Data availability 3. Fast, large local storage
- zwieback 10y agoDecades ago I played around with neural nets but was frustrated because I either had to preprocess and normalize my inputs to the point where I didn't need a network anymore or I had to train a large network with so much data that it was not practical. Having a cookbook approach with a catchy name and orders of magnitude more processing power have revived neural nets and now they are finally doing something useful. Now everyone is jumping on the bandwagon so the field is progressing very quickly. Just because it's hyped doesn't mean it's not worth giving it a second look (although I'm still on the sidelines myself.)
- dougabug 10y agoThe basic CNN structure was in place, but as the saying goes, "The Devil's in the details." Early CNN's were applied to problems such as handwritten character recognition with rows of small grayscale image cells as inputs, and were much shallower, smaller models. Today's CNN's operate on full resolution, multi-channel images and video, and can be orders of magnitudes deeper and larger. For instance, ResNets have been proven to demonstrate monotonic performance improvements out to 1200 layers on benchmark datasets. This would have been unthinkable even a couple years ago. By way of comparison, even the state of the art VGG network architecture of a couple years ago originally had to be trained in stages to reach 16 and 19 layers for submission to ILSVRC 2014 (Xavier / MSRA initialization makes this unnecessary now). At the time, VGG and GoogleNet (22 layers) were considered to be extraordinarily deep CNN's.
- zackmorris 10y agoI argued back in 2000 (a year after I got my computer engineering degree) that AI wouldn't take off until computing moved from single threaded/single core to multithreaded/ multicore processing. The fact that we are only hearing about this stuff 15 years later makes me feel that that assertion was largely right. The biggest problem I see in AI is that the algorithms are generally fairly straightforward, but people haven't had the computing power to explore the problem space. We are seeing drastic improvement in things like video cards (routinely 1000+ cores) and data processing locality (map reduce). But processors have stagnated. If we really want AI in any reasonable timescale, we need large arrays of general-purpose cores with a sane communication protocol that doesn't fixate on things like caching, we need a hybrid between Go and Erlang to do concurrent functional programming in a readable way with automagic scaling over a network, and we need all this yesterday. The fancy schmancy AI algorithms will become apparent when processing power is no longer the primary limitation, and at that point we can optimize them.
- bilbobeer 10y agoIn 1965 Cooley-Tukey published their famous FFT paper, funny thing is that at Standard Oil they were using FFT in the 1950's, because it could be used for analyzing 'Russian' atomic bomb seismic signatures in real-time it took some 10+ years before it became public ( openly published ), but it was routine in BIG-OIL in 1950's. Now jump to CNN, we were doing CNN on Cray's in the 1970's on Seismic 3d data ( acoustic sound waves pressure/sheer ) in order to find oil deposits, its the same stuff I see in CNN algo's of today, and its the exact same stuff we were doing in the 1970's in visualizing seismic data (P-S wave ratios were the color that told you what kind of rock/substance you propagated ). We used to call this 'around the wheel', I suspect that its good to make the kids think this stuff is all something new, most of the foundation of computational ML was designed in 1950's, hell Von Neuman wrote the book on Cybernetic's, automating the human mind with computers, again 1950's. Another perspective is given the quality ( poor ) that's being leaked, and given corporate history, I suspect that BIG-OIL, and BIG-NSA of today have stuff that is super good and advanced and most of what they leak to GIT-HUB is just garbage. In summary keep it simply, learn how it really works, and then dial in your own algorithm, and you too will discover the holy-grail, but remember that edoocation is always 15-20 years lagging industry. It's always been this way, and always will. NSA/CIA has always mandated a competitive edge over the foreign gov's, never assume that anything you get for free, or academia is cutting-edge. ( Lastly George Green developed the Math here, back in 1816-ish, thus the Convolutional in CNN is 200 years old to be exact. )