5 ms·
>Neural networks? See The Vanishing Gradient Problem. For what it's worth, Geoff Hinton and others have had a lot of success in the last few years by intention
by thisisdave 13y ago
>Neural networks? See The Vanishing Gradient Problem.
For what it's worth, Geoff Hinton and others have had a lot of success in the last few years by intentionally injecting noise into their neural network computations. The best-known example is [0], but he's also had some success just adding noise to the calculation of the gradient in good old sigmoid MLPs--even reducing the information content of the gradient below a single bit in some cases.
Going back further, stochastic gradient descent and other algorithms have shown for decades that trading accuracy for speed can be very desirable for neural network computations.
[0] http://arxiv.org/pdf/1207.0580.pdf http://arxiv.org/pdf/1207.0580.pdf
- varelse 13y agoExcept the degree of error introduced by techniques like "knockout" and "maxout" are a hyperparameter under user control as are the many possible implementations of stochastic gradient descent. What's being suggested here seems equivalent to saying that because you love the crunchy fired top of creme brulee, why not go ahead and torch the whole thing? As for contrastive divergence with Restricted Boltzmann Machines overcoming the vanishing gradient problem. That's true, but has anyone even tried doing this with 7-bit floating point and demonstrated it even works? I'm assuming recurrent neural networks relied on at least 32-bit point (correct me if I'm wrong). A reply to me flickered on here briefly indicating they haven't built a processor based on this yet which the author then deleted. I think this would be a much more interesting architecture if it would just stop trying to reinvent floating point or "if it ain't broke, don't "fix" it."
- thisisdave 13y ago>As for contrastive divergence with Restricted Boltzmann Machines overcoming the vanishing gradient problem. That's true, but has anyone even tried doing this with 7-bit floating point and demonstrated it even works? I wasn't referring to RBMs here, although it does look like I misremembered the details. See slides 74-76 of this presentation for a brief sketch [0]. In these feed-forward MLPs, the information content from each neuron is capped at one bit (which is quite a bit less than seven). I seem to recall him saying something similar (and with more detail) in a Google Tech Talk a few months later [1], although I'm not sure. [0]: http://www.ipam.ucla.edu/publications/gss2012/gss2012_10743.pdf http://www.ipam.ucla.edu/publications/gss2012/gss2012_10743.... [1]: http://www.youtube.com/watch?v=DleXA5ADG78 http://www.youtube.com/watch?v=DleXA5ADG78
- varelse 13y agoIndeed, but I quote: "The pre-training uses stochastic binary units. After pre-training we cheat and use backpropagation by pretending that they are deterministic units that send the REAL-VALUED OUTPUTS of logistics." So without trying to sound like a pedant, the manner in which the error is introduced is effectively a hyper parameter, no? The skills involved in doing this correctly will land you the big bucks these days: http://techcrunch.com/2014/01/26/google-deepmind/ http://techcrunch.com/2014/01/26/google-deepmind/ My problem is that I think the programming experience of an entire hardware stack based on 7-bit floating point would be abysmal and that companies that tried to port their software stack to it would die horribly (to be fair they'd deserve it though). OTOH If they'd just stick to the limited domains where this sort of thing is applicable, my perception of it would flip 180 degrees - it's probably a pretty cool image preprocessor. What it's not is the replacement for traditional CPUs. But the people behind this seem to have utter contempt for the engineers tasked with programming their magical doohickey, or putting this in TLDR terms: 1. Scads of 7-bit floating point processors 2. ??? 3. PROFIT!!!