11 ms·
Lessons from Optics, the Other Deep Learning
- metakermit 9y agoI like the parallel between optical lens systems and deep learning. I'm also kind of disappointed by the "arcane lore" status hyper-parameters have in different ML domains. I think it would be healthier for the community to make it a habit to explicitly document why a certain topology and layer sizes were selected. It's like providing documentation with your open source project – yes, it would be possible for knowledgable people to use it without it, but much more difficult and beginner unfriendly.
- manux 9y agoI wonder how documentable the space of hyperparameters really is (which is I think what the OP is poking at) with the current way we conceive of them, and also with how experiments currently happen. Often, people either reuse other people's architectures, or simply try 2 or 3 and stick with the best one, only changing the learning rate and such. I also wonder if there's a computation issue (training is long, we can only try so many things), or if it really is that we are working in the wrong hyperparameter space. Maybe there is another space we could be working in, where the HPs that we currently use (learning rate, L2 regularization, number of layers, etc.) are a projection from that other HP space where "things make more sense".
- azag0 9y agoIn this regard, it is similar to how natural sciences are done. The hyperparameter space of possible experiments is immense, they are expensive, so one has to go with intuition and luck. Reporting this is difficult. [edit:] In this analogy, deep learning currently misses any sort of a general theory (in the sense of theories explaining experiments).
- saguro 9y agoA DNN might be more effective at exploring the hypyerparameter space than people are with their intuition and luck. Rumor is Google has achieved this.
- yorwba 9y agoGoogle simply has the computational resources to cover thousands of different hyperparameter combinations. If you don't have that, you won't ever be able to do systematic exploration, so you might as well rely on intuition and luck.
- posterboy 9y agoThis is not accurate. Chess alone is so complex, brood force would still take an eternity, and they certainly don't have a huge incentive to waste any money just to show off (because that would reflect negatively on them). But how does it work? It's enough to outpace other implementations, alright. But the model even works on a consumer machine, if I remember correctly. I have only read a few abstract descriptions and I have no idea about deep learning specifically. So the following is more musing than summary: They use the Monte Carlo method to generate a sparse search space. The data structure is likely highly optimized to begin with. And it's no just a single network (if you will, any abstract syntax tree is a network, but that's not the point), but a whole architecture of networks --modules from different lines of research pieced together, each probably with different settings. I would be surprised if that works completely unsupervised; after all it took months from beating go to chess. They can run it without training the weights, but likely because the parameters and layouts are optimized already, and to the point of the OP, because some optimization is automatic. I guess what I'm trying to say is, if they extracted features from their own thought process (ie. domain knowledge) and mirrored that in code, than we are back at expert systems. PS: Instead of letting processors run small networks, take advantage of the huge neural network experts have in their head and guide the artificial neural network into the right direction. Mostly, information processing follows insight from other fields, and doesn't deliver explanations. The explanations have to be there already. It would be particularly interesting to hear how the chess play of the developers involved has evolved since and how much they actually do understand the model.
- joe_the_user 9y agoIn this regard, it is similar to how natural sciences are done. The hyperparameter space of possible experiments is immense, they are expensive, so one has to go with intuition and luck. Reporting this is difficult. I'd agree it's done in a sort-of scientific way. But I don't think you can say it's done the way natural science is done. A complex field, like oceanography or climate science, may be limited in the kind of experiments it can do and may require luck and intuition to produce a good experiment. But such science is always aiming to reproduce an underlying reality and the experiment aim to verify or not a given theory. The process of hyperparameter optimization doesn't involve any broader theory of reality. It is essentially throwing enough heuristics at a problem and tune enough that they more or less "accidentally" work. You use experiment to show this heuristic approximation "works" but this sort of approach can't be based on a larger theory of the domain. And it's logical that there can't be a set theory of how any approximation to any domain works. You can have a bunch of ad-hoc descriptions of approximation each of which works with a number of common domains but it seems logical these will remain forever not-a-theory.
- Cacti 9y agoExploring even a tiny, tiny, tiny part of the hyperparam space takes thousands of GPUs. And that is for a single dataset and model---change anything and you have to redo the entire thing. I mean, maybe some day, but right now, we're poking at like 0.00000000001% of the space, and that is state-of-the-art progress.
- Nydhal 9y agoA step in the right direction would be to encourage sharing negative results. It's important to know what to avoid too.
- romaniv 9y ago>I think it would be healthier for the community to make it a habit to explicitly document why a certain topology and layer sizes were selected. Also, which other topologies were tried and failed to produce good results. It's amazing that this information is missing from most modern ML papers.
- Cacti 9y agoIt is not often the case that someone actually knows why a hyperparam or architecture choice works. We pretend, sometimes, but frankly, it's mostly made up junk to cover the fact that most ML research involves a huge amount of intuitive guesswork and trial-and-error. And the loss surfaces vary. Even just changing the dataset or even the input size alters the loss surface and can easily break a model. It's not called Gradient Descent by Grad Student for nothing.
- mabbo 9y agoScary possibility: what if there is no good formal theory to explain how it works? What if intelligence, both animal and machine, is purely random trial and error and "this thing seems to work"? I don't believe that's true necessarily, but it will sure hamper the authors hopes.
- azag0 9y agoIn the same way that there is no formal theory for the exact shapes of proteins? I think it’s possible. But as with proteins, there are probably some general aspects of the problem that can be explained in simpler terms.
- canes123456 9y ago> what if there is no good formal theory to explain how it works? We don't need a theory that is perfect. Each theory was partially wrong but still lets you make useful predictions about the world. We need useful models that let you reason about the world. All models are wrong, some are useful. > What if intelligence, both animal and machine, is purely random trial and error and "this thing seems to work"? Evolution could be just considered random trail and error. However until we reach the singularity, we need people to speed up the evolution process by adapting and remixing pieces that worked before. We need models for what each level does so have ideas of what to try for a new application.
- jfoutz 9y agoI think the gp is referring to something a little more general. Let's pretend for a moment that our minds can be modeled by a formal system. Every thought has a chain of axioms grounding it. So, the scary part is, Godel showed us there are true things that can't be represented by a formal system. Maybe the useful models exist, but we can't comprehend them, because they're true outside of the set of rules we happened to get built into our minds? generally though, i'm on board with you. all models are wrong, some models are useful.
- eli_gottlieb 9y ago
- stochastic_monk 9y agoThis reminds me of ACDC, a Deep-fried Convnets-like[0] approach by some NVIDIA employees. [1] See section 1.1, where they state they could perform the operation in analog. [0]: https://arxiv.org/abs/1412.7149 https://arxiv.org/abs/1412.7149 [1]: https://arxiv.org/pdf/1511.05946 https://arxiv.org/pdf/1511.05946
- eb0la 9y agoI really miss having the building block rationale of all the perception/classification/segmentation networks out there. The only thing I've found really useful until now, is to put 2-fully connected layers if the classifier does now handle well classification... just because you needed a hidden perceptron layer for the XOR case. I hope to find more examples like that. If you know them, please share!!
- gaze 9y agoThe difference is the character of the non-linearity.
- dekhn 9y agoIf this is interesting to you, it will also be interesting to you that lenses perform the physical equivalent of an analog fourier transform, and physicists exploited this to compute wave spectra well before digital computers existed.
- 6502nerdface 9y agoSimilarly, the human cochlea performs a physical fourier transform of sound waves, and acoustic engineers used similar principles to create paper spectrograms back in the 1950s (check out the Kay Electric Co. Sona-Graph).
- sgt101 9y agoThe flaw (geddit?) in this is that deep networks are processing data from different domains, with different characteristics (which change in chaotic ways). In ML the choice was always "the simpler the better" - we used statistics and information theory to apply Occam's razor. Deep networks don't work that way and they do work well in some real domains, nature does not always prefer simple domain theories. If the laws of the universe suddenly changed designing optics would suddenly be difficult as well.
- TYPE_FASTER 9y agoI worked for a couple years on control systems software without any kind of experience or formal training. We would get source code, with some initial tuned parameters, then take our robot out into the field, and re-tune our control loops based on reality. Here was my takeaway: an engineer has to understand the domain and algorithms involved at a deep level, or they will not be productive. Or, you will need to have both an engineer and somebody with the domain knowledge and experience. It doesn't really matter what your problem domain is. If you're an engineer, and it's your job to make changes to a system, whether code or config, you need to understand it at a deep level. And your manager needs to understand this requirement. Otherwise, you will be guessing at changes, so your productivity will be horrible or non-existent.
- zwieback 9y agoVery true but the post makes a more subtle point: you don't have to have a model that explains everything as long as your model is predictive for the problem you're trying to solve. You can build a good telescope without understanding quantum mechanics. I guess that's the difference between science and engineering, broadly speaking.
- deleted 9y ago[deleted]
- amelius 9y ago> I guess that's the difference between science and engineering, broadly speaking. Science also doesn't need to have a model that explains everything. Because if it did, we wouldn't have science.
- IamNotAtWork 9y agoI thought based on your opening sentence that you were going to say having no understanding of the problem worked out okay. So are you saying you were unproductive at your old job?
- TYPE_FASTER 9y ago
- deleted 9y ago[deleted]
- TeMPOraL 9y agoTangential to the topic, but one of the references in this article is this amusing paper: http://nyus.joshuawise.com/batchnorm.pdf http://nyus.joshuawise.com/batchnorm.pdf ... which references an even better one: http://pages.cs.wisc.edu/~kovar/hall.html http://pages.cs.wisc.edu/~kovar/hall.html We've been having a solid laughfest in the office for the past 10 minutes or so.
- eonwe 9y agoI can feel for Mr. Kovar (the writer of the latter link). Looking at his resume, he did wisen up and did his master's thesis in computer science. I trust he's happier now than as a undergrad student.
- wrycoder 9y agoGe is easily "soldered" using indium and an ultrasonic soldering iron. If you don't have those, good luck. The standard technique is to set up a "Kelvin probe", with four contacts on the Ge sample. Pass a current from a constant current source (an IC or FET these days) between the outer two contacts and measure the voltage across the inner ones. It doesn't sound like his lab assistant set up something at which he could succeed, and that's a shame. He couldn't even repeat the room temperature reading.
- vog 9y agoAs funny as this may sound, there's some deep value in this: We deeply need more incentives in the academic world to report non-findings like that. We have a strong publication bias towards positive results, which has already become a huge problem. Moreover, we should have a much stronger focus on repeatability of experiments.
- deleted 9y ago[deleted]
- jimbokun 9y agoAnd maybe this is exactly the trick, make it low key and have a sense of humor about it. Sounds like a site dedicate to "My Ass" results would be extremely popular with grad students and real world researchers. Being able to know "it's not just me" and maybe even avoid some of the stumbling blocks others have run into, or to not just blindly use some approach that happened to work for one experiment, but seems to fail for many others.
- zwieback 9y agoThanks for introducing me to this blog! This is the money quote for me: There’s a mass influx of newcomers to our field and we’re equipping them with little more than folklore and pre-trained deep nets, then asking them to innovate.
- tinymollusk 9y agoAs one of the recent newcomers, should I feel defensive when I read something like this? I understand there are people with much more knowledge. Isn't this true of everyone, in every field? The message I've gotten is "try things out". Innovation isn't necessarily improving specific techniques, but applying them to new fields. To apply techniques to things that are more mundane like data processing in non-AI-focused companies, you're gonna need bodies who know how to apply these newer programming techniques to solve problems. Not every electrician has to understand electrical engineering.
- nlowell 9y agoI am new as well. I think the folklore aspect is not because the established people in the field are bad teachers, but because even they don't have much rationale besides "this is what seems to work". It's a new field, that's fine. Innovations are still as simple as "Oh we used a cyclical learning rate instead of constant" and boom.
- currymj 9y agoi don't think so. the author here also helped co-write the "deep learning is alchemy" talk that was somewhat controversial at NIPS. i think this is especially important if you purely want to do applications. we have a bag of tricks (dropout, batchnorm, different optimizers and learning rate schedules). we have no real theory for why any of this should work; often a proposed explanation will later turn out not to make sense. so the choice of how to train things comes down to "folklore", the community's collective experience. and there's no guarantee that folklore will generalize to your new architecture or dataset, and no way to know whether it even should. the presentation seems to have struck a nerve and there's papers and talks floating around now examining the performance of common architectures in very simple settings. it's probably worth paying attention to these at least in the background, as it will hopefully crystallize into a body of knowledge that will be useful for someone trying to decide on architectures and optimization techniques.
- twtw 9y agoAnother field that might be interesting to compare is analog design. There is a similar stack of theories: lumped element -> transmission line -> maxwell's equations. And yet analog IC design depends heavily on inherited mental models from mentors and modifications of well known topologies. Outsiders think it is black magic. The physics is all understood (nearly) perfectly, and yet knowing the details of QM that explains MOSFET operation helps not at all (or very little) when designing actual useful circuits. The real world considerations of parasitics, coupling, etc. dominate, and extensive formal analysis is not terribly useful. The general methodology is to make changes to the design based on intuition, simple predictive models that give you a direction, and previous experience, and then simulate to see how you did. A ton of high-quality engineering is done based on intuition, mental models, and patterns learned over years of experience. My hunch is that deep learning will be the same. EDIT: Just reread, and I want to clarify. I'm not saying that analog design is at the same stage of development as deep learning, or that it is anywhere near as ad hoc. Deep learning probably has a long way to go, but it could potentially end up in a similar state where years of experience is critical and intuition rules.
- wrycoder 9y agoA couple of times, I've heard Gerry Sussman at MIT (SICP author) give a talk on how bipolar transistor circuits are actually designed. You don't use SPICE or mesh analysis except in unusual circumstances (e.g. non-linear circuits) or to fine tune a completed design. As an example, a bias design goes something like this: "Let's see, I'll pin the base at five volts with a resistor divider. The emitter will be 0.6V below that. Then the emitter current will be (5.0 - 0.6) divided by the emitter resistor. The collector current will be essentially the same, so I can pick the collector load resistor to give me an appropriate quiescent point and make sure the output impedance is less than a tenth of the input impedance of the following stage (so I can ignore the latter)."
- taeric 9y agohttps://www.infoq.com/presentations/We-Really-Dont-Know-How-To-Compute https://www.infoq.com/presentations/We-Really-Dont-Know-How-... is the version of the talk I've seen. Highly recommended to everyone. I have yet to get his classical mechanics book. Gonna have to pull the trigger on that soon.
- nashashmi 9y agoLook at how many times this article was submitted and how long ago it was submitted and never gained traction. https://news.ycombinator.com/from?site=argmin.net https://news.ycombinator.com/from?site=argmin.net
- Houshalter 9y agoHNs algorithm could be improved if new storied were briefly shown on the front page so a few people could see them and have a chance of voting. Like how new comments start a the top of the thread and fall if they don't get votes.
- agitator 9y agoIt could also be that the younger engineers are engineers by training, while the more senior members of the team are Phd's in the field. What I'm trying to say is that Phd's come from an academic research background, while engineers come from a product focused background. The deep learning field is still dealing with a lot of unknowns, counter-intuitive responses to modifications, and pure experimentation. The engineers might just not realize the need for continued experimentation, and, for them, it may just feel like an undesirable waste of time to fiddle with parameters (as in, taking away time from developing the actual product). It's an alternate point of view, but something that I experienced.
- meri_dian 9y agoA comment below makes an important point that I think is worth repeating: "The power of digital computing is the power of modular expansion of objects. Analog circuits and computers don't have that. And current trained deep learning models don't have combinability and modularity either." A point I'd like to make: the brain exhibits properties of both digital and analog computers. It also exhibits repeating units in the neocortex which do vary but are uniform enough that neuroscientists are comfortable classifying them as discrete units within the brain. I believe we must look to how the brain implements effective modularity in the context of analog computation in order to replicate the success of digital computers with deep nets.
- gyom 9y agoOne of the problems with coming up with a good theory is that, at the end of day, we're building a system that's particularly suited for a certain kind of patterns. If you're building a facial-recognition convnet, there is something about the dataset of faces that is going to influence what works and what doesn't. When you're building digital circuits, they're expected not to care about what the bits mean, which patterns are more likely. It works for all possible inputs, with equal quality. There are things in common with how you would process faces and how you would recognize other visual objects, and that's why there are design patterns such as "convolutional layers come before fully-connected layers". In a way, the "no free lunch" theorem says that you are always paying a price when you specialize to a certain kind of patterns. It comes at the detriment to other patterns. So, any kind of stack of theories on ML/DL is going to be incomplete unless you say something about the nature of your data/patterns. (That doesn't mean that we can't anything useful about DL, but it just puts a certain damper on those efforts.)
- amelius 9y agoDeep learning is the new alchemy. So perhaps we can learn from alchemists? PS: The parallel between DL and optics is (if viewed historically) a bit misleading, because for building lenses we first had a theory.
- lexy0202 9y agoWhat the author refers to as "randomization strategies" is in fact regularisation: a set of techniques to prevent over fitting.
- deleted 9y ago[deleted]