22 ms·
A non-technical explanation of deep learning
- great_wubwub 3y agoAs someone who knows barely enough to be dangerous, I like this. I'm sure it leaves enough out to make most experts angry, but it makes a lot of sense to me.
- time_to_smile 3y ago> I'm sure it leaves enough out to make most experts angry It's not that it leaves out details, it's that the articles metaphors are not actually correct in regards to the way deep learning works. This post mostly confuses both reinforcement learning and ensemble models with deep learning. If you only enough "enough to be dangerous" then this post will steer your intuition in the wrong direction.
- time_to_smile 3y ago> This is how neural networks work: they see many examples and get rewarded or punished based on whether their guesses are correct. This description more closely describes reinforcement learning, rather than gradient based optimization. In fact, the entire metaphor of a confused individual being slapped or rewarded without understanding what's going on doesn't really make sense when considering gradient optimization because the gradient wrt the to loss function tells the network exactly how to change it's behavior to improve it's performance. This last point is incredibly important to understand correctly since it contains one of the biggest assumptions about network behavior: that the optimal solution, or at least good enough for our concerns solution, can be found by slowing taking small steps in the right direction. Neural networks are great at refining their beliefs but have a difficult time radically changing them. A better analogy might be trying to very slowly convince your uncle that climate change is real, and not a liberal conspiracy. edit: it also does a poor job of explaining layers, which reads much more similar to how ensemble methods work (lots of little classifiers voting) than how deep networks work.
- cshimmin 3y agoWell said re: gradient optimization vs. "getting slapped". However, note that since NN optimization is almost always nonconvex, we are NOT guaranteed to arrive an a optimal (or even close-enough) solution. A major limitation of gradient based optimization on nonconvex problems is that they are very susceptible to getting trapped in local minima. But, for now it's the best tool we have, so we just have to hope that we get close enough, or just empirically run lots of times to find the best local minimum we can. Incidentally, this actually is more like a brute-force approach, but at the ensemble level, which is quite different than the article means it.
- nailer 3y agoDoes this article imply there are circumstances where a spreadsheet is a cat? What a poor example of technical writing.
- teacpde 3y agoNot the author, but to the author's defense, it is meant to be non-technical. And the first paragraph reads interesting to me.
- nailer 3y agoMost non technical people would think there are zero circumstances where a spreadsheet could be a cat.
- ricardobeat 3y agoThat's part of the explanation. It might not make sense at first, but you'll figure something out to avoid being slapped.
- adrianmonk 3y agoIt's obvious from context that it's the content of the media. To me at least. If I play you a song on Spotify and say, "Is this a saxophone?", you wouldn't say, "No, it's a iPhone running Spotify." If a policeman holds up a photograph of a person and says, "Is this the person who attacked you?", the victim doesn't say, "No, it's an 8 by 10 glossy print."
- giardini 3y agoNothing about LLMs?!
- fifteen1506 3y agoYeah, I need something to explain me about those Transformers things. I know it was published by Google in 2017 and that it is 'magic'. End of knowledge. Maybe I should ask ChatGPT?
- jedberg 3y ago> Maybe I should ask ChatGPT? You actually should, it spits out a pretty good explanation (sometimes).
- Izkata 3y ago2-hour video posted a month or two ago in a comment here: "Let's build GPT: from scratch, in code, spelled out." https://www.youtube.com/watch?v=kCc8FmEb1nY https://www.youtube.com/watch?v=kCc8FmEb1nY (I haven't gotten around to watching it yet)
- onikolas7 3y agoFunny. In the game black&white you would slap or pet your avatar to train it. The lead AI programmer on that was Demis Hassabis of deepmind fame.
- redog 3y agoSomehow he knew AI would be our Gods.
- Maultasche 3y agoThe description made me think of Black & White as well. I still have memories of smacking my creature around every time he ate someone.
- amelius 3y agoThe problem with deep learning is opposite. You can understand most of it with just high school math. Advanced math is mostly useless because of the dimensionality of neural nets.
- Utkarsh_Mood 3y agocan you elaborate further on what you mean by 'dimensionality of neural nets'? Thanks!
- amelius 3y agoYes, I mean the huge number of trainable parameters.
- deleted 3y ago[deleted]
- mach1ne 3y agoYes, it’s really rather like alchemy in some sense. Stuff works, and often nobody knows exactly why.
- uoaei 3y ago"I don't follow the latest ML scaling and theory research" does not in any way equate to "these things are unknowable".
- lhnz 3y agoHm, watching Neel Nanda videos recently and I do get the feeling that there are lots of unknowns in ML and also in what trained networks have learnt.
- deleted 3y ago[deleted]
- uoaei 3y agoThat's like saying you understand state-of-the-art CFD code because you can read Fortran. There are many aspects to learning systems that we still don't have any kind of grasp on, and will take more than a little advanced math (statistics/probability theory, transport theory, topology, etc.) to understand as a community. Dunning-Kruger is probably more common in spaces like this one, where people carry social capital for being able to "spin up quickly". But the true meta-skill of upskilling is turning unknown unknowns (UU) into known unknowns (KU), and then into known knowns (KK). It's not enough to just jump from UU to KK through osmosis by reading blog posts on a news aggregator, because there will still be a huge space of unknowns not covered by that approach.
- wrs 3y agoThis is the funniest refutation of the Chinese Room argument that I’ve seen. Note that at the end, it’s still the case that none of these people can recognize a cat.
- pringk02 3y agoDoesn't that mean it supports the Chinese room argument? I'm not sure I follow your reasoning. (also, popular conciousness forgets that technically the Chinese Room argument is only arguing against the much narrower, and now philosophically unfashionable, "Hard AI" stance as it was held in the 70s)
- wrs 3y agoI understand the Chinese Room argument to be that because the human in the room doesn’t understand Chinese, the system doesn’t understand Chinese. In this case, none of the humans can recognize cats, but the collective can.
- cscurmudgeon 3y agoThats not the Chinese Room argument. The argument says just because a system processes X doesn't imply it has consciousness of X.
- wrs 3y agoThe flaw is the unsupported assertion that the whole system being conscious of X depends on a part of the system being conscious of X. The same assertion would fail here in the same way.
- mensetmanusman 3y agoAt these levels of discussion everything is asserted as axioms to see what the consequences are. If a human isn’t conscious of Chinese, but the arrangement of paper rules is, one would have to assert that paper + human is conscious.
- hgsgm 3y agoNon-technical, non-accurate. "Truthy", buzzfeed/huffpo quality.
- pkdpic 3y agoI love this, but Im always confused in these kinds of analogies what the reward / punishment system really equates to... Also reminds me of Ted Chiang warning us that we will torture innumerable AI entities long before we start having real conversations about treating them with compassion.
- time_to_smile 3y agoDon't love it, it's not correct. > what the reward / punishment system really equates to Nothing, and least as far as neural network training goes. This is an extremely poor analogy regarding how neural networks learn. If you've ever done any kind of physical training and have had a trainer sightly adjust the position of your limbs until what ever activity you're doing feels better, that's a much closer analogy. You're gently searching the space of possible correct positions, guided by an algorithm (your trainer) that knows how to move you towards a more correct solution. There's nothing analogous to a "reward" or "punishment" when neural networks are learning.
- GaggiX 3y ago>There's nothing analogous to a "reward" or "punishment" when neural networks are learning. Well deep reinforcement learning.
- cshimmin 3y agoYeah but even in that case, "reward" is just the thing a NN is trying to predict. The NN itself is not receiving the reward (or any punishment). Instead, it's following gradient signals to improve that estimate of reward, which is then used as a proxy for an optimal policy decision.
- commandlinefan 3y ago> what the reward / punishment system really equates to Well, in the article, it says the punishment was a slap. On the other hand, he just says "she gives you a wonderful reward"... so you're left to use your imagination there.
- charcircuit 3y agoWhy is violence and praise being used to illustrate gradient descent? Why does each person get to see the entire input data?
- jstx1 3y agoDoes stuff like this help anyone? I still haven’t forgiven CGP Grey for changing the title to his 2017 ML video to “How AIs, like ChatGPT, learn”. The video is about genetic algorithms and has nothing to do with ChatGPT. (or with anything else in modern AI)
- deleted 3y ago[deleted]
- SnooSux 3y agoI've barely forgiven him for explaining genetic algorithms and acting like they have any relevance to contemporary ML research. The footnote video was an alright explanation of backprop. If that were part of the main video that would have been reasonable. I really like his history/geography videos but anything technical leave a lot to be desired. And don't get me started on Humans Need Not Apply.
- jstx1 3y ago> And don't get me started on Humans Need Not Apply. Well now you have to tell us. :) Many of the concrete examples in that video are exaggerated and/or misunderstood but the general question it asks - what to do when automation makes many people unemployable through no fault of their own - seems valid.
- bolyarche 3y ago[dead]
- musicale 3y ago> what to do when automation makes many people unemployable through no fault of their own - seems valid Unfortunately the video doesn't answer its own question directly. The answer for the past 40 years or so seems to be "move them to lower-paying service jobs, or out of the job market entirely."
- 3y ago
- vrglvrglvrgl 3y ago[dead]
- zvmaz 3y agoI have met people who think they understand a particular topic I am versed in, but actually don't. Similarly, I am often wary that I get superficial knowledge about a topic I don't know much about through "laymen" resources, and I doubt one can have an appropriate level of understanding mainly through analogies and metaphors. It's a kind of "epistemic anxiety". Of course, there are "laymen" books I stumbled upon which I think go to appropriate levels of depth and do not "dumb down" to shallow levels the topics, yet remain accessible, like Gödel's Proof, by Ernest Nagel. I'd be glad to read about similar books on all topics, including the one discussed in this thread. Knowledge is hard to attain...
- sainez 3y agoI find the best way to learn technical topics is to build a simplified version of the thing. The trick is to understand the relationship between the high level components without getting lost in the details. This high level understanding then helps inform you when you drill down into specifics. I think this book is a shining example of that philosophy: https://www.buildyourownlisp.com/ https://www.buildyourownlisp.com/. In the book, you implement an extremely bare-bones version of lisp, but it has been invaluable in my career. I found I was able to understand nuanced language features much more quickly because I have a clear model of how programming languages are decomposed into their components.
- joe_the_user 3y agoI find the best way to learn technical topics is to build a simplified version of the thing. The trick is to understand the relationship between the high level components without getting lost in the details. This high level understanding then helps inform you when you drill down into specifics. I agree but that's a good guide to build a technical understanding of a complex subject, not sufficient-in-itself tool set for considering questions in that complex subject. Especially, I'll people combining some "non-technical summary" of quantum-mechanics/Newton Gravity/genetic engineer/etc with their personal common sense are constant annoyance to me whenever such topics come here.
- 3y ago
- clarle 3y agoTotally aware that this isn't a fully formal definition of deep learning, but one interesting takeaway for me is realizing that in a way, corporations with their formal and informal reporting structures are structured in a way similar to neural networks too. It seems like these sort of structures just regularly arise to help regulate the flow of information through a system.
- 0xBABAD00C 3y agoThere is research claiming the entire universe is a neural network: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7712105/ https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7712105/
- __loam 3y agoFrankly, stuff like this makes me more skeptical of the ML community. Remember when people thought Brains were just really complicated hydraulic systems?
- 0xBABAD00C 3y agoVanchurin is a physics professor, but ok https://twitter.com/vanchurin?lang=en https://twitter.com/vanchurin?lang=en I actually think that interdisciplinary work like this can shed light into common structures across physics, biology, neuroscience, CS, etc. If anything, I wish there were more attempts to explore the connections between these disciplines.
- opportune 3y agoThe author is basically a crank, it looks like they held some teaching positions in an unrelated subject at various universities yielding the “professor” title, and originally studied/published actual research in cosmology decades back. Might as well be skeptical of the math community because of circle squarers
- 0xBABAD00C 3y ago
- sainez 3y agoIf anyone is looking for a quick overview of how LLMs are built, I highly recommend this video by Steve Seitz: https://www.youtube.com/watch?v=lnA9DMvHtfI https://www.youtube.com/watch?v=lnA9DMvHtfI. It does an excellent job of taking you from 0 to a decent understanding without dumbing down the content or abusing analogies.
- MandieD 3y agoThis really was excellent and just what I was looking for to explain what LLMs are to non-CS people. Thanks!
- _gmax0 3y agoThe most concise and intuitive line of explanation I've been given goes along the lines of this: 1 - We want to model data, representative of some system, through functions. 2 - Virtually any function can be expressed by a n-th order polynomial. 3 - We wish to learn the parameters, the coefficients, of such polynomials. 4 - Neural networks allow us to brute-force test candidate values of such parameters (finding optimal candidate parameters such that error between expected and actual values of our dataset are minimized) Whereas prior, methods (e.g. PCA) could only model linear relationships, neural networks allowed us to begin modeling non-linear ones.
- jacksnipe 3y agoExcept gradient descent is about as far from brute force as it gets
- _gmax0 3y agoSure, under the assumption that your parameter space is convex.
- deleted 3y ago[deleted]
- teruakohatu 3y ago> Whereas prior, methods (e.g. PCA) could only model linear relationships, Prior methods also allowed modelling of non-linear relationships, eg. Random Forests.
- pedrosorio 3y agoMentioning polynomials is a pretty poor way to explain it for two reasons: - It requires some mathematical understanding so will exclude some part of the non-technical audience - It is the incorrect analogy. Non-linearities in neural networks have nothing to do with polynomials. In fact, polynomial regression is a type of linear regression, and for the most part, it sucks. Also, as someone mentioned, all the “serious” alternative ML methods prior to the deep learning revolution allow modeling non linearities (even if just through modification of linear regressions, like polynomial regression).
- lxe 3y agoThis has Three-Body Problem vibes :)
- apomekhanes 3y ago[flagged]
- Myrmornis 3y ago> they see 3 spreadsheets of numbers representing the RGB values of the picture. This needs expanding: it's the sort of thing that's easy for a programmer to say, but few non-{programmer,mathematically trained person} are going to see that an RGB value has 3 parts and so a collection of RGB values could be sliced into 3 sheets.
- romwell 3y ago...or know what Ruth Ginsburg Bader has anything to do with it all. The RGB color model and representation of images in it is already technical. Anyone who knows what it means also wouldn't need to be told the following quip: >Also note that computers see things as multi-dimensional tables of data. They don't look at a "picture" - they see 3 spreadsheets of numbers representing the RGB values of the picture. ...which is the only time RGB is mentioned in the article. That's before we get to the part that "multidimensional" here is extraneous, and doesn't even match the typical usage (where RGBA is stored as a single 32-bit value). Everything is a tape of 1's and 0's, "multidimensionality" comes from interpretation of data. The dimension of image data is still 2: each pixel is a sample a 2D projection of a 3D world, and is related to other pixels in a way that's different than, say, those of letters in a line of text, or voxels (letters don't have a a well-defined "up" neighbor, voxels have more well-defined neighbors than pixels do).
- lhnz 3y agoI have a few funny analogies that I think kind of work. 1. "gradient descent" is like tuning a guitar by ear and listening to the beat frequencies ("loss") and then decreasing these by tuning a string up or down. 2. the best I can come up with for "backpropagation" is to imagine a clever device that can tirelessly optimize a Rube Goldberg machine for you but as a science, not an art.
- tartakovsky 3y agoMade me laugh because it's true, funny. Well done!
- dukeofdoom 3y agoI asked ChatGPT: Sure, I can explain deep learning using the analogy of mixing paint colors. When you mix different colors of paint, you can create new colors that are a combination of the original colors. Similarly, in deep learning, you can combine different types of data, such as images, text, and audio, to create a more comprehensive understanding of the information. In a neural network, which is the basic building block of deep learning, each layer is like a different color of paint. The input layer represents the initial data, which is like a blank canvas. As the data passes through each layer, the network extracts increasingly complex features, similar to how mixing different colors creates more complex hues. Each layer in a neural network can be thought of as a different color that is mixed with the previous layer's color. As the colors mix, the resulting color becomes more complex and nuanced, just as the data becomes more sophisticated and detailed as it passes through the layers of the neural network. At the end of the process, the output layer represents the final result, which is like a finished painting. The deep learning algorithm can then use this final output to make predictions or classifications based on the input data. So, in essence, deep learning is like mixing paint colors to create new and more complex hues, but instead of colors, the mixtures are made up of data and the neural network layers are the different colors being mixed.
- dukeofdoom 3y agoCertainly! Here's an explanation of how a neural network is trained using the analogy of mixing paint colors. Initialization: The weights and biases of the neural network are like the starting colors of the paint mixture. Just like how you need to start with some base colors to mix new ones, the neural network needs some initial weights and biases to begin the training process. Forward Pass: During the forward pass, the input data is fed into the neural network, and the network produces an output. This output is like the mixed color of paint that results from combining different colors. Backward Pass: During the backward pass, the error between the predicted output and the actual output is calculated using a loss function. This error is like a measure of how close the mixed color is to the desired color. Just like how you would adjust the mixture of colors to get the desired hue, the neural network adjusts the weights and biases of the neurons to minimize the error between the predicted output and the actual output. Repeat: Steps 2 and 3 are repeated many times with different input data points, and the weights and biases of the neural network are updated after each iteration. This process is like adding more colors to the mixture and adjusting the amounts until the desired color is achieved. Validation: After several epochs, the neural network is evaluated using a validation set of data to determine if it's overfitting or underfitting the training data. This step is like checking the mixture of colors to see if it matches the desired color under different lighting conditions. Testing: Once the neural network has been trained and evaluated, it can be used to make predictions on new data. This step is like using the paint mixture to create a painting. In summary, just like mixing paint colors to achieve a desired hue, the neural network adjusts its weights and biases to minimize the error between the predicted output and the actual output until it accurately predicts the output for new data.
- deleted 3y ago[deleted]
- kagaw 3y agoI see, so neural network works like Naruto's Kage Bunshin technique where the learning of the clones will pass to its origin
- hyperdimension 3y ago> They respond that this sounds very convoluted and they'll only agree to do it if you call them "colonel". Cute.