8 ms·
It would also be nice to remove the "magically thinking" around machine learning. It's mathematically related to all prior signal processing techniques (mostly
by xyzzy21 4y ago
It would also be nice to remove the "magically thinking" around machine learning. It's mathematically related to all prior signal processing techniques (mostly a proper superset) but it also have fundamental limits that no one talks about seriously. ML et al. are NOT MAGIC but they are treated as if they were.
And that is in itself a dangerous moral and ethical lapse.
- wolverine876 4y agoPeople will (and I'm sure do) use this magical thinking politically, persuading people to trust the computer and therefore, unwittingly, trust the persons who control the computer. That, to me, is the greatest threat - it is an obvious way to grab power, and most people I know don't even question it. It's a major consequence of mass public surveillance.
- heavyset_go 4y agoBureaucracies would love for a blackbox to delegate all of their decisions and responsibilities to in an effort to shift liability away from themselves. You can't be liable for anything, you were just doing what the computer told you to do, and computers aren't fallible like people are.
- bell-cot 4y agoIn my wishful thinking, by far the best way to do that would be for the courts to stick companies with full legal liability for the shortcomings of their "machine learning" systems. And if it's fairly easily demonstrate that GiantCo's ML decision making system is a sexist, racist, ageist...then GiantCo is not just guilty, but also presumed to have known in advance that they were systematically and deliberately on the wrong side of the law.
- arcticfox 4y ago> ML et al. are NOT MAGIC but they are treated as if they were. They're not magic - nothing is, but what are they? > but it also have fundamental limits that no one talks about seriously What are these fundamental limits? 20 years ago I imagine skeptics in your camp would have set these "fundamental" limits at lower than DALL-E 2, GPT-3, AlphaStar etc. Or are you talking about limits today? In which case, sure, but I think "fundamental" is the wrong word to use there given they change continuously. > It's mathematically related to all prior signal processing techniques (mostly a proper superset) And human brains are what if not signal processing machines?
- amelius 4y ago> They're not magic - nothing is, but what are they? Emergent magic.
- travisgriggs 4y ago> And that is in itself a dangerous moral and ethical lapse. Agreed. Is it an original lapse, or derivative though? When researchers/engineers oversell their story to get the funding they wouldn't otherwise get, where is the collapse? With the engineers/researchers? Or with the forces that built a system where that was the only way forward for them? When a hungry thief steals to eat, is the thief morally bankrupt? Or is those that engineered the shortage?
- savanaly 4y agoIs it the ones who engineered the shortage, or the ones who engineered the system in which the ones who engineered the shortage operated when they were designing the other system?
- OccamsRazr 4y ago> And that is in itself a dangerous moral and ethical lapse. More importantly, it can be a dangerous business lapse.
- ravi-delia 4y agoMaybe this is just my soft, theory-laiden pure math brain talking, but I'd be a lot less impressed with machine learning if we had a decent formal understanding of them. As is they're way weirder than I think most engineering types give them credit for. But then again, that's how I feel about a lot of applied stuff, it all feels a little magic. I can read the papers, I can mess around with it, but somehow it's still surprising how well it can work.
- SleekEagle 4y agoUltimately it comes down to gradient-based descent (which is pretty magical in its own right), but what's most surprising to me is that the loss landscape is actually organized enough to yield impressive results. Obviously the difficulties of training large NNs are well-documented, but I'm surprised it's even that easy
- SemanticStrengh 4y agoNNs are just glorified logistic regression. People should simply understand that neural networks cannot emulate a dumb calculator accurately, this simple fact is enough to realize being an universal approximator is in practice a fallacy, and true Causal NLU or AGI is essentially out of reach of neural networks, by design. Only a brain fidel architecture would have hope however C.elegans retro engineering is underfunded and spiking neural networks are untrainable.
- Der_Einzige 4y agoUsing gradient based techniques does a LOT to force neural network weights to resemble surfaces that they do not at all look like when using global optimization and gradient free techniques to optimize them. Most of the stupid crap that people give about degenerate cases where deep learning doesn't work (cartpoll in reinforcement learning, sine/infinite unbounded functions) are showcasing how bad gradient based training is - not how bad deep learning is at solving these problems. I can within seconds solve cartpoll with neural networks using neuroevolution of weights....
- hoseja 4y ago>cannot emulate a dumb calculator accurately Neither can people, for the most part.
- mpfundstein 4y agoknock knock. some critic from the 70s arrived. hows gofai going?
- SemanticStrengh 4y agoOh yes it's not GOFAI that has won the ARC challenge it's neural networks, right? right? https://www.kaggle.com/c/abstraction-and-reasoning-challenge https://www.kaggle.com/c/abstraction-and-reasoning-challenge I have more expertise in deep learning than anyone else here and the delusions of the incoming transformer winter will be painful to watch. In the meantime, enjoy your echo chamber.
- woopwoop 4y agoHonestly at this point it kind of is magic. These things are knocking out astonishing novel tasks every month, but the state of our knowledge is "why does sgd even work lol". There is no coherent theory.
- srean 4y ago> "why does sgd even work lol" I find this hand a little over played. It depends on the degree of fidelity we demand of the answer and how deep we want to go questioning the layers of answers. However, if one is happy with a LOL CATS fidelity, which suffices in many cases, we do have a good enough understanding of SGD -- change the parameters slightly in the direction that makes the system work a little bit better, rinse and repeat. No one would be astonished that using such a system leads to better parameter settings than ones starting point, or at least not significantly worse. Its only when we ask more questions, ask deeper questions that we get to "we do not understand why SGD works so astonishingly well"
- Filligree 4y agoWhy are there so few local minima, you mean? I think it’d have to be related to the huge number of dimensions it works on. But I have no idea how I’d even begin to prove that.
- srean 4y agoIts not even certain that they are few. Whats rather unsettling is that with these local moves of SGD the parameters settle on a good enough local minima in spite of the fact that we know that many local minima exists that have zero or near zero training loss. There are glimmers or insight here and there but the thing is yet to be fully understood
- woopwoop 4y agoYeah I didn't mean to imply "Why does SGD result in lower training loss than the initial weights" is an open question. But I don't think even lolcatz would call that a sufficient explanation. After all if the only criterion is "improves on initial training loss" you could just try random weights and pick the best one. The non-convexity makes sgd already pretty mysterious, and that is without even getting into the generalization performance, which seems to imply that somehow sgd is implicitly regularizing.
- vincentmarle 4y agoWhen you have a complex system that produces nth-order effects, then the only approach is to treat it as empirical phenomena (aka black box magic), and that is what most research papers in this field do.
- throwawaygh 4y agoIn the 80s and 90s it was really common to anthropomorphize spaghetti code. Just because something is difficult to analyze doesn't mean it has limitless power.
- nonrandomstring 4y ago> It would also be nice to remove the "magically thinking" around machine learning. To be honest it would be a morally and ethical less dangerous world if we could get our feet back on the ground in relation to digital technologies in general. > fundamental limits that no one talks about seriously. I am starting to touch and stumble into the invisible cultural walls that I think make people "afraid" to talk about limitations. I am not yet done analysing that, but suspect it has something to do with the maxim that people are reluctant to question things on which their salary depends. That seems to be a difference between "scientists" and "hackers" in some way. Going back to Hal Abelson's philosophy, "magic" is a legitimate mechanism in coding, because we suppose that something is possible, and by an inductive/deductive interplay (abduction) we create the conditions for the magic to be true. The danger comes when that "trick" (which is really one of Faith) is mixed with ignorance and monomaniacal fervour, and so inflated to a general philosophy about technology.
- time_to_smile 4y ago> suspect it has something to do with the maxim that people are reluctant to question things on which their salary depends. I once worked on a team that spent a lot of time building models to optimize parts of the app for user behavior (trying to intentionally remain vague for anonymity reasons). Through an easy experiment I ran I ended up (accidentally) demonstrating that the majority of DS work was not adding more than minimal improvements, and so little monetary value and it did not justify any of the time spend on this. I was let go not long after this, despite having help lead the team to record revenues by using a simple model (which ultimately was what proved the futility of much of the work the team did). Just a word of caution as you > start to touch and stumble into the invisible cultural walls that I think make people "afraid" to talk about limitations
- nonrandomstring 4y agoGood story. I guess you had done with your work there. Sometimes teams/places have a way of naturally helping us move to the next stage. Competences work at multiple levels, visible and invisible. Being good at your job. Showing you're good at your job. Believing in your job. Getting other people to believe in your job. Getting other people to believe that you believe in your job... and so on ad absurdum. Once one part of that slips the whole game can unravel fast.
- godelski 4y agoI think this is a common problem and comes because we stressed how these models are not interpretable. It is kinda like talking about Schrödinger's cat. With a game of telephone people think the cat is both alive and dead and not that our models can't predict definite outcomes, only probabilities. Similarly with ML people do not understand that "not interpretable" doesn't mean we can't know anything about the model's decision making, but that we can't know everything that the model is choosing to do. Worse though, I think a lot of ML folks themselves don't know a lot of stats and signal processing. They just aren't things that aren't taught in undergrad and frequently not in grad.
- mirntyfirty 4y agoAlong with that it becomes remarkably more difficult to distinguish causation vs correlation although I’m sure that point is heavily debated
- godelski 4y ago> difficult to distinguish causation vs correlation I mean this is an extremely difficult thing to disentangle in the first place. It is very common for people in one breath to recite that correlation does not equate to causation and then in the next breath propose causation. Cliches are cliches because people keep making the error. People really need to understand that developing causal graphs is really difficult, and that there's almost always more than one causal factor (a big sticking point for politics and the politicization of science, to me, is that people think there are one and only one causal factor). Developing causal models is fucking hard. But there is work in that area in ML. It just isn't as "sexy" because they aren't as good. The barrier to entry is A LOT higher than other type of learning, so this prevents a lot of people from pursuing this area. But still, it is an necessary condition if we're ever going to develop AGI. It's probably better to judge how close we are to AGI with causal learning than it is for something like Dall-E. But most people aren't aware of this because they aren't in the weeds. I should also mention that causal learning doesn't necessitate that we can understand the causal relationships within our model, just the data. So our model wouldn't be interpretable although it could interpret the data and form causal DAGs.
- Jenk 4y agoSupplant "magic" with "not understood" Suddenly it all becomes a lot more palatable that many don't know how it works.