10 ms·
I recognize the author Jascha as an incredibly brilliant ML researcher, formerly at Google Brain and now at Anthropic. Among his notable accomplishments, he a
by refibrillator 2y ago
I recognize the author Jascha as an incredibly brilliant ML researcher, formerly at Google Brain and now at Anthropic.
Among his notable accomplishments, he and coauthors mathematically characterized the propagation of signals through deep neural networks via techniques from physics and statistics (mean field and free probability theory). Leading to arguably some of the most profound yet under-appreciated theoretical and experimental results in ML in the past decade. For example see “dynamical isometry” [1] and the evolution of those ideas which were instrumental in achieving convergence in very deep transformer models [2].
After reading this post and the examples given, in my eyes there is no question that this guy has an extraordinary intuition for optimization, spanning beyond the boundaries of ML and across the fabric of modern society.
We ought to recognize his technical background and raise this discussion above quibbles about semantics and definitions.
Let’s address the heart of his message, the very human and empathetic call to action that stands in the shadow of rapid technological progress:
> If you are a scientist looking for research ideas which are pro-social, and have the potential to create a whole new field, you should consider building formal (mathematical) bridges between results on overfitting in machine learning, and problems in economics, political science, management science, operations research, and elsewhere.
[1] Dynamical Isometry and a Mean Field Theory of CNNs: How to Train 10,000-Layer Vanilla Convolutional Neural Networks
http://proceedings.mlr.press/v80/xiao18a/xiao18a.pdf http://proceedings.mlr.press/v80/xiao18a/xiao18a.pdf
[2] ReZero is All You Need: Fast Convergence at Large Depth
https://arxiv.org/pdf/2003.04887 https://arxiv.org/pdf/2003.04887
- throw10920 2y agoThis is a really manipulative way to categorically hand-wave away objections without actually responding to their content, in addition to having several logical fallacies (such as the appeal to emotion and the argument from authority). This is not in the spirit of intellectual curiosity that HN is for.
- LarsDu88 2y agoAdding to my reading list!
- salawat 2y ago>> If you are a scientist looking for research ideas which are pro-social, and have the potential to create a whole new field, you should consider building formal (mathematical) bridges between results on overfitting in machine learning, and problems in economics, political science, management science, operations research, and elsewhere. Translation to laymen: ML is being analogized to the mathematical structure of signaling between entities and institutions in society. Mathematician proposes problem that plagues one (overfitting in ML, the phenomena by which a neural network's ability to generalize is negatively impacted by overtraining so the functions it can emulate are tightly coupled to the training data), must plague the other. In short, there must be a breakdown point at which overdevelopment of societal systems or signaling between them makes things simply worse. I personally think all one need do is look at what would happen if every system were perfectly complied with to see we may already be well beyond that breakpoint in several industrial verticals.
- lubujackson 2y agoThe exciting thing about this idea is if you can correlate, say, economics with the works of ML, that means a computer program which you can run, revise and alter can directly give you measurable data about these complex system interactions that mostly have existed as a platonic idea since reality is too nuanced and multiple to validate concepts formally. With the idea that there is some subset of logic that sits below economics that is provable and exact. That is a powerful idea worth pursuing!
- nerdponx 2y agoThis idea has been pursued several times in the past, and it always ends up producing lots of interesting academic results and no practical conclusions. It's certainly an interesting perspective on the development of complex systems. The idea that an economy can be somehow overfitted to its own incentives and constraints I don't think is entirely new, cf the Beer Game. But as a general concept, it's certainly not something that usually finds its way into policy discussion, beyond some very specific talk about reshoring of certain critical industries. However, I think the most important benefit of this perspective is going to be providing yet another counterargument against the Austrian economics death cult.
- ahartmetz 2y agoIt seems to me that something similar to Adam Smith happened to the Austrians: their ideas have been cherry-picked. According to German Wikipedia, their main things were / are a focus on individual preferences, marginal utility, and a rejection of mathematical modeling(!) There was also something about lower state expenditures (...taxes...) giving better results for the people - that's the one that seems to be very popular with rich people for some reason. Go figure.
- jampekka 2y agoAustrian economics also rejects empirical assesment of its claims. Instead, universal thruths are derived "logically" (formal logic banned though) from "obviously true" axioms using a method called praxeology. It seems a lot like Scientology: the more you learn about it, the more bizarre it gets. And of course it's used to extract a lot of money for few benefactors.
- tablatom 2y agoInteresting timing for me! Just a couple of days ago I discovered the work of biologist Olivier Hamant who has been raising exactly this issue. His main thesis is that very high performance (which he defines as efficacy towards a known goal plus efficiency) and very high robustness (the ability to withstand large fluctuations in the system) are physically incompatible. Examples abound in nature. Contrary to common perception evolution does not optimise for high performance but high robustness. Giving priority to performance may have made sense in a world of abundant resources, but we are now facing a very different period where instability is the norm. We must (and will be forced to) backtrack on performance in order to become robust. It’s the freshest and most interesting take on the poly-crisis that I’ve seen in a long time. https://books.google.co.uk/books/about/Tracts_N_50_Antidote_to_the_cult_of_perf.html?id=XvQMEQAAQBAJ&redir_esc=y https://books.google.co.uk/books/about/Tracts_N_50_Antidote_...
- jfim 2y agoWe've seen this during the COVID pandemic supply chain disruptions as well, where just in time supply chain management doesn't work as expected when operating in an abnormal environment.
- soulofmischief 2y agoI'd always thought this conclusion was just a given. Highly optimized systems take full advantage of their environment and rely on a high degree of predictability in order to avoid redundant operations. These systems minimize the free energy in the system, and so very little free energy is available to counteract new forces introduced to the environment which act on the system. You'll find parallels in countless domains, since the very basis for learning and stabilization of a system revolves around becoming more or less sensitive to a given stimulus. Examples could be attention, supply chain economics, institutions, etc.
- jimkleiber 2y agoI was gonna come here to say that, especially how there was a shortage on toilet paper. I remember reading it was becuase factories were so efficient that when people started using the toilet at home instead of the office, it was hard to switch the factories from making commercial to residential toilet paper. I think someone even made the pun of paper-thin margins.
- thomasahle 2y agoI love the idea of ReZero, basically using a trainable parameter, alpha, in residual layers like this: Deep Network | xi+1 = F(xi) Residual Network | xi+1 = xi + F(xi) Deep Network + Norm | xi+1 = Norm(F(xi)) Residual Network + Pre-Norm | xi+1 = xi + F(Norm(xi)) Residual Network + Post-Norm | xi+1 = Norm(xi + F(xi)) ReZero | xi+1 = xi + αi F(xi) However, I haven't actually seen this used in practice. The papers we have on Gemma and Llama all still seem to be using layer norms. Am I missing something?
- immibis 2y agoIsn't this already part of F?
- thomasahle 2y agoI should add that alpha is initialized to 0.
- aoeusnth1 2y agoYour sound system has a volume dial to turn up and down the gain of the track even though you could get the same effect by re-recording the track at a higher volume; isn’t that curious?
- immibis 2y agoBut I don't optimise my track to have an ideal volume. I do optimise my AI like that.
- mrfox321 2y agoMore importantly, he invented diffusion models: http://proceedings.mlr.press/v37/sohl-dickstein15.pdf http://proceedings.mlr.press/v37/sohl-dickstein15.pdf
- RGamma 2y agoBrilliant enough to know he's helping build another atom bomb (presumably for peanuts)? And the nuclear briefcase is gonna be controlled by the ultrarich.