7 ms·
Making deep learning go brrrr from first principles (2022)
- lockhouse 3y ago[flagged]
- IvanMilatForPM 3y ago[flagged]
- totetsu 3y ago[flagged]
- itisit 3y agoTo what are you referring? Has the post been edited after your comment?
- O_nlogn 3y agoI think gp's referring to this meme [1] embedded in the article. I believe a charitable interpretation would be that it's satirising the various entities who have voiced opinions against large scale deep learning (from skeptics and theorists, to social activist types that want to slow or stop DL's march). No need to take it too seriously. [1] https://horace.io/img/perf_intro/gpus_go_brrr.webp https://horace.io/img/perf_intro/gpus_go_brrr.webp
- tysam_and 3y agoH.He has some overlap with the EAI community which has more crab-bucket tendencies when picking topics and/or mocking other groups. Even the more positive sentiments seem to be somehow crab-bucketed a bit (and, bizarrely at least to me, a large amount of politicking). I'll visit there occasionally, but it's just too toxic a sphere for me to want to contribute research to. I get an itchy feeling whenever I try to seriously investigate an opportunity for contributing open source research in that group.
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- tysam_and 3y ago[flagged]
- jamesblonde 3y agoGreat point about Sutton's law. I always use Hinton's Capsule Networks as an example of an algorithm that may work in theory, but doesn't scale on existing hardware, due it's voting stage, which is not parallelizable on existing hardware accelerators.
- tysam_and 3y agoThat's a great point, I hadn't really thought about that a lot, and definitely had not made the hard connection. I do love how he continues to spin out new potential modalities for Deep Learning. While underrecognized due to the hype today, I think we darn well need far more of it these days! :D :))))
- ssivark 3y ago> Experiment noise has taught me so much over several thousands of manually-run (yes, I'm a masochist, it's a personal preference) experiments converging in a few seconds. Is there any resource that makes legible how to go about this? Or shares insights into the process of iterative learning through experiments, more broadly?
- tysam_and 3y agoI'd like to create it in the future, but I can share a rough version of what I have. Basically, as long as you gain more information than noise by scaling your experiments down (i.e., the smaller models create similar directions in the needle), then it's easier to run lots of experiments, and much more cheaply oftentimes too. Remember that performance oftentimes runs on a log(1+p) curve in terms of cost/time/complexity vs reward. If you write your codebases to minimize the expected time to make a particular change (fewer files, simpler proxy problem [as long as it transfers to your larger-scale problem!], more 'flat' structure, simple dataloaders, etc), then that is extremely valuable as well. What you are optimizing is the average time from idea-to-answer. What is nice is that experiments go from strongly-planned, biased (from human beliefs, etc) things to a more high-variance process where there isn't as much requirement to double-and-triple check everything before running it. There is often an _extremely_ poor mismatch between the very first impression of what many of us think will work and will not and what actually works when the dust settles. Sorta like how SGD bounces around the loss landscape before settling in. One thing that also is easier with fast experiments is to make the 'opposite' change if a planned change doesn't work out well, just to see how it goes. A tiny thing, but if you're going sequentially instead of batched (in terms of batched experiments) it can be useful. I use as few libraries as possible. A lot of the bulky monitoring solutions can be good for some enterprise things, I don't find they're good for innovating, but instead, integrating. You innovate on small test problems that are appropriately representative. When done, you integrate and see what difference it makes. Doing dumb things is one of the best ways to learn. Having access to everything to be able to print it out/log it/look at it in a chart is great too. But everything is bottlenecked by speed. Increase your experiment speed, where (critically!) your speed includes the time from idea to first answer coming back (and not just plain ol' runtime), and you'll find oftentimes it's easy to get 5-10x research speed improvements. We just don't do research all that efficiently these days, I don't think. But we can. And I'm sure we will, eventually. :3 There are a few things like scaling that can be difficult, but math is still math, and if you pick your proxies right, it should be an explored-translation-step after the proxy problem behaves as well as it is reasonably able. And if you have a secondary, slightly larger proxy problem, that's an even better thing in my experience, since it prevents the sticker shock of broken graphs and things failing for an unknown reason. All about how well you can pass the peach from one person to another. :')))) <3 <3 :')))) :'D :'D