3 ms·
Nice experiment! Not to minimize your work, in fact some ideas might be complementary to yours, a couple years ago there were already some approaches reaching 3
by m3at 4y ago
Nice experiment! Not to minimize your work, in fact some ideas might be complementary to yours, a couple years ago there were already some approaches reaching 30s on a single V100:
https://myrtle.ai/learn/how-to-train-your-resnet-8-bag-of-tricks/ https://myrtle.ai/learn/how-to-train-your-resnet-8-bag-of-tr...
- tbalsam 4y agoYes, David Page's work is lovely. This initial release is almost a bit for bit remake of the functionality of the original code, but built to be linear and hackable at basically any stage of the pipeline. Page gets a ton of respect for me for all of the novel stuff introduced, I spent like 80-90 hours plus trying to debug the minutiae it takes to get things working properly at that accuracy -- and doing that thing has been my career. It's a seriously impressive accomplishment to me and the ease of which he presents some of those changes in the blog feels like one of those baking shows where you get a good idea of how truly difficult the achievement is when you do it yourself. I wanted to start with his baseline but as that hackable workbench for my own purposes to explore some information theory concepts w.r.t. deep learning and etc. His code is beautiful but also a framework-within-a-framework and nearly purely functional so quick hacks are basically impossible beyond a certain point. There are tradeoffs of course. Continuing to drop bit depth and a few other improvements will probably carry things surprisingly far, so long as hardware compatibility with said hacks remains Gucci. There's also some Triton kernel hacks that we could dip into but it would taint some of the "pure simple python" goals for the project. But yes -- this is a port of David Page's work designed for a 1-2 hour quick-sketch experimenting researcher. I've found a few other improvements that I hope to refine and contribute to the repo at some point -- after I fix a few basic, glaring bugs like the console printing the whole progress chart again each time. But yes, we're well on our way and thank you so much for linking that -- I'm a rather large fan of his work and I truly hope that some more of his wizardry comes to public light for us to glean from. :D :)))) <3
- m3at 4y agoNow I realise that I _somehow_ totally missed your link to Page's work right in the README… Thanks for the detailed comment, I definitely will dive into your script soon! One early suggestion, you might want to try torch dynamo [1], anecdotally I had good speedups (~20%) on some image models; though not sure how significant the impact might be at this (relatively) small scale. [1] https://pytorch.org/docs/master/dynamo/ https://pytorch.org/docs/master/dynamo/
- tbalsam 4y agoThank you! And good feedback, the README could be condensed a bit more to be more clean and readable. Thanks so much for the suggestion, I'll take a look at it! And much appreciated to a huge degree on interest in the script, it's still not perfectly polished stylistically but now that the baseline checks out, we definitely have a lot of performance gains to be had in the next release! :D :) Feel free to ping me if you ever need anything, I'm not the most active on GitHub but if you ever need to reach me by email for questions/comments/thoughts/etc, hi [ period ] tysam [ the at symbol ] gmail [ period ] com is my email address. :D
- m3at 4y agoWill do if I have feedback :)