9 ms·
PyTorch 2.0
- mdaniel 4y agodiscussion from (presumably) the PyTorch Conference announcement: https://news.ycombinator.com/item?id=33832511 https://news.ycombinator.com/item?id=33832511
- brucethemoose2 4y agoI'm hoping torch.compile is a gateway to "easy" non-Nvidia accelerator support in PyTorch. Also, I have been using torch.compile for the Stable Diffusion unet/vae since February, to good effect. I'm guessing similar optimizations will pop up for LLaMA.
- voz_ 4y agoIs there somewhere I can see your Stable Diffusion + torch.compile code? I am interesting in how you integrated.
- brucethemoose2 4y agoIn `diffusers` implementations (like InvokeAI) its pretty easy: https://github.com/huggingface/diffusers/blob/42beaf1d23b5ccb5db18f3f1e7555918ce40434c/docs/source/en/optimization/torch2.0.mdx#using-accelerated-transformers-and-torchcompile https://github.com/huggingface/diffusers/blob/42beaf1d23b5cc... But I also compile the VAE and some other modules, I will reply again later when I can look at my local code. Some modules (like face restoration or the scheduler) still dont like torch.compile. For the Automatic1111 repo (and presumably other original Stability AI implementations), I just add `m.model = torch.compile(m.model)` here: https://github.com/AUTOMATIC1111/stable-diffusion-webui/blob/master/modules/sd_hijack.py#L184 https://github.com/AUTOMATIC1111/stable-diffusion-webui/blob... I tried changing the options in the config dict one by one, but TBH nothing seems to make a significant difference behind the default settings in benchmarks. I haven't messed with compiling LORA training yet, as I dont train much and it is sufficiently fast, but I'm sure it could be done.
- brucethemoose2 4y agoHere is the InvokeAI code, minus the codeformer/gfpgan changes that dont work yet: https://gist.github.com/brucethemoose/ea64f498b0aa51adcc88f543bf80ed97 https://gist.github.com/brucethemoose/ea64f498b0aa51adcc88f5... I intend to start some issues for this on the repo soon(TM).
- datadeft 4y agoCould you give a bit more details about this? Do you have a link?
- brucethemoose2 4y agoSee the above reply ^
- simonw 4y ago"the MPS backend" - that's the thing that lets Torch run accelerated on M1/M2 Macs!
- sebzim4500 4y agoBased on George Hotz's testing it is very broken. It's possible it has improved since then, I guess but he streamed this a few weeks ago.
- dagmx 4y agoIt supports a subset of the operators (as mentioned in the release notes). I don’t think it’s broken for the ones that it does support though.
- norgie 4y agoYes, this is my experience. Many off the shelf models still don't work, but several of my own models work great as long as they don't use unsupported operators.
- mochomocha 4y agoThat's been my experience. However when fallback to CPU happens, it sometimes end up making a specific graph execution slower. But that's explicitly mentioned by the warning and pretty much expected.
- deleted 4y ago[deleted]
- mardifoufs 4y ago>Python 3.11 support on Anaconda Platform >Due to lack of Python 3.11 support for packages that PyTorch depends on, including NumPy, SciPy, SymPy, Pillow and others on the Anaconda platform. We will not be releasing Conda binaries compiled with Python 3.11 for PyTorch Release 2.0. The Pip packages with Python 3.11 support will be released, hence if you intend to use PyTorch 2.0 with Python 3.11 please use our Pip packages. It really sucks that anaconda always lags behind. I know the reasoning*, and I know it makes sense for what a lot of teams use it for... but on our side we are now looking more and more into dropping it since we are more of an R&D team. We already use containers for most of our pipelines, so just using pip might be viable. *Though I guess Anaconda chewed more than it can handle w.r.t managing an entire Python universe, and keeping up to date. Conda-forge is already almost a requirement but using the official package (with pip, in this case) has its own benefits for very complex packages like pytorch.
- brucethemoose2 4y agoThe Arch Linux PyTorch 2.0 packages are great if you are looking for "cutting edge," as they are compiled against CUDA 12.1 now, instead of 11.8 like the official nightly releases. You can also get AVX2 patched Python and optimized C Python packages through CachyOS or ALHP. But even Arch is still stuck on Python 3.10
- DreamFlasher 4y agoAfaik NumPy, SciPy, SymPy and Pillow are not managed/owned by Anaconda? At least here: https://numpy.org/about/ https://numpy.org/about/ Anaconda isn't mentioned.
- DreamFlasher 4y agoAh, yeah they do have a Python 3.11 release, just not on anaconda. Okay, yeah, for a couple of years now there isn't a good reason anymore to use anaconda anyways.
- mardifoufs 4y agoYes that's the issue! Most of the software is already ready, usable and just works... unless you use anaconda. Now that I think about it, is there some technical reason for that? I always thought it was mostly about stability, but I can't imagine python 3.11 being so unstable as to warrant waiting a whole year before even porting.
- fpgaminer 4y agoThe thing I'm looking forward to most is having Flash Attention built-in. Right now you have to use xformers or similar, but that dependency has been a nightmare to use, from breaking, to requiring specific concoctions of installing dependencies or else conda will barf, to being impossible to pin because I have to use -dev releases which they constantly drop from the repositories. PyTorch 2.0 comes with a few different efficient transformer implementations built-in. And unlike 1.13, they work during training and don't require specific configurations. Seemed to work just fine during my pre-release testing. Also, having it built into PyTorch might mean more pressure to keep it optimized. As-is xformers targets A100 primarily, with other archs as an afterthought. And, as promised, `torch.compile` worked out of the box, providing IIRC a nice ~20% speed up on a ViT without any other tuning. I did have to do some dependency fiddling on the pre-release version. Been looking forward to the "stable" release before using it more extensively. Anyone else seeing nice boosts from `torch.compile`?
- saiojd 4y agoI really wish compiling cuda extensions worked better out of the box. Is there a reason they can't bundle nvcc alongside pytorch outside of complexity/expense?
- brucethemoose2 4y agoLegal reasons. Filesize. Platform compatibility.
- saiojd 4y agoInteresting, I had not considered these points outside of file size! Do you think it is possible they will be overcome or is the chance 0?
- joseph_grobbles 4y ago[dead]
- bobbygunderson 4y ago
- yumraj 4y agoNo CUDA 12 support unfortunately..
- brucethemoose2 4y agoArch Linux builds it for CUDA 12.1
- VadimPR 4y agoAgree, it's a bit dated already due to this.
- tormeh 4y agoHopefully the AMD support doesn't just come in the form of ROCm...
- singularity2001 4y ago100% backward compatible That's (for me) the biggest reason why tensor flow fell out of flavor: the API broke too often (not just between tf 1 and 2)
- lucasap 4y agoIf anyone can edit it, I found a typo: > Python 1.8 (deprecating Python 1.7) > Deprecation of Cuda 11.6 and Python 1.7 support for PyTorch 2.0 It is clearly supposed to be python 3.8 and 3.7 respectively.
- marviel 4y ago> As an underpinning technology of torch.compile, TorchInductor with Nvidia and AMD GPUs will rely on OpenAI Triton deep learning compiler to generate performant code and hide low level hardware details. OpenAI Triton-generated kernels achieve performance that’s on par with hand-written kernels and specialized cuda libraries such as cublas.