3 ms·
I have the same problem with anything over 512x512 on my M1 Ultra with 128GB. VRAM must be capped.
by enduser 4y ago
I have the same problem with anything over 512x512 on my M1 Ultra with 128GB. VRAM must be capped.
- moneycantbuy 4y agoThanks for the Ultra data point. I'm able to get 768x896 to run, but the output image is still white noisy at 50 ddim steps, perhaps related to the phenomena of being trained/windowed on 512x512 images as sibling squeaky-clean described. RAM usage at various sizes: 512x512 14 GB, 768x768 26 GB, 768x896 32 GB
- squeaky-clean 4y agoThose values seem really high compared to my setup, windows/nvidia/lstein repo. For me 512x512 uses 6.1GB. Random guess but I think your pipeline is running with full-precision floats (32bit), while by default the repo should be using autocast() which will try to use half-precision floats wherever possible. I know an optimizedSD repo exists and one of the steps they take is explicitly setting precision to half. (And other changes that reduce memory usage but decrease iteration speed). However I don't know how M1/Metal handles half-precision, hopefully it doesn't just cast them back to 32bit. Also white noisy images at 50 steps seems off to me. At 50 steps in a large image I definitely get a visible product. It's just often non-euclidean or very scattered bits of organization and chaos.