3 ms·DiffusionBlocks: Training Neural Networks One Block at a Time4 points by sebg 4mo agobillconan 4mo agoI do not understand. how is this different from building smaller transformer layers, and each layer just denoises less?