Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
erwannmillon
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
erwannmillon
1y ago
hahahah i'm doing well tianpei good to hear from you!
2.
▲
by
erwannmillon
1y ago
yeah data really is everything that was the number one lesson from this whole project
3.
▲
by
erwannmillon
1y ago
yoo i'm also a researcher on the krea 1 project and happy to answer any questions :)
4.
▲
by
erwannmillon
3y ago
LCMs are actually exactly SD architecture. LCM is initialized from a regular SD unet and finetuned on a new objective. We are already compiling to get to these times. A lot of other people getting sub-100ms times are using fewer inference s
5.
▲
by
erwannmillon
3y ago
i worked on the gpu/infra side of this, so feel free to AMA Ultimately the LCM is just a SD Unet trained with a new objective, so a lot of SD optimizations are transferable to LCMs
6.
▲
by
erwannmillon
3y ago
I actually worked on this feature at krea, happy to answer any technical questions about how this works / is trained
7.
▲
by
erwannmillon
3y ago
In the decoder, the features from the unet blocks get concatenated with features from the encoder layer through 'skip connections'. The paper discusses how rescaling the backbone features (element-wise multiplication by some scala
8.
▲
Improve Stable Diffusion Quality with Skip Connection Rescaling (No Training)
(arxiv.org)
2 points
by
erwannmillon
3y ago
|
4 comments
9.
▲
by
erwannmillon
3y ago
Improve SD image quality and reduce artefacts without any additional training, simply by reweighting skip connections in the decoder stage of a diffusion Unet decoder
10.
▲
by
erwannmillon
3y ago
"Native refiner swap inside one single k-sampler. The advantage is that now the refiner model can reuse the base model's momentum (or ODE's history parameters) collected from k-sampling to achieve more coherent sampling. In A
11.
▲
by
erwannmillon
3y ago
I work at krea.ai. We are for sure making art extremely accessible, but we consider it enhancing creativity rather than replacing. I fully agree that being able to generate an aesthetically pleasing image with an AI that has been optimized
12.
▲
by
erwannmillon
3y ago
So we're using a color space that has two channels dedicated entirely to color, which is the only thing the model needs to learn. The model doesn't need to touch the lightness channel at all, only predict the noised added to the c
13.
▲
by
erwannmillon
3y ago
Think inference time was on the order of 4-5seconds per image on a v100, which you can rent for like .80 cents an hour, though you can get way better gpus like a100s for ~1.1 usd/h now. But ofc this is at 64px res in pixel space. If yo
14.
▲
by
erwannmillon
3y ago
Yeah, if you have a high res image, you can get color info at super low-res and then regenerate the colors at high res with another model. (though this isn't an efficient approach at all) https://github.com/TencentARC&#
15.
▲
by
erwannmillon
3y ago
Depends, given the low res, the 3x64x64 pixel space image is smaller than the latents you would get from encoding a higher-res image with models like VQGAN or the stablediff VAE at their native resolutions. It's easier to get a sense o
16.
▲
by
erwannmillon
3y ago
Think the final training run was only a couple hours on a Colab V100
17.
▲
by
erwannmillon
3y ago
Took a lot of failed experiments, the model would keep converging to greyscale / sepia images. Think one of the ways I fixed was by adding an greyscale encoder to the arch. Used its output embedding as additional conditioning. Can'
18.
▲
by
erwannmillon
3y ago
Btw, I did this in pixel space for simplicity, cool animations, and compute costs. Would be really interesting to do this as an LDM (though of course you can't really do the LAB color space thing, unless you maybe train an AE specifica
19.
▲
by
erwannmillon
3y ago
Fair enough. Honestly this was just a fun side project. I actually coded this up last october when I was doing a deep dive to learn about diffusion models, and saw that no one had ever applied them to colorization. This was just a fun oppor
20.
▲
by
erwannmillon
3y ago
temporal coherence is def an issue with these types of models, though I haven't tested it out with ColorDiffusion. Assuming you're not doing anything autoregressive (from frame to frame) to do temporal coherence, you can also para
21.
▲
by
erwannmillon
3y ago
Technically yes, the encoder and unet are convolutional and support arbitrary input sizes, but the model was trained at 64x64px bc of compute limitations. You could probably resume the training from a 64x64 resolution checkpoint and train a
22.
▲
by
erwannmillon
3y ago
Yeah, the model is racist for sure. That's a limitation of the dataset though (celeb A is not known for its diversity, but it was easy for me to work with, I trained this model on Colab) And plausibility is a feauture, not a bug. There
23.
▲
by
erwannmillon
3y ago
hahaha it reminded me of some "zoom and enhance" stuff when I was making the animations
24.
▲
by
erwannmillon
3y ago
You can do this with spatial palette t2i or controlnet. Give a super lores spatial palette as conditioning like this: https://camo.githubusercontent.com/8e488996fd309165fb065b0cd... https://github.com/Tencen
25.
▲
by
erwannmillon
3y ago
touché, nevertheless, colors go brrrrrrrr
26.
▲
by
erwannmillon
3y ago
basically the training works as follows: Take a color image in RGB. Convert it to LAB. This is an alternative color space where the first channel is a greyscale image, and two channels that represent the color information. In a traditional
27.
▲
by
erwannmillon
3y ago
trained on celebA, so no, but you could for sure train this on a more varied dataset