3 ms·
I don't understand what this does. Examples mention "human prompt" but I don't see it anywhere.
by MitPitt 4y ago
I don't understand what this does. Examples mention "human prompt" but I don't see it anywhere.
- okamiueru 4y agoThe readme.md has some image examples. If I interpret them correctly, the first one is the input image, the second the segmented output, and the rest are example outputs using the prompt text shown above the collage. The visual quality of the output images is not particularly impressive compared to what we've become used to. What (IMO) it attempts to showcase is how the input image segmentation is used to guide the final image generation. That part is quite impressive. The shapes, and "segments" are very well preserved from input to output.
- sp332 4y ago"The human prompt and BLIP2 generated prompt build the text instruction." The examples only show the BLIP2 prompt.
- GaggiX 4y agoIt's a controlnet model trained on SAM segmentation maps. The final model takes a prompt and a segmentation map as input and generates an image conditioned on them.
- Sharlin 4y agoAnd I thought I was moderately "with it" with regard to AI… Could I get that in ELI5?
- GaggiX 4y agohttps://github.com/lllyasviel/ControlNet https://github.com/lllyasviel/ControlNet Controlnet is a neural network added to an already trained model so they can be conditioned on new stuff like canny edge, depth map, segmentation map. Controlnet let you train this model on the new condition "easily", without catastrophic forgetting and without a huge dataset. In the repo linked by OP, they have trained a controlnet model on the segmentation map generated by SAM: https://segment-anything.com/ https://segment-anything.com/