3 ms·
Is this going to end up into a single model, where its trained on text and images and audio and videos and 3d models, and it can do anything to anything dependi
by ItsMonkk 4y ago
Is this going to end up into a single model, where its trained on text and images and audio and videos and 3d models, and it can do anything to anything depending on what you ask of it? Feels like the cross-training would help yield stronger results.
- minimaxir 4y agoThese diffusion models are using a frozen text encoder (e.g. CLIP for Stable Diffusion, T5 for Imagen), which can be used in other applications. StabilityAI trained a new/better CLIP for the purpose of better Stable Diffusions.
- CuriouslyC 4y agoProbably not. We're actually headed towards many smaller models that call each other, because VRAM is the limiting factor in application, and if the domains aren't totally dependent on each other it's easier to have one model produce bad output, then detect that bad output and feed it into another model that cleans up the problem (like fixing faces in stable diffusion output). The human brain is modularized like this, so I don't think it'll be a limitation.
- deleted 4y ago[deleted]