3 ms·
TextDiffuser-2: Unleashing the power of language models for text rendering
- opjjf 3y agoKind of crazy that they use Breath of the Wild even in the examples. Why can this be generated if not by obviously stealing Nintendo IP?
- Oras 3y agoThe problem, from the paper: > Several methods alleviated this issue by incorporating explicit text position and content as guidance on where and what text to render. However, these methods still suffer from several drawbacks, such as limited flexibility and automation, constrained capability of layout prediction, and restricted style diversity. Looking at the diagram provided, they use GPT-4 to suggest the position following the text prompt. I see it as very useful for making sure to have the text in the right position without doing manual work trying to find the right position. I'm not an expert, but doesn't this method add another cost and overhead for calling Text-to-Image models?
- alextheparrot 3y agoLLM + Text-to-Image model is exactly how DALL·E 3 is deployed, fwiw
- zaptrem 3y agoIncluding the text positioning generation part? What’s the source on that?
- alextheparrot 3y agoThe comment was directed at “doesn't this method add another cost and overhead for calling Text-to-Image models”
- wahnfrieden 3y agoNo
- blixt 3y agoIt’s very smart, though using bounding boxes will most likely limit it to 2D contexts (and some head-on 3D contexts) since the text won’t follow the bounding box when perspective is involved. I’m sure it can be improved to support bounds that have 3D transforms though.
- marban 3y agoRecent comparison of what's out there: https://www.reddit.com/r/StableDiffusion/comments/18o1ole/apparently_not_even_midjourney_v6_launched_today/ https://www.reddit.com/r/StableDiffusion/comments/18o1ole/ap...
- whywhywhywhy 3y agoDoes Midjourney v6 use something similar to this because they both have a weird look to the text like amateurishly photoshopped look where it’s almost has different aliasing to the rest of the image looking not truly integrated. Impressive it’s legible but some work is needed to get it to normal production quality.
- grork 3y agoI’m assuming the type foundry legal departments are getting ready to come for the image generators when they find out their typefaces have been vacuumed up and are now generating new content without licensing the typeface for use?
- kelseyfrog 3y ago> In the United States, the shapes of typefaces are not eligible for copyright but may be protected by design patent. In the US only the font files themselves are copyrightable. It would be a difficult case to make that the model weights, or output bears any copyright-violating relation to the font files.