3 ms·
Have you seen Open AI's DALL•E? https://openai.com/blog/dall-e/ https://openai.com/blog/dall-e/ Kind of does something similar to what you describe, right? Thi
by ckosidows 6y ago
Have you seen Open AI's DALL•E? https://openai.com/blog/dall-e/ https://openai.com/blog/dall-e/ Kind of does something similar to what you describe, right?
This project blows my mind, by the way. I still think it's the coolest thing to come out of AI research so far.
- chrisweekly 6y agoWOW!! I'm astounded. Zero-shot visual reasoning?? As an emergent capability?? Um. wat. >"GPT-3 can be instructed to perform many kinds of tasks solely from a description and a cue to generate the answer supplied in its prompt, without any additional training. For example, when prompted with the phrase “here is the sentence ‘a person walking his dog in the park’ translated into French:”, GPT-3 answers “un homme qui promène son chien dans le parc.” This capability is called zero-shot reasoning. We find that DALL·E extends this capability to the visual domain, and is able to perform several kinds of image-to-image translation tasks when prompted in the right way."
- tablespoon 6y ago> Have you seen Open AI's DALL•E? https://openai.com/blog/dall-e/ https://openai.com/blog/dall-e/ Is there any way to try it out easily? Everything on that page looked like it was pre-rendered (constrained-choice mad libs).
- ckosidows 6y agoI don't think so. It does appear pre-rendered and my guess is the blog might only showcase the use cases that worked the best. Or it might take a really long time to generate results.
- drdeca 6y agoIt is also my understanding that it as a whole is not released. However, iirc one of its two main components, has a smaller version of it released. This component is called CLIP . Given some text options and an image, CLIP will predict/guess which of the texts is a description of the image. It handles images in various art styles well? See : https://openai.com/blog/clip/ https://openai.com/blog/clip/
- Retric 6y agoVery cool and interesting that the green pentagon clock had several images of non pentagons. It’s getting the general flat sides correct, but not counting the number of vertices.
- ummwhat 6y agoI feel like that's true to how the average person thinks. My aunt told me she got octagonal tiles for her bathroom. I said you can't tile octagons in two dimensions. After a short argument she showed me a picture. They were hexagons.
- jonplackett 6y agoThis is amazing! The reasoning required for making the building blocks image is really impressive. It's funny isn't it how we're so amazed at these abilities, but even a fairly slow-witted child could do all this stuff.
- danielbarla 6y agoThat's insanely cool. I wonder, if similarly, it could be set up to pump out 3D models of things. We'd be pretty close to having an "AI storyteller" for games, movies, etc - and one that can improvise!