3 ms·
Google claims to have a new DALL-E like model that does much better in precise instruction, due to using text models in different ways: https://imagen.research.
by native_samples 4y ago
Google claims to have a new DALL-E like model that does much better in precise instruction, due to using text models in different ways: https://imagen.research.google/ https://imagen.research.google/
Unfortunately the impact of Imagen will be low like all the other Google models because, as keeps happening, they don't trust people enough to actually release it, not even in demo form:
"Imagen relies on text encoders trained on uncurated web-scale data, and thus inherits the social biases and limitations of large language models. As such, there is a risk that Imagen has encoded harmful stereotypes and representations, which guides our decision to not release Imagen for public use without further safeguards in place."
This seems to be emerging as a clear pattern - OpenAI trains an AI and then opens it to a group of select invitees (the usual suspects) with lots of opaque T&Cs that restrict what they're allowed to say about it [1]. For example, OpenAI forbids users from sharing any outputs that contain realistic faces even though the AI generates such pictures all the time, even when not requested [2]. This is reminiscent of how RDBMS vendors forbid users from publishing benchmarks and usually results in people giving OpenAI some stick for not living up to their name.
But OpenAI is still light years more open than Google, which routinely announces just months later that they've trained an AI far better than anything OpenAI produced, but doesn't provide any evidence beyond their own cherry picked examples. The justification given is always the same: the AI has learned un-woke things like the fact that certain jobs tend to be done by specific genders, and would generate examples based on that learnings. And because the Google researchers are crazy they think that seeing pictures of white male builders or young female nurses would be actually dangerous to the public, so they just don't provide any API access at all, not even to little select lists of friends.
IMO this is turning into a serious problem for AI research. It's already notoriously non-reproducible, but OpenAI's work is at least auditable. People can do experiments on DALL-E and see that it's real for themselves, explore the limits and come up with ways to handle them, as is happening here. They could in future build apps that use them. In contrast Google, who should by all rights be at the forefront of this space, keeps getting eclipsed by OpenAI again and again because they're now so woke that they've retreated into a tiny little purity bubble. They claim that they'll make their AIs available when they figured out how to brainwash them to have the right views, but they were claiming this is an "open problem" for years and never seem to release anything. So all the talk is about GPT-3 and DALL-E and Google's equivalents just get forgotten.
[1] https://www.lesswrong.com/posts/uKp6tBFStnsvrot5t/what-dall-e-2-can-and-cannot-do?commentId=vAxnyEJFBKMeA26a2 https://www.lesswrong.com/posts/uKp6tBFStnsvrot5t/what-dall-...
[2] https://www.lesswrong.com/posts/uKp6tBFStnsvrot5t/what-dall-e-2-can-and-cannot-do#Realistic_human_faces https://www.lesswrong.com/posts/uKp6tBFStnsvrot5t/what-dall-...
- rjh29 4y agoIs it non-reproducible though? Google did not release their model but they did publish their method, which they didn't have to do. People like lucidreams are already at work replicating them so everyone can benefit: https://github.com/lucidrains/DALLE2-pytorch https://github.com/lucidrains/DALLE2-pytorch
- native_samples 4y agoYes, that's absolutely a good point worth repeating - Google is obviously under no obligation to publish anything or make any demos available. They can keep their research entirely proprietary if they want. There is maybe a question of whether this is what their researchers imagined, as many were brought to Google partly by the promise of being able to do their research in the open. But that's their problem. The link you posted seems to reinforce my point though. It's a reimplementation of DALL-E, not anything done by Google. That seems to be how it goes, everyone tries to duplicate the OpenAI work even though Google's models are (or claim to be) more advanced. The mindshare value of making their demos available appears to be phenomenal. Is it reproducible - well, only for flexible definitions of reproducible. If you try to follow the method in the paper and don't get results as good, is the issue the method or that you didn't follow it well enough? You can't follow it all that closely because they often aren't detailed enough and the training sets are constantly changing. The issue here (for Google) is a bit different. If they were keeping their research proprietary because they wanted to turn it into a product and sell it, that's one thing and totally understandable. It's very expensive and needs some way to financially justify the cost. Their justification is quite different though and raises questions about their whole AI initiative. What's the point of creating SOTA AI trained on the internet if you're afraid of what people will use it for? Their efforts to make ideologically acceptable AI don't seem to have worked yet, nor has OpenAI been embarrassed by abuse. So OpenAI is powering ahead here and everyone talks about their models, whilst Google's languish in obscurity. Like, Scott Alexander is talking about the limits of DALL-E 2 that Google claim they already solved, but it's irrelevant because Scott can play with DALL-E and not Imagen.