14 ms·
Show HN: I stripped DALL·E Mini to its bare essentials and converted it to Torch
- WheelsAtLarge 4y agoGood Job. What are the hardware requirements?
- deleted 4y ago[deleted]
- coding123 4y agoCloning....
- smcleod 4y agoAlmost worked on my 2021 Macbook Pro (M1 Pro) - https://github.com/kuprel/min-dalle/issues/1#issue-1286765237 https://github.com/kuprel/min-dalle/issues/1#issue-128676523...
- bart3r 4y agoI had exactly the same issue
- sam1r 4y agoSeems like there is a fix. https://github.com/kuprel/min-dalle/issues/1#issuecomment-1168228797 https://github.com/kuprel/min-dalle/issues/1#issuecomment-11...
- smcleod 4y agoWorking now!
- julien_c 4y agowhat's the inference time on M1?
- tomduncalf 4y agoRunning the "alien life" example from the README took 30 seconds on my M1 Max. I don't think it uses the GPU at all. I couldn't get the "mega" option to work, I got an error "TypeError: lax.dynamic_update_slice requires arguments to have the same dtypes, got float32, float16" (looks like a known issue https://github.com/kuprel/min-dalle/issues/2 https://github.com/kuprel/min-dalle/issues/2) Edit: installing flax 0.4.2 fixes this issue, thank all!
- klohto 4y agoThe thread now has a fix. As for the GPU, it's possible to get it working with some extra steps https://github.com/google/jax/issues/8074 https://github.com/google/jax/issues/8074 Macbook Pro M1 Pro numbers (CPU): python3 image_from_text.py --text='court sketch of godzilla on trial' --mega 640.24s user 179.30s system 544% cpu 2:30.39 total
- tomduncalf 4y agoPretty much identical on M1 Max python3 image_from_text.py --text='a comfy chair that looks like an avocado' 612.30s user 180.72s system 552% cpu 2:23.52 total
- tomduncalf 4y agoFrom reading that thread it didn't sound like GPU was fully supported yet, were you able to get it working?
- deleted 4y ago[deleted]
- NhanH 4y ago
- godmode2019 4y agoThat dinosaur image has fantastic meme potential
- fartcannon 4y agoSurly wandb is not a bare essential?
- Borgz 4y agoYou can download the models yourself if you don't want to use it.
- capableweb 4y agoWhere?
- chrisa 4y agoThere are links on the readme there, or you can run: > wandb artifact get --root=./pretrained/dalle_bart_mini dalle-mini/dalle-mini/mini-1:v0 and select option (3)
- jjallen 4y agoI wish this requirement was in the README
- deleted 4y ago[deleted]
- sydthrowaway 4y agoGamechanger Clone before kill.
- etaioinshrdlu 4y agoI don't think this is infringing on anyone's rights.
- sydthrowaway 4y agoopenai?
- ShamelessC 4y agoThis project is done by volunteers unrelated to open ai.
- c0n5pir4cy 4y agoThey've started changing the name in some places as well to avoid this kind of confusion - they've renamed the app to Craiyon as OpenAI have asked them to. (https://www.craiyon.com/#headlessui-disclosure-button-7 https://www.craiyon.com/#headlessui-disclosure-button-7)
- etaioinshrdlu 4y agoAnyone have some stats on inference time and RAM requirements? (on specific hardware)
- bart3r 4y agoI have a 2019 MacBook Pro 2.4Ghz Quad-core i5, 8GB RAM with Intel graphics card python3 image_from_text.py --text='a happy giraffe eating the world' --seed=7 154.61s user 22.18s system 262% cpu 1:07.40 total WARNING:absl:No GPU/TPU found, falling back to CPU. (Set TF_CPP_MIN_LOG_LEVEL=0 and rerun for more info.) As you can see, it took 1min 7seconds to complete. I assume it would be much faster with a grunty graphics card
- chrisa 4y agoUsing an rtx 3090 (NVIDIA gpu with 24GB of RAM): Mini = 5.33 s Mega = 14.7 s Update: about 1/2 that time is just loading the model, so if you load the model and then generate multiple images, it drops to: Mini = 3.91 s Mega = 8.86 s
- daenz 4y agoThis stuff is so cool and it makes me happy that we're democratizing artistic ability. But I can't help but think RIP to all of the freelance artists out there. As these models become more mainstream and more advanced, that industry is going to be decimated.
- mdp2021 4y agoNot differently to translators etc.: not required for every small task, still required for doing things professionally.
- mola 4y agoYeah, so a lot less work...
- mdp2021 4y agoI will rephrase it: if people today are available to eat dirt instead of nourishment - etc. for innumerable instances -, to get contented with lack of quality (with the akin acceptance of consequential decline of the general perception of quality), the fault is more in decadence than in instruments. You need well cultivated intelligence to obtain a good product: if "anything goes" is the motto, if "cheap" is the "mandate", there lies the issue.
- londons_explore 4y agoLots of professional translators moved into language tuition. I guess lots of artists will move into teaching art. I see tools like this might increase interest by the public into making their own art with the help of new tools, and some will want to be taught.
- jokethrowaway 4y agoTranslators at least have official documents (aka the only times I see a translator in my life across 3 countries) because the government is retarded and needs someone with a title to translate "Name" and "Surname" on a birth certificate. There is no equivalent for illustrators. My friends who studied some specific language are all unemployed or doing unqualified jobs. Their peers from a generation before are teachers or work in some embassy. That said, before some unicorn really start doing some serious polishing, you'll still want some illustrators to piece art together. Taking the output of these models won't deliver a ready made product easily.
- JanSt 4y agoWhat is the license of the generated artwork?
- witheld 4y agoYou pressed the button, it's owned by you, all rights are reserved by default.
- Dangeranger 4y agoThis statement is not accurate[0]. > But copyright law only protects “the fruits of intellectual labor” that “are founded in the creative powers of the [human] mind.” COMPENDIUM (THIRD) § 306 (quoting Trade-Mark Cases, 100 U.S. 82, 94 (1879)); see also COMPENDIUM (THIRD) § 313.2 (the Office will not register works “produced by a machine or mere mechanical process” that operates “without any creative input or intervention from a human author” because, under the statute, “a work must be created by a human being”). So Thaler must either provide evidence that the Work is the product of human authorship or convince the Office to depart from a century of copyright jurisprudence. [0] https://www.copyright.gov/rulings-filings/review-board/docs/a-recent-entrance-to-paradise.pdf https://www.copyright.gov/rulings-filings/review-board/docs/...
- witheld 4y agoThis is literally irrelevant, this is about attempting to register art as owned BY the computer If you give the computer the instructions, such as “avocado chair”, the avocado chair is yours. It wouldn’t be yours if it was something like a deep dream- if you ran the program with no input and generated a “random” work.
- Dangeranger 4y agoPlease cite a source that confirms your claim, rather than stating it as a fact without evidence. You may be right, but I have no way of confirming that given the content of your comment. The report I’ve cited makes a compelling argument against your claim, and several prominent organization’s copyright policies align with it.
- amar-laksh 4y agoPeople might also like this one: https://github.com/saharmor/dalle-playground https://github.com/saharmor/dalle-playground Really easy to work with
- enlyth 4y agoI couldn't get this to pick up my graphics card when running it with WSL 2, it's just says no cuda devices found or something so I gave up, not sure if anyone had any luck
- lastdong 4y agoNow we just need to containarise it (there are a few docker python nvidia images)
- epicureanideal 4y agoFree idea: Same but for making short video clips, and then eventually producing entire movies.
- jcims 4y agoThe google collab link works if you replace the computed path to flax_model.msgpack on line 10 in load_params.py with ‘/content/pretrained/vqgan/flax_model.msgpack’ Edit: actually it's easier to open a terminal and move /content/pretrained/vqgan to /content/min-dalle/pretrained/vqgan
- kuprel 4y agoThanks for figuring this out. The problem was that the vqgan repository was being cloned to the wrong directory. The updated colab should work now
- DivineTraube 4y agoThe updated version gives me this (after successful setup with the example alien thing): UnfilteredStackTrace Traceback (most recent call last) <ipython-input-2-0e20e3adf861> in <module>() 2 ----> 3 image = generate_image_from_text("alien life", seed=7) 4 display(image) 67 frames UnfilteredStackTrace: TypeError: lax.dynamic_update_slice requires arguments to have the same dtypes, got float16, float32. The stack trace below excludes JAX-internal frames. The preceding is the original exception that occurred, unmodified. -------------------- The above exception was the direct cause of the following exception: TypeError Traceback (most recent call last) /content/min-dalle/min_dalle/models/dalle_bart_decoder_flax.py in __call__(self, decoder_state, keys_state, values_state, attention_mask, state_index) 38 keys_state, 39 self.k_proj(decoder_state).reshape(shape_split), ---> 40 state_index 41 ) 42 values_state = lax.dynamic_update_slice( TypeError: lax.dynamic_update_slice requires arguments to have the same dtypes, got float16, float32.
- jcims 4y agoYou need to install flax 0.4.2. If you're using collab you just open a terminal (icon in the bottom left of the screen) and run: pip3 install flax==0.4.2
- 4y ago
- deleted 4y ago[deleted]
- mg 4y agoWhat is the maximum resolution possible with this? If it depends on the hardware, what would be the limit when one rents the biggest machine available in the cloud?
- jwitthuhn 4y agoFixed size of 256x256. It cannot go any bigger or smaller.
- 01acheru 4y agoOut of curiosity: why it cannot be changed? I know nothing about this field so... thanks!
- freemint 4y agoBecause that is how the network is trained. You could modify the network size and retrain to get different resolutions.
- petercooper 4y agoIt's already being scaled up to 256x256 from something smaller anyway. You could add an extra upscaler to go further which I've tried with moderate success, but you're basically doing CSI style 'enhance' over and over.
- ShamelessC 4y agoTransformers output fixed-length sequences. For this transformer they chose 256 pixels, or 32 "image tokens" that each decode to an 8-by-8 pixel "patch". You can technically increase or decrease this - or use a different aspect ratio by using more or fewer image tokens, but this is static after you start training. It will also require more "decodes" from the backbone VQGAN model (responsible for converting pixels to image tokens), and thus take longer to run inference on. CLIP-guided VQGAN can get around this by taking the average CLIP score over multiple "cutouts" of the whole image allowing for a broad range of resolutions and aspect ratios.
- kertoip_1 4y agoDoes it just download pre-trained DALL-E Mini models and generate images using them? Because I can't seem to find any logic in that repo other than that. I'm not into that field, just curious if I'm missing something.
- ShamelessC 4y agoThey converted the original JAX weights to the format that Pytorch uses. Because JAX is still fairly new, it can be a lot easier to get Pytorch to run on e.g. CPU. I do find the number of upvotes interesting and I imagine many people just upvote things that have DALLE in the title, to a degree. Not to discourage the OP of course, great work.
- OJFord 4y agoLook how much easier it is to install & run, people are interested in and up-voting the result, not how much work was (or wasn't) required to achieve it.
- ShamelessC 4y agoFair enough - that makes sense, just haven't had a chance to dive into it yet.
- avhon1 4y agoIt still seems to require JAX somewhere to work. On my desktop, running the example > python image_from_text.py --text='alien life' --seed=7 results in > RuntimeError: This version of jaxlib was built using AVX instructions, which your CPU and/or operating system do not support. You may be able work around this issue by building jaxlib from source. Unfortunately, following the instructions to build JAXlib from source (https://jax.readthedocs.io/en/latest/developer.html#building-from-source https://jax.readthedocs.io/en/latest/developer.html#building...) result in several 404 not found errors, which later cause the build to stop when it tries to do something with the non-existent files. Unfortunately, it looks like I won't be running this today.
- jwitthuhn 4y agoLove this, before I only ever saw code that ran this model through jax. This seems to perform much better on my m1.
- pja 4y agoNB: Needs a weights & biases account in order to download the models.
- lars_francke 4y agoYou can CTRL+C that prompt and it'll download them anyway but it'll tell you that you can't visualize your results then.
- mikewarot 4y agoWhat is wandb.ai, and Why does it keep asking for an API key? It's not listed in the requirements I've posted it as an issue
- h0mie 4y agoSaas dashboard for monitoring/mertics
- gjs278 4y ago
- nl 4y ago> What is wandb.ai Weight & Biases > Why does it keep asking for an API key From the README: the Weight & Biases python package is used to download the DALL·E Mini and DALL·E Mega transformer models It might not be obvious you need an account if you aren't in the field though.
- Aardwolf 4y agoWhy is this needed to download the model? I'd prefer to download it myself and choose where I put it too. It now uses some hashed filename in some config directory in your homedir for this, I dislike this and want control over where I put models, make it more self contained instead of random directories spread all over your OS, and give them as input by file path. This feedback is about dalle mini playground instead but it does the same thing. If this one is stripped to bare essentials I'd expect this type of dependencies stripped too. Edit: I don't want to seem like complaining too much though and am very happy with these open models and tooling for them. Thanks!
- carnitine 4y agoIt’s not, it just makes it easier. Should be pretty simple to modify the code to work the way you want.
- deleted 4y ago
- gjs278 4y ago
- bjarneh 4y agoTypeError: lax.dynamic_update_slice requires arguments to have the same dtypes, got float32, float16.
- langitbiru 4y agoI guess we will have text-to-image startups in the next batch of YC.
- nudpiedo 4y agowandb forbids reusing the models and other information, independently of their usage, so they should find another source for their models EDIT: as I am being accused of inventing it I will quote the terms of agreement and license, since maybe its own founder seems to not have read it or someone without training on how to write proper terms and agreements made it for them and the restrictive usage of "Material" does apply to its hosted software. Note that there is no formal definition of "Materials" or "Service", so that it applies to all the contents of the webpage including the software stored there: https://wandb.ai/site/terms https://wandb.ai/site/terms I quote it: 2. Use License Whether you are accessing the Services for personal, non-commercial transitory viewing only (our free license for individuals only), for academic use, or for commercial purposes (our subscription package for businesses), permission is granted to temporarily download one copy of the information or software (the “Materials”) from our website. This is the grant of a license, not a transfer of title, and under this license, you may not: a. Modify or copy the Materials; b. Use the Materials for any commercial purpose, or for any public display (commercial or non-commercial); c. Attempt to decompile or reverse engineer any software contained in the Materials; d. Remove any copyright or other proprietary notations from the Materials; or e. Transfer the Materials to another person or "mirror" the Materials on any other server. This license shall automatically terminate if you violate any of these restrictions and may be terminated by us at any time. Upon terminating your viewing of these materials or upon the termination of this license, you must destroy any downloaded materials in your possession whether in electronic or printed format. f. Utilize our personal license for individuals for commercial purposes and any such use of our personal license for commercial purposes (e.g. using your corporate email) may result in immediate termination of your license.
- slewis 4y agoFounder of Weights & Biases here (wandb). We don’t forbid anything, models are property of the people who created them. Why do you think that? EDIT: I'll edit respond, since you did. Look at sections 3b and 3c in the terms, they cover Models and other user content specifically. Those are user property, not our property. But I can see how this is confusing. We will clarify it.
- JacobiX 4y agoInteresting when testing with inputs like "Oscar Wilde photo" or "marilyn monroe photo" and comparing to a Google image search. After some iterations we can have quite similar images but the faces are always blurry.
- deleted 4y ago[deleted]
- punk_ihaq 4y agoComment deleted
- buf 4y agoI always generates an image of a moon for me.
- mensetmanusman 4y agoThat’s a feature not a bug…
- pilotneko 4y agoI think there might be some session leakage. I typed “A pig with a bowler hat.” and the model returned a picture of a half moon.
- ubertaco 4y agoNo matter what I typed, it always generated the same image of a moon half-covered in shadow. I think something might be a bit buggy with this.
- sAbakumoff 4y agoThe results are amazingly poor. Try "biden plays chess against napoleon"
- tgv 4y agoI tried a few non-descriptive statements from random tweets. As it turns out, nobody's made a random tweet since 2016, but for the few that exist, the results are great. E.g. "Good Morning Everyone , Happy Nice Day :D" generates something that can only be described as bored-ape meets Picasso in kindergarten. Probably the next-gen 1M$ NFT. If anybody needs proof that these models don't think, this is it.
- Linda703 4y ago[dead]
- andybak 4y agoIf people just want to run text to image models locally by far the easiest way I've found on Windows is to install Visions of Chaos. It was originally a fractal image generation app but it's expanded over time and now has a fairly foolproof installer for all the models you're likely to have heard of (those that have been released anyway). https://softology.pro/tutorials/tensorflow/tensorflow.htm https://softology.pro/tutorials/tensorflow/tensorflow.htm
- leereeves 4y agoThank you, I've been looking for something like that and it looks very cool, judging by this tutorial that shows it in action, as it creates an image and displays the results in progress: https://youtu.be/4_LgrAL7EWg?t=163 https://youtu.be/4_LgrAL7EWg?t=163
- ccbccccbbcccbb 4y agonot really surprising, but <caveat emptor> bare minimum GPU: NVIDIA 2080 with 8GB VRAM 300 Gb of disk space </caveat emptor>
- leereeves 4y agoIt should probably have an option to download just one model instead of all of them. 300GB is a lot.
- ccbccccbbcccbb 4y agoAlas, the very first point was already a no-go for me
- shreyshnaccount 4y agowhere did you get these? couldnt find it, mightve missed smth
- ccbccccbbcccbb 4y ago[dead]
- ramesh31 4y agoHas anyone got this running on M1?
- ausbah 4y agohas anyone applied compression techniques to large models like dalle-2?
- dalle-world 4y agoI spun up an aws ubuntu ec2 with 2 Tesla M60. When I run python3 image_from_text.py --text='alien life' --seed=7 I get this error detokenizing image Traceback (most recent call last): File "/home/ubuntu/work/min-dalle/image_from_text.py", line 44, in <module> image = generate_image_from_text( File "/home/ubuntu/work/min-dalle/min_dalle/generate_image.py", line 74, in generate_image_from_text image = detokenize_torch(image_tokens) File "/home/ubuntu/work/min-dalle/min_dalle/min_dalle_torch.py", line 107, in detokenize_torch params = load_vqgan_torch_params(model_path) File "/home/ubuntu/work/min-dalle/min_dalle/load_params.py", line 11, in load_vqgan_torch_params params: Dict[str, numpy.ndarray] = serialization.msgpack_restore(f.read()) File "/usr/local/lib/python3.10/dist-packages/flax/serialization.py", line 350, in msgpack_restore state_dict = msgpack.unpackb( File "msgpack/_unpacker.pyx", line 201, in msgpack._cmsgpack.unpackb msgpack.exceptions.ExtraData: unpack(b) received extra data.
- dalle-world 4y ago
- pmarreck 4y agoI get a similar error running it locally (not sure if related, but it also can't find my GPU, which is a 3080ti and should be sufficient): Traceback (most recent call last): File "/home/pmarreck/Documents/min-dalle/image_from_text.py", line 44, in <module> image = generate_image_from_text( File "/home/pmarreck/Documents/min-dalle/min_dalle/generate_image.py", line 75, in generate_image_from_text image = detokenize_torch(torch.tensor(image_tokens)) File "/home/pmarreck/Documents/min-dalle/min_dalle/min_dalle_torch.py", line 108, in detokenize_torch params = load_vqgan_torch_params(model_path) File "/home/pmarreck/Documents/min-dalle/min_dalle/load_params.py", line 12, in load_vqgan_torch_params params: Dict[str, numpy.ndarray] = serialization.msgpack_restore(f.read()) File "/home/pmarreck/anaconda3/lib/python3.9/site-packages/flax/serialization.py", line 350, in msgpack_restore state_dict = msgpack.unpackb( File "msgpack/_unpacker.pyx", line 202, in msgpack._cmsgpack.unpackb msgpack.exceptions.ExtraData: unpack(b) received extra data.
- mahastore 4y agoWTF is Wandb.ai? This seems like a sneaky way to get people to sign up for this wandb thingy.