13 ms·
Stability.ai – Introducing Stable Video 3D
- throwaway743 3y agoAnyone know of anything that'll auto rig/add weights?
- ImHereToVote 3y agoThere are numerous tools that auto-rig humanoid figures. The obvious one: https://www.mixamo.com/#/ https://www.mixamo.com/#/
- Filligree 3y agoIf the animations shown are representative, then the mesh output may very well be good enough to use in a 3d printer. Looking forward to experimenting with this.
- neom 3y agoI don't know much about 3D printing, would be very interested in learning more about this idea if you'd be so kind as to expand on it. Could I have AI spend all day auto scanning what teens are doing on instagram, auto generate toys based on it, auto generate advertisements for the toys, auto 3D print on demand?
- SirSourdough 3y agoHypothetically, sure, assuming the parent comment that these meshes are sufficient for modelling is correct and that you can find any teens who want a non-digital toy. I think a good hobbyist application for this would be something like modelling figurines for games, which is already a pretty popular 3D printing application. This would allow people with limited modelling skills to bring fantastical, unique characters to life “easily”.
- Filligree 3y agoPretty much. We're already generating images of monsters and characters for a D&D campaign; being able to print those in 3D would be pretty amazing.
- CobrastanJorji 3y agoI think their suggestion was more "I have a photo of a cool horse, and now I would like a 3D model of that same horse." Another way of looking at it, 3D artists often begin projects by taking reference images of their subject from multiple angles, then very manually turning that into a 3D model. That step could potentially be greatly sped up with an algorithm like this one. The artist could (hopefully) then focus on cleanup, rigging, etc, and have a quality asset in significantly less time.
- bobba27 3y agoThe question is whether this actually "creates a 3d model based on the picture", or if it "finds an existing model that looks similar to the picture and texture map it".
- maicro 3y agoOP is suggesting that this (AI model? I honestly am behind on the terminology) could replace one of the common steps of 3D printing - specifically, the step where you create a digital representation of the physical object you would want to end up with. There are other steps to 3D printing in general, though; a super rough outline: - Model generation - "Slicing" - processing the 3D model into instructions that the 3D printer can handle, as well as adding any support structures or other modifications to make it printable - Printing - the actual printing process - Post-processing - depending on the 3D printing technology used, the desired resulting product, and the specific model/slicing settings, this can be as simple as "remove from bed and use" to "carefully snip off support structures, let cure in a UV chamber for X minutes, sand and fill, then paint" As I said before, this AI model specifically would cover 3D model generation. If you were to use a printing technology that doesn't require support structures, and handles color directly in the printing process (I think powder bed fusion is the only real option here?), the entire process should be fairly automatable - a human might be needed to remove the part from the printer, but there might not be much post-processing to do. The rest of your desired workflow is a bit more nebulous - I don't know how you would handle "scanning what teens are doing on instagram", at least in a way that would let you generate toys from the information; generating and posting the advertisement shouldn't be too hard - have a standardish template that you fill in with a render from the model, and the description; printing on demand again is possible, though you'll likely need a human to remove the part, check it for quality and ship it. You could automate the latter, but that would probably be more trouble than it's worth.
- neom 3y agoInteresting, to be clear I don't think this is a good idea and it's kinda my nightmare post capitalism hell. I just think it's interesting this could be done now. On finding out what teens want, that part is somewhat easy-ish, I guess you'd need a couple of agents, one that is scanning teen blogs for stories and then converting them to key words, then another agent that takes the key words (#taylorswift #HaileyBieberChiaPudding #latestkdrama etc) into Instagram, after a while your recommend page will turn into a pretty accurate representation of what teens are into, then just have an agent look at those images and generate difs of them. I doubt it would work for a bunch of reasons, but it's an interesting thought experiment! Thanks!
- jsheard 3y agoWith previous attempts at this problem the shaded examples could be quite misleading because details that appeared to be geometric were actually just painted over the surface as part of the texture, so when you took that texture away you just had a melted looking blob with nowhere near as much detail as you thought. I'd reserve judgement until we see some unshaded meshes. What they show in the demo: https://i.imgur.com/9bZNTcd.jpeg https://i.imgur.com/9bZNTcd.jpeg What comes out of the 3D printer: https://i.imgur.com/MZrzsfh.png https://i.imgur.com/MZrzsfh.png
- SV_BubbleTime 3y agoIt’s always been this. None of these ever show the untextured model. When I see a demo where they are showing wireframes I know it’ll be good enough.
- jsheard 3y agoSeems like a tougher nut to crack than image generation was, since there isn't a bajillion high quality 3D models lying around on the internet to use as training data, everyone is trying to do 3D model generation as a second-order system using images as the training data again. The things that make 3D assets good, the tiny geometric details that are hard to infer without many input views of the same object, the quality of the mesh topology and UV mapping, rigging and skinning for animation, reducing materials down to PBR channels that can be fed into a renderer and so on aren't represented in the input training data, so the model is expected to make far more logical leaps than image generators do.
- refulgentis 3y agoIt almost seems easier, in that you have an arbitrary # of real world objects to scan and the hardware is heavily commoditized (IIRC iPhones have this built in at highres now?)
- polygamous_bat 3y agoHow is building a dataset easier than using a prebuilt dataset?
- ionwake 3y agoIm sorry for dumb lazy question. But would the input require more than one image? Is there a demo url to test this? I think it might jsut be time to buy a 3d printer. EDIT> Does "single image inputs" mean more than one image?
- kylebenzle 3y agoSingle image means one image.
- dartos 3y agoCan confirm the word single means 1
- ionwake 3y agolol cmon guys don't be too hard on me it does say "inputs"
- stavros 3y agoI do see how "single image inputs" can be conflated with "multiple inputs of a single image each time", as opposed to "video".
- ionwake 3y agoTBH I always look at the worst case scenario. I was worried it meant it need 3 images inputted as a single image at direct steps of the process, so requiring different angles. I wasn't sure, but thought best to check. I feel like it would have been clearer to have said something like " generates a 3d models from a single image". ( not exact wording but you catch my drift ). Sorry I am over analysing but all feedback is good right?
- ganeshkrishnan 3y agoDescribe in single words only the good things that come into your mind about... your mother.
- airstrike 3y agothat demo animation is so clever and satisfying
- amelius 3y agoBut it doesn't look very realistic, tbh.
- dreadlordbone 3y agoit doesn't break Euclidian space at least
- deleted 3y ago[deleted]
- itsgrimetime 3y agoI can’t get them to play
- ddtaylor 3y agoDoes anyone know what hardware inference can run on or memory requirements?
- Mathnerd314 3y agoIn the repo the model weights file is 9.37GB, whereas sdxl turbo is 13.9GB, and I don't see any mention of huge context windows, so probably it just needs a decent graphics card.
- kouteiheika 3y agoIt crashes with an out-of-memory error on my 24GB 4090, so at least when it comes to their sample script the answer is "a lot". Maybe it's just an inefficient implementation though.
- dragonwriter 3y agoPretty much every initial Stability release has been inefficient and has resources drop a lot when optimized for real consumer hardware community engines appeared for running the model. OTOH, with their shift to a less open licensing structure, community tooling probably won’t emerge with the same level of energy.
- canadiantim 3y agoI can't wait until we can use something like this for architectural design
- whywhywhywhy 3y agoSDXL+Controlnet and then feeding it just blocked out depth maps are probably more useful for that.
- andybak 3y agoSomething like https://dust3r.europe.naverlabs.com/ https://dust3r.europe.naverlabs.com/ might be more appropriate?
- kouteiheika 3y agoJust tried to run this using their sample script on my 4090 (which has 24GB of VRAM). It ran for a little over 1 minute and crashed with an out-of-memory error. I tried both SV3D_u and SV3D_p models. [edit]Managed to generate by tweaking the script to generate less frames simultaneously. 19.5GB peak VRAM usage, 1 min 25 secs to generate at 225 watts.[/edit]
- ganeshkrishnan 3y ago4090 is in weird spot. High speed but low RAM. Theoretically everything should run in ai but practically nothing runs
- LoganDark 3y ago4090 has more VRAM than most computers have system RAM. Surprised this is considered "low RAM" in any way except for relative to datacenter cards and top-spec ASi.
- samplatt 3y agoYou're comparing RAM amounts to other RAM amounts without considering requirements. 24GB is more than (most) current games would ever require, but is considered a uncomfortably-constrictive minimum for most industrial work. Traditional CPU-bound physics/simulation models have typically wanted all the RAM they could get; the more RAM the more accurate the model. The same is true for AI models. I can max out 24GB just using spreadsheets and databases, let alone my 3D work or anything computational.
- LoganDark 3y agoGood to know. I've only been running LLMs at home and most of the open-source ones have been more than small enough to fit in my measly 12GB. But I guess most workloads want so much that 24GB won't fit them at all.
- jokethrowaway 3y ago
- bugbuddy 3y ago[flagged]
- hansonpeter 3y ago[dead]
- londons_explore 3y agoAll the examples resemble plastic children's toys... How would it handle other objects? (People, fabrics, buildings, plants, mountains, mechanical parts, etc)
- programjames 3y agoIt's hard to get camera position tracking for random objects, so it looks like they used simulations. There's probably a lot more plastic children's toy models in Blender than people, fabrics, buildings, &c.
- issung 3y ago> Stable Video 3D (SV3D) is a generative model based on Stable Video Diffusion that takes in a still image of an object as a conditioning frame, and generates an orbital video of that object. So can it actually output a 3d model? Or just images of what it thinks the object would look like from other angles?
- krebby 3y agoThe reference video (https://youtu.be/Zqw4-1LcfWg https://youtu.be/Zqw4-1LcfWg) says they use a NeRF / structure from motion and then create a mesh with marching cubes from the generated radiance field. This is how most soa text-to-object generators work now as well
- 2StepsOutOfLine 3y agoI'm also struggling to find any examples of how to actually get a 3D model output. Very few references to this capability outside of the blog post.
- dubin 3y agoI'd like to play around with something like this, but from my understanding my machine (Macbook, 2021 M1) isn't nearly powerful enough (right?). Are there remote/cloud environments where I can run models like this?
- ilaksh 3y agoI suggest just using Stability's API. You aren't allowed to use it locally for commercial use anyway. You could set something up on RunPod or AWS, but I doubt it's worth the effort.
- dubin 3y agoAwesome, thank you! It does look like SV3D is not a part of the API currently, but only a matter of time I imagine.
- thrdbndndn 3y agoThe emphasis here is Single Image, but can this model generate with multiple images too? We know that a single image of an object physically can't cover all the sides of it, so it's all guesswork in AI. This is totally fine for certain scenario, but in lots of other cases, it's trivial to have multiple images of the same object, and if that offers higher fidelity, it's totally worth it. I'm aware there are many algorithms or AI models that already do that. I'm asking about Stability's one specifically because if they have impressive Single Image result, surely their multi-image results would also be much better than state-of-the-art?
- pksebben 3y agoIf it's not there yet, I'm willing to bet it will be soon enough given folks hacking it apart and injecting their own solutions.
- leesec 3y agoI wonder when Emad will be outed as a fed or a fraud. He's certainly leaving a trail of nasty behavior in the industry.
- TibbityFlanders 3y ago[dead]
- dheera 3y agoThey compare against Zero123-XL, but they should compare against MVDream instead. MVDream is quite good. If you fiddle with the loss you can get even better results.
- abdellah123 3y agoDid you write the blog post using AI ?
- nbzso 3y agoBillions purred into technology with minimal use case application. What is the direct implication of this tech? Porn on demand?
- ecronant 3y agoAvante-garde/experimental film making is the main benefactor of all this. Basically, cool looking video no one watches. I say that as a huge fanboy/artist myself. It is like Christmas every other day right now. All this VC money being set on fire to make better Avante-garde film tools is just wonderful. A dream come true.
- akanet 3y agoThere are many direct consequences of people being able to directly transform text into textured 3d models, and even vaster indirect consequences if one pauses to reflect. A tight feedback system with good cohesion could revolutionize, art, design, mechanical engineering, video games, etc, etc.
- GeoAtreides 3y agoExtracting object from an image, transforming the object (rotating, for example), re-blended it into the original image. Or making 3d game assets from objects you have around. Imagine: take your phone, go around town, into shops, into churches, come back, press a button, get huge library 3d assets to populate your game. Or, something like this, for IKEA: couple of photos of a room --> extract objects --> let user re-arrange furniture. The room could be either the user's room or an IKEA showroom. You can do it with existing tools, but this kind of technology reduces it to pressing a couple of buttons.
- Ultimatt 3y agoHow has no one drawn attention to this was science fiction in Enemy of the State in 1998 now a trivial reality https://www.youtube.com/watch?v=4AjLXZV46eE https://www.youtube.com/watch?v=4AjLXZV46eE