5 ms·
> you'll describe something and you'll get a full 3D scene, with 3D models, source of lights set up, etc. I'm always confused why I don't hear more about proje
by epr 3y ago
> you'll describe something and you'll get a full 3D scene, with 3D models, source of lights set up, etc.
I'm always confused why I don't hear more about projects going in this direction. Controlnets are great, but there's still quite a lot of hallucination and other tiny mistakes that a skilled human would never make.
- boppo1 3y agoBlender files are dramatically more complex than any image format, which are basically all just 2D arrays of 3-value vectors. The blender filetype uses a weird DNA/RNA struct system that would probably require its own training run. More on the Blender file format: https://fossies.org/linux/blender/doc/blender_file_format/mystery_of_the_blend.html https://fossies.org/linux/blender/doc/blender_file_format/my...
- mikepurvis 3y agoBut surely you wouldn't try to emit that format directly, but rather some higher level scene description? Or even just a set of instructions for how to manipulate the UI to create the imagined scene?
- BirdieNZ 3y agoI've seen this but producing Python scripts that you run in Blender, e.g. https://www.youtube.com/watch?v=x60zHw_z4NM https://www.youtube.com/watch?v=x60zHw_z4NM (but I saw something marginally more impressive, not sure where though!)
- bsenftner 3y agoMy god that is an irritating video style, "AI woweee!"
- deleted 3y ago[deleted]
- mikebelanger 3y agoYeah I'd imagine that's the best way. Lots of LLMs can generate workable Python code too, so code that jives with Blender's Python API doesn't seem like too much of a leap. The only trick is that there has to be enough Blender Python code to train the LLM on.
- arcticbull 3y agoMaybe something like OpenSCAD is a good middle ground. Procedural code-like format for specifying 3D objects that can then be converted and imported in Blender.
- lightedman 3y agoI tried all the AI stuff that I could on OpenSCAD. While it generates a lot of code that initially makes sense, when you use the code, you get a jumbled block.
- regularfry 3y agoThis. I think problem is that the LLMs really struggle with 3d scene understanding, so what you would need to do is generate code that generates code. But also I suspect there just isn't that much openscad code in the training data, and the semantics are different enough to python or any of the other languages that are well-represented that it struggles.
- numpad0 3y agoIt sure feels weird to me as well, that GenAI is always supposed to be end-to-end with everything done inside NN blackbox. No one seems to be doing image output as SVG or .ai.
- HammadB 3y agoThere is a fundamental disconnect between industry and academia here.
- maccard 3y agoOver the last 10 years of industry work, I'd say about 20% of my time has been format shifting, or parsing half baked undocumented formats that change when I'm not paying attention. That pretty much matches my experience working with NN's and LLM's
- metanonsense 3y agoImo the thinking is that whenever humans have tried to pre-process or feature-engineer a solution or tried to find clever priors in the past, massive self-supervised-learning enabled, coarsely architected, data-crunching NNs got better results in the end. So, many researchers / industry data scientists may just be disinclined to put effort into something that is doomed to be irrelevant in a few years. (And, of course, with every abstraction you will lose some information that may bear more importance than initially thought)
- fy20 3y agoThe way that website builders using GenAI work is they have a LLM generate the copy, then find a template that matches that and fill it out. This basically means the "visual creativity" part is done by a human, as the templates are made and reviewed by a human. LLMs are good at writing copy that sounds accurate and creative enough, and there are known techniques to improve that (such as generating an outline first, then generating each section separately). If you then give them a list of templates, and written examples of what they are used for, the LLM is able to pick one that's a suitable match. But this is all just probability, there's no real creativity here. Earlier this year I played around with trying to have GPT-3 directly output an SVG given a prompt for a simple design task (a poster for a school sports day), and the results were pretty bad. It was able to generate a syntantically coreect SVG, but the design was terrible. Think using #F00 and #0F0 as colours, placing elements outside the screen boundaries, layering elements so they are overlapping. This was before GPT-4, so it would be interesting to repeat that now. Given the success people are having with GPT-4V, I feel that it could just be a matter of needing to train a model to do this specific task.
- Keyframe 3y agoScene layouts, models and their attributes are a result of user input (ok and sometimes program output). One avenue to take there would be to train on input expecting an output. Like teaching a model to draw instead of generate images.. which in a sense we already did by broadly painting out silhouettes and then rendering details.
- deleted 3y ago[deleted]
- guyomes 3y agoVoxel files could be a simpler step for 3D images.
- deleted 3y ago[deleted]
- bozhark 3y agoOne was on the front page the other day, I’ll search for a link
- jowday 3y agoThere's a lot of issues with it, but perhaps the biggest is that there aren't just troves of easily scrapable and digestible 3D models lying around on the internet to train on top of like we have with text, images, and video. Almost all of the generative 3D models you see are actually generative image models that essentially (very crude simplification) perform something like photogrammetry to generate a 3D model - 'does this 3D object, rendered from 25 different views, match the text prompt as evaluated by this model trained on text-image pairs'? This is a shitty way to generate 3D models, and it's why they almost all look kind of malformed.
- sterlind 3y agoIf reinforcement learning were farther along, you could have it learn to reproduce scenes as 3D models. Each episode's task is to mimic an image, each step is a command mutating the scene (adding a polygon, or rotating the camera, etc.), and the reward signal is image similarity. You can even start by training it with synthetic data: generate small random scenes and make them increasingly sophisticated, then later switch over to trying to mimic images. You wouldn't need any models to learn from. But my intuition is that RL is still quite weak, and that the model would flounder after learning to mimic background color and placing a few spheres.
- skdotdan 3y agoDeepmind tried something similar in 2018 https://deepmind.google/discover/blog/learning-to-write-programs-that-generate-images/ https://deepmind.google/discover/blog/learning-to-write-prog...
- dragonwriter 3y ago> I'm always confused why I don't hear more about projects going in this direction. Probably because they aren't as advanced and the demos aren't as impressive to nontechnical audiences who don't understand the implications: there’s lots of work on text-to-3d-model generation, and even plugins for some stable diffusion UIs (e.g., MotionDiff for ComyUI.)
- lairv 3y agoI think the bottleneck is data For single 3D object the biggest dataset is ObjaverseXL with 10M samples For full 3D scenes you could at best get ~1000 scenes with datasets like ScanNet I guess Text2Image models are trained on datasets with 5 billion samples
- bsenftner 3y agoOh, I don't know about that. Working in feature film animation, studios have gargantuan model libraries from current and past projects, with a good number (over half) never used by a production but created as part of some production's world building. Plus, generative modeling has been very popular for quite a few years. I don't think getting more 3D models then they could use is a real issue for anyone serious.
- senseiV 3y agoWhere can you find those? I'm in the same situation as him, I've never heard of a 3d dataset better than objaverse XL. Got a public dataset?
- bsenftner 3y agoThese are not public datasets, but with some social engineering I bet one could get access. I've not worked in VFX for a while, but when I did the modeling departments at multiple studios had giant libraries of completed geometries for every project they ever did, plus even larger libraries of all the pieces and parts they use as generic lego geometry whenever they need something new. Every 3D modeler I know has their own personal libraries of things they'd made as well as their own "lego sets" of pieces and parts and generative geometry tools they use when making new things. Now this is just a guess, but do you know anyone going through one of those video game schools? I wager the schools have big model libraries for the students as well. Hell, I bet Ringling and Sheridan (the two Harvards of Animation) have colossally sized model libraries for use by their students. Contact them.
- insanitybit 3y agoI assume because it's still extremely early.
- eigenvalue 3y agoI think this recent Gaussian Splatting technique could end up working really well for generative models, at least once there is a big corpus of high quality scenes to train on. Seems almost ideal for the task because it gets photorealistic results from any angle, but in a sparse, data efficient way, and it doesn’t require a separate rendering pipeline.
- sanitycheck 3y agoFrom my very clueless perspective, it seems very possible to train an AI to use Blender to create images in a mostly unsupervised way. So we could have something to convert AI-generated image output into 3D scenes without having to explicitly train the "creative" AI for that. Probably much more viable, because the quantity of 3D models out in the wild is far far lower than that of bitmap images.