4 ms·
> not very useful for most work We seem to be on a timeline where most of the significant use cases that the model doesn't handle well today is less than 2 yea
by erickj 3y ago
> not very useful for most work
We seem to be on a timeline where most of the significant use cases that the model doesn't handle well today is less than 2 years away from significant improvement.
My (completely baseless) guess is that within 2 years we begin to see "high budget" feature length productions beginning to move towards a cost saving model which fully allocate the production budget to primarily virtual content.
In less than a few years time there will almost certainly be a vast ecosystem of production and post production tools to give creators the controls to reliably create and fine tune their shots.
- renegade-otter 3y agoEven image generators are still in that phase where they are excellent at generating faces and dream-like sequences but suck at details. Pair that with the increasing legal copyright headwinds in terms of sucking in the world's data, and I can see this flatlining. All that processing power is expensive, and these companies are yet to have a clear path to monetization. Stability AI is already circling the drain. This has the "let's back the cash truck up to lock in the market share" of the streaming wars vibes, until they realized it's just not sustainable.
- hexage1814 3y agoI agree with you, and just a few more observations about where do I think the current bottleneck might be: I wonder how well the model handles with re-using objects/people/scenes. Like, can I create a character and then use him again along 10 different shots? Also, I'm pretty curious about how the user interface looks like. Cause they the text-to-video model interfaces seem pretty limited compared to the freedom a person has using Unreal Engine or Blender or shooting a movie in real life. How would the golden standard text-to-video user interface would look? And I have been thinking on this for years, even before the current generative AI boom, and I wonder if it could generate like a 3D representation of the scene that you described, like there would be a file where you could very easily change things around, as if that thing had been created on Blender or whatever, but very very user-friendly and easy to edit things. It will seem silly what I'm going to say, but the ideal interface, it reminds those movies people did using the game "The Sims", and how you could very easily move objects, and move the camera, and so on. What I'm trying to say here is that I would imagine these models creating a 3D representation of the scene, and the movie-making process ends up being somewhat similar to how could you could customize objects/people in that game.
- TomaszZielinski 3y agoI have only vague idea about this (I worked on small 3d games many many years ago), but I imagined something similar to what you described. Basically you use Sora to generate a promising scene, then you ask it (or another model) to turn that scene into a scene graph in a text file. It will make mistakes, but it could work similarly to the Python interpreter in ChatGPT--it can iterate until everything is OK. Maybe there could even be some adversarial stuff where the scene graph is rendered on the fly to compare it to the generated clip, etc. And then you can use you standard toolset to edit it, probably enhanced with a copilot model to automate as much as possible.
- whiplash451 3y agoThe cool demos from OpenAI, Figure and the like make us hallucinate a future that will take much (much) longer to pan out because they ignore the domain-specific knowledge that is inherent to the domain they pretend to disrupt. I’ll be impressed when ILM talks about it.
- commakozzi 3y agothis'll age well...
- CamperBob2 3y agoIt's "God of the Gaps" all the way down with these folks.