7 ms·
It Took Me 30 Years to Solve This VFX Problem – Green Screen Problem [video]
- superjan 7mo agoWatched this a few days ago. The video is light on technical details, except maybe that they used CGI to generate training data.
- rhdunn 7mo agoThe idea behind a greenscreen is that you can make that green colour transparent in the frames of footage allowing you to blend that with some other background or other layered footage. This has issues like not always having a uniform colour, difficulty with things like hair, and lighting affecting some edges. These have to be manually cleaned up frame-by-frame, which takes a lot of time that is mostly busy work. An alternative approach (such as that used by the sodium lighting on Mary Poppins) is that you create two images per frame -- the core image and a mask. The mask is a black and white image where the white pixels are the pixels to keep and the black pixels the ones to discard. Shades of gray indicate blended pixels. For the mask approach you are filming a perfect alpha channel to apply to the footage that doesn't have the issues of greenscreen. The problem is that this requires specialist, licensed equipment and perfect filming conditions. The new approach is to take advantage of image/video models to train a model that can produce the alpha channel mask for a given frame (and thus an entire recording) when just given greenscreen footage. The use of CGI in the training data allows the input image and mask to be perfect without having to spend hundreds of hours creating that data. It's also easier to modify and create variations to test different cases such as reflective or soft edges. Thus, you have the greenscreen input footage, the expected processed output and alpha channel mask. You can then apply traditional neural net training techniques on the data using the expected image/alpha channel as the target. For example, you can compute the difference on each of the alpha channel output neurons from the expected result, then apply backpropagation to compute the differences through the neural network, and then nudge the neuron weights in the computed gradient direction. Repeat that process across a distribution of the test images over multiple passes until the network no longer changes significantly between passes.
- Springtime 7mo agoIn an earlier video they made a couple years back about Disney's sodium vapor technique Paul Debevec suggested he was considering creating a dataset using a similar premise: filming enough perfectly masked references to be able to train models to achieve better keying. So it was interesting seeing Corridor tackle this by instead using synthetic data.
- somat 7mo agoWith regards to the sodium vapor process, an idea has been percolating in the back of my head ever since I saw that video. But I don't really have the budget to try it out. theory: make the mask out of non-visable light illuminate the backing screen in near Infra-Red light. (after a bit of thought I chose near-IR as opposed to near-UV for hopefully obvious reasons) point two cameras at a splitting prism with a near IR pass filter(I have confirmed that such thing exists and is commercially available) Leave the 90 degree(unaltered path) camera untouched, this is the visible camera. Remove the IR filter from the 180 degree(filter path) camera, this is the mask camera. Now you get a perfect non-color shifting mask(in theory), The splitting prism would hurt light intake. It might be worth it to try putting the cameras really close together , pointed same direction, no prism, and see if that is close enough.
- diacritical 7mo agoDon't humans and other warm objects also radiate IR?
- somat 7mo agoThat is far-IR, thermal stuff, Near-IR, 700 nanometer-ish is right below red in human vision. Camera sensors can pick up a little near-IR so they have have a filter to block it. If that filter was removed and a filter to block visable light was used in place you would have a camera that can only see non-visable light. Poorly, the camera was not engineered to operate in this bandwidth, but it might be good enough for a mask. A mask that does not interfere with any visible colors.
- vsviridov 7mo agoThe community has managed to drastically lower hardware requirements, but so far I think only Nvidia cards are supported, so as an AMD owner I'm still missing out :(
- mouth 7mo agoThis works on macOS, as well, via Apple Silicon.
- Computer0 7mo agoLooking forward to trying it out, 8gb of vram or unified memory required!
- IshKebab 7mo agoPretty impressive results! Seems like someone has even made a GUI for it: https://github.com/edenaion/EZ-CorridorKey https://github.com/edenaion/EZ-CorridorKey Still Python unfortunately.
- BoredPositron 7mo agoLike 90% of the other tooling in VFX...
- IshKebab 7mo agoIs it? That's a shame. I assumed this is Python because of Pytorch.
- diacritical 7mo agoFrom ~04:10 till 05:00 they talk about sodium-vapor lights and how Disney has the exclusive rights to use it. From what I read the knowledge on how to make them is a trade secret, so it's not patented. Seems weird that it would be hard to recreate something from the 1950's. I also wonder how many hours were wasted by people who had to use inferior technology because Disney kept it secret. Cutting out animals and objects from the background 1 frame at a time seems so mindnumbingly boring.
- jasonwatkinspdx 7mo agoYeah, that's just nonsense. We used sodium vapor monochromatic bulbs in my high school physics class to duplicate the double slit experiment. I suspect the real reason is that digital green screen in the hands of experienced people is "good enough" vs the complication of needing a double camera and beam splitting prism rig and such.
- meatmanek 7mo agoThe lights are relatively easy to get. iirc (it's been a bit since I watched their full video on the subject[1]) the hard part to find was the splitter that sends the sodium-vapor light to one camera and everything else to another camera. 1. https://www.youtube.com/watch?v=UQuIVsNzqDk https://www.youtube.com/watch?v=UQuIVsNzqDk
- diacritical 7mo agoYup, I wanted to say that the prisms are hard to recreate, not the light itself.
- aidenn0 7mo agoIt would seem to me to be relatively easy to build something like that if you're okay shooting with effectively a full stop less light (just split the image with a half-silvered reflector and use a dichroic filter to pass the sodium-vapor light one one side. The splitter would have to be behind the lens, so it would require a custom camera setup (probably a longer lens-to-sensor distance than most lenses are designed for too), but I can't think of any other issues.
- comex 7mo agoSee also this video comparing Corridor Key to traditional keyers: https://www.youtube.com/watch?v=abNygtFqYR8 https://www.youtube.com/watch?v=abNygtFqYR8
- Sniffnoy 7mo agoSummary: He created 4 hard-to-key shots, and on each of them tried KeyLight, IBK, and Corridor Key. Overall on 3 of them he judged that Corridor Key had done the best job, on one of them he judged that IBK had done the best job. I think on all of them he judged that more work was still necessary, none of them was fully usable as-is.
- ralusek 7mo agoI'm a software engineer that, like the vast majority of you, uses AI/agents in my workflow every day. That being said, I have to admit that it feels a little weird to hear someone who does not write code say that they built something, without even mentioning that they had an agent build it (unless I missed that).
- jrm4 7mo agoI mean, the heading of the video says "he solved the problem," which I think is wise to pay a lot of attention to.
- tekacs 7mo agoWorth bearing in mind that people in VFX are often relatively technical. From their own 'LLM handover' doc: https://github.com/nikopueringer/CorridorKey/blob/main/docs/LLM_HANDOVER.md https://github.com/nikopueringer/CorridorKey/blob/main/docs/... > Be Proactive: The user is highly technical (a VFX professional/coder). Skip basic tutorials and dive straight into advanced implementation, but be sure to document math thoroughly.
- steve_adams_86 7mo agoBack when I played with animation and post pipelines, I was writing a decent amount of python. It's part of how I got into programming. At the time I would have said I can't program, and I suspect this guy is similar.
- adamtaylor_13 7mo agoThis is interesting. I had the exact opposite reaction. You don't hear architects get hounded because they say they "built" some building even though it was definitely the guys swinging hammers that built it. But yet, somehow because he didn't artisanally hand-craft the code, he needs to caveat that he didn't actually build it?
- kalaksi 7mo agoMaybe it's a language thing. Architects saying they built something sounds a bit off to me. In my native language, and in everyday language, I don't think people would use "built" like that. I don't know how architects talk with each other, though.
- deleted 7mo ago[deleted]
- amelius 7mo agoThere's still a bug: the glass with water does not distort the checker pattern in the background at 24:12.
- CharlesW 7mo agoWhen you watch the video it becomes pretty clear why it wouldn't be able to do that, although it's fun to think about how a future iteration or alternative might be able to credibly (if you don't look too hard) mimic that someday.
- jweir 7mo agoTrue, but with visual art there is what is correct and what looks correct. When things are moving and the area small no one is going to notice. But now that is problem is solved a director will come along and say... I want a scene with a big glass of water and the camera will zoom in on it and will see the monster refracted through the glass.
- gmueckl 7mo agoAt that point it's better to do the glass entirely in post.
- orbital-decay 7mo agoSure, because they used monotone backgrounds and never really captured any distortion.
- DrewADesign 7mo agoYou’d have to track it, render it, and comp it in. It’s not ridiculously difficult, but there’s no way that’s going to happen automatically.
- orbital-decay 7mo ago>there’s no way that’s going to happen automatically They train their model in a pretty straightforward way, it can also be used to capture the distortion as well, just use a non-monochrome (possibly moving) background optimized for this. It's a matter of effort and attention to detail during training (uneven green screen lighting, reflections, etc), not fundamental impossibility
- dylan604 7mo agoThe sad thing about this is the problems encountered during post from the production team saying "fix it post" during the shoot. I've been on set for green screen shoots where the lighting was not done properly. I watched the gaffer walk across the set taking readings from his meter before saying the lighting was good. I flip on the waveform and told him it was not even (which never goes down well when camera dept tells the gaffer it's not right). He put up an argument, went back and took measurements again before repeating it was good. I flipped the screen around and showed him where it was obviously not even. A third set of meter readings and he starts adjust lights. Once the footage was in post, the fx team commented about how easy the keys were because of the even lighting. The problem is that the vast majority of people on set have no clue what is going on in post. To the point, when the budget is big enough, a post supervisor is present on production days to give input so "fixing it in post" is minimized. When there is no budget, you'll see situations just like in the first 30 seconds of TFA's video. A single lamp lighting the background so you can easily see the light falling off and the shadows from wrinkles where the screen was just pulled out of the bag 10 minutes before shooting. People just don't realize how much light a green screen takes. They also fail to have enough space so they can pull the talent far enough off the wall to avoid the green reflecting back onto the talent's skin. TL;DR They solved something to make post less expensive because they cut corners during production.
- weinzierl 7mo agoI fully agree but I think for them making it possible to cut corners during production is the whole point. Think about it: The choice is between 5 minutes of work plus a one time purchase of a decent GPU and a big room with a complex lighting setup with a post supervisor present. Now, quality of the end result will not be the same, for sure. You and me would opt for the quality setup whenever we can, but many others won't.
- dylan604 7mo agoIf you're on such a low production budget that you just physically do not have the lamps to light a screen, then you really have to ask if green screen is the right option. Maybe flip it and shoot black limbo so you do not need lights, and the lights you do have can be better used as key lights for separation. You also don't have to worry about the color cast from your light screen. Essentially, you just need a garbage matte for the key, and then clean up what might be getting keyed that you don't actually need. Detecting foreground subject from background is so capable now that a screen isn't necessary, and matte clean up is pretty much unnecessary. Of course you lose street cred of not being able to say you used green screen, but who cares as long as the shot works out.
- amelius 7mo agoIs it a coincidence that the result is stable between subsequent frames?
- jayd16 7mo agoAs far as alternatives, I wonder if anyone has tried a screen that cycles through colors in a known sequence. Using this modulating-color screen, it might actually be easier to separate the subject because you get around the "green shirt over green screen" problem. You might even be able to use a time sampling to correct the light cast on the subject from the screen as you would have a full spectrum of response. I could also imagine using polarized light as the backdrop as well.
- dgently7 7mo agothe general problem with any technique that isnt just throw some vaugely green thing behind our actors is that setting up complicated tech like this on an actual film set is extremely expensive. both the time it would take and the risk of it not working. so you end up with a dedicated permanent stage install but now you need to get the actors and crew to that place. better keys isnt a bad enough problem to justify that effort/cost. even the highly touted "virtual production" mandalorian stuff where you just put a big led wall behind the actors has shown to be more expensive than traditional vfx unless you tightly control the creative or approach.
- mcurist 7mo agoAnd all the people aware of the production technique watch it and imagine the characters saying "We can't run from the monster in different direction, our virtual production stage is precisely this big!" Green screens just more flexible
- MrVitaliy 7mo agoAnyone tried using lidar and just cut/measure distance to the object?
- wizzledonker 7mo agoThat would require calibration with the camera, and even then the camera and lidar sensor can’t be in exactly the same place. I doubt results would be better.
- summarity 7mo agoWell sort of, the industry tried to go way beyond that by capturing the entire light field: https://techcrunch.com/2016/04/11/lytro-cinema-is-giving-filmmakers-400-gigabytes-per-second-of-creative-freedom/ https://techcrunch.com/2016/04/11/lytro-cinema-is-giving-fil...
- rcxdude 7mo agoApparently they used something similar for production on avatar: stereo cameras for depth estimation which allowed realtime depth composition of CG characters onto the shots they were taking, which makes it a lot easier to get everyone on the same page about the scene, especially with characters that are outside normal human proportions. But it wasn't good enough for the final shots.
- dgently7 7mo agoper pixel depth does not solve for semi-transparency.
- Coeur 7mo agoWell there was ZCam, which was a time-of-flight add-on for ENG cameras. It couldn't really handle edges perfectly and existing bluescreen tech was good enough for TV production, so they pivoted into gaming and sold to Microsoft for the Kinect. https://en.wikipedia.org/wiki/ZCam https://en.wikipedia.org/wiki/ZCam (Demo: https://www.youtube.com/watch?v=s7Kcmx29RCE https://www.youtube.com/watch?v=s7Kcmx29RCE )
- swframe2 7mo agoThe model in this repo seems pretty good: https://github.com/xuebinqin/DIS https://github.com/xuebinqin/DIS
- tempaccountabgd 7mo ago[dead]
- qingcharles 7mo agoI use the Adobe version of this in Photoshop every day and I assumed that Adobe solved this the same way, but used professionals to cut out the subjects from the backgrounds then fed both versions into their AI. Since they added it a year or so ago it has been game-changing. I'm cutting out portraits every day and having a magical tool that cuts out the subject with perfect hair cut out with a single click is sci-fi. Here's a demo of Photoshop's tool: https://www.youtube.com/watch?v=SNVJN6PKeGQ https://www.youtube.com/watch?v=SNVJN6PKeGQ (the other magical Photoshop tool is the one that removes reflections from windows, which is even more insane when you reverse it and tell it you only want the reflection and not what's on the other side of the glass)
- lynnharry 7mo agoIt’s fascinating to see the bridge between academic research and industry application here. While Image Matting is a massive research area in Computer Vision, academia often focuses on solving perfect 'benchmarks.' Corridor Crew effectively took that foundational research, like neural unmixing and synthetic training, and adapted it to solve the 'messy' reality of production, like tracking markers and motion blur. It’s a great example of using open-source deep learning resources to build a tool that prioritized workflow over just a high accuracy score.
- tempaccountabcd 7mo ago[dead]
- mk_stjames 7mo agoIn case anyone wanted technical details of the NN, I dug into the repo: Its a transformer, with a CNN refiner after. Specifically, a ViT using the Hiera architecture (https://github.com/facebookresearch/hiera https://github.com/facebookresearch/hiera) The Hiera ViT has dual decoder heads, one for the alpha and one for the RGD foreground, and then a small CNN refiner network to solve some artifacting in the output from the Hiera model. I'd be very interested to see a long form tech talk of Niko explaining his process of learning ML ropes and building this model.
- jcmoscon 7mo agoIt is refreshing to see problems being solved by AI that are not LLMs. There are so many day-to-day challenges that we could solve using data, machine learn and some creativity.
- anfogoat 7mo agoThis is technically true I guess but assuming the YT comment I just read represented it truthfully, it was an LLM that wrote it.
- voxic11 7mo agoIf you look at the github page it was an LLM that wrote all the code. Makes sense as corridor are not software developers.
- grishka 7mo agoIt's well established that machine learning excels at solving classification problems, or those that can be reduced to one. It saddens me that we're wasting so much of that potential on those stupid stochastic parrots that solve all those non-problems that no one has ever had. It saddens me even more that so many people are absolutely sure that LLMs are "smart", or that they can "think", or even that they're somewhat conscious. And that even if they're not quite that, one more order of magnitude of scale will definitely give us an AGI. Oh that didn't help? Then one more, that will definitely be it. One real problem that LLMs have solved is that they made natural language processing as a discipline obsolete. They also usually don't suck at summarizing long texts, except when they sometimes do. But that's it, really.