4 ms·
The reason why is because there is no training data of the sort you describe out there. By using MVS-based approaches, they are able to get over the data hurdl
by Fission 7y ago
The reason why is because there is no training data of the sort you describe out there.
By using MVS-based approaches, they are able to get over the data hurdle by compiling a dataset of your average YouTube video, instead of creating 3D renderings that include dynamic people. Importantly, MVS is really quite accurate, and in many cases can be considered ground truth.
Being able to forgo 3D renderings to use video only is almost certainly a reason why their results are so good.