4 ms·
fair, in that they aren't perfect photorealism, but if their comparisons with the regular codecs and techniques are correct, I'd take their modelled faces over
by ACow_Adonis 6y ago
fair, in that they aren't perfect photorealism, but if their comparisons with the regular codecs and techniques are correct, I'd take their modelled faces over the comparable- level of digital artifacts for the same bandwidth.
After all, it doesn't need to be said that when you're on a regular videoconferencing call and bandwidth starts to suffer, the resulting images don't really look anything like a photorealistic person either. I think this is actually a really good use of NN.
- mehrdadn 6y agoThe thing is you only want to make this trade-off when the bandwidth is actually starting to suffer. It'd be nice if there was a nice way to make this adaptive and use NNs only when throughput is low, but the nonlinearity of the distortions makes me think this would be really hard. [1] I know what I don't want is for a normal conversation in an uncongested network to look unnatural or for facial expressions to get distorted unnecessarily. Edit: [1] I meant to say doing a mixture of these (with the NN image as the "base", with H.264 to improve accuracy) seems really hard. On the other hand, just a hard switch from H.264 to NN when quality degrades is probably quite practical?
- algieg 6y agoAdaptive use would be awesome. Especially on conference calls at work there is always that one (or ten) person whose connection is absolutely awful and looks like a giant pixelhead.
- ACow_Adonis 6y agoPerhaps it's just me, and my philosophical bent, but I can actually see coming at it from the exact opposite end. I don't want bandwidth and things spent/wasted without it providing a significant benefit (I'm probably one of those 1080p/720p is good enough for most things type guys). I definitely don't want work making large bandwidth or resource claims on my connection when I'm working at home. And if any of this remote working has taught me anything, it's that most of my colleagues don't have steady/reliable tech or connections, so i'd almost want it used pre-emptively as a default so we can spend the rest of those resources on robustness or other qualities. (I realise of course that at the moment none of them have high-grade Nvidia graphics cards, but I'm talking hypothetically in the far off future). In short, I want a world where the cost to benefit ratio of things is orders of magnitude larger, because things like this let us spend network/resources on things which matter. Yes, when I'm calling my parent/grandparent one on one I might want to upgrade the signal, but I don't need to see random colleague's face in all their HD glory, or remote people whom I have no idea who they look or sound like anyway (i believe that's also been one of the findings with deepfakes, that you don't notice the eerieness/falseness as much if it's a reference of a face that you don't have pre-determined knowledge of).