13 ms·
Perceptually lossless (talking head) video compression at 22kbit/s
- andrewstuart 2y agoThe more magic AI makes, the less magical the world becomes.
- andai 2y ago?
- andai 2y agoWhy am I downvoted for asking parent to clarify? Was I impolite for not using a full sentence?
- EarlKing 2y agoClearly Sauron is a jealous ringmaker and doesn't like hobbits using his ring to shitpost.
- Joel_Mckay 2y agoProbably just disappointed at the wasted bandwidth: 24fps * 52 facial 3D marker * 16bit packed delta planar projected offsets (x,y) = 19.968 kbps And this is done in Unreal games on a potato graphics card all the time: https://apps.apple.com/us/app/live-link-face/id1495370836 https://apps.apple.com/us/app/live-link-face/id1495370836 I am sure calling modern heuristics "AI" gets people excited, but it doesn't seem "Magical" when trivial implementations are functionally equivalent. =3
- scotty79 2y agoI think the point here is to make it photorealistic which everything apart from AI still fails at superhard.
- Joel_Mckay 2y agoTake a minute to look something up first, and then formulate a more interesting opinion for us to discuss: https://www.unrealengine.com/en-US/metahuman https://www.unrealengine.com/en-US/metahuman The artifacts in raster image data is nowhere near what a reasonable model can achieve even at low resolutions. =3
- scotty79 2y agoI know metahuman. As impressive as it is, when you judge by the standards of game graphics, if you are ever mislead into thinking metahumans are real humans or even real physically existing things it's time to see your eye doctor (and/or do MRI head scan). On the other hand AI videos can be easily mistaken for people or hyper realistic physical sculptures. https://img-9gag-fun.9cache.com/photo/aYQ776w_460svvp9.webm https://img-9gag-fun.9cache.com/photo/aYQ776w_460svvp9.webm There's something basic about how light works that traditional computer graphics still fails to grasp. Looking at its productions and comparing it to what AI generates is like looking at output of amateur and an artist. Sure, maybe artist doesn't always draw all 5 fingers but somehow captures the essence of the image in seemingly random arrangement of light and dark strokes, while amateur just tries to do their best but fails in some very significant ways.
- Joel_Mckay 2y ago"AI" videos make many errors all the time, but most people are not aware of what to look for... Undetectable CGI is done in film/games all the time, and indeed it takes talent to hide the fact it is fake. One could rely on the media encoder to garble output enough to look more plausible (people on potato devices are used to looking at garbage content.) However, at the end of the day the "uncanny valley" effect takes over every-time even for live action data in a auto-generated asset, as the missing data can't be "Magically" recovered with 100% certainty. Bye =3
- scotty79 2y agoUndetectable CGI in games ... right. I don't think you are a gamer. In movies it can be done with enough of manual tweaking by artists and a lot of photographic content around to borrow sense of reality from it. "Potato" devices by which I assume you mean average phones, currently have better resolutions than PCs had very recently and a lot still do (1080p). And a photo on 480p still looks more real than anything CGI (not AI). Your signature is hilarious. I won't comment about the reasons because I don't want this whole thread to get flagged.
- satvikpendem 2y ago> Any sufficiently advanced technology is indistinguishable from magic. - Arthur C. Clarke
- HPsquared 2y agoThis is the power of numerical methods.
- andrewstuart 2y agoThere’s a finite amount of magic and if AI borrows it here then it must be repaid there.
- psychoslave 2y agoThe greatest feat ever: let magic disappear before wonder of understanding.
- xyzsparetimexyz 2y agoOh shut up. There's plenty of awful uses for ai but this isn't one of them
- andai 2y agoWhat did you mean by this?
- AndrewVos 2y agoElon weirdly looks more human than usual in the AI version!
- LeoPanthera 2y agoThis is very impressive, but “perceptually lossless” isn’t a thing and doesn’t make sense. It means “lossy”.
- high_byte 2y agowhy not? if you change one pixel by one pixel brightness unit it is perceptually the same. for the record, I found liveportrait to be well within the uncanny valley. it looks great for ai generated avatars, but the difference is very perceptually noticeable on familiar faces. still it's great.
- codeflo 2y agoGP is correct, that’s the definition of “lossy”. We don’t need to invent ever new marketing buzzwords for well-established technical concepts.
- AndrewDucker 2y agoGP is incorrect. There is "Is identical", "looks identical" and "has lost sufficient detail to clearly not be the original." - being able to differentiate between these three states is useful.
- Rygian 2y agoLossless means "is identical". The other two are variations of lossy. Calling one of them "perceptually lossless" is cheating, to the disadvantage of algorithms that honestly advertise themselves as lossy while still achieving "looks identical" compression.
- protimewaster 2y agoIt's a well established term, though. It's been used in academic works for a long time (since at least 1970), and it's basically another term for the notion of "transparency" as it relates to data compression.
- red0point 2y ago> But one overlooked use case of the technology is (talking head) video compression. > On a spectrum of model architectures, it achieves higher compression efficiency at the cost of model complexity. Indeed, the full LivePortrait model has 130m parameters compared to DCVC’s 20 million. While that’s tiny compared to LLMs, it currently requires an Nvidia RTX 4090 to run it in real time (in addition to parameters, a large culprit is using expensive warping operations). That means deploying to edge runtimes such as Apple Neural Engine is still quite a ways ahead. It’s very cool that this is possible, but the compression use case is indeed .. a bit far fetched. A insanely large model requiring the most expensive consumer GPU to run on both ends and at the same time being limited in bandwidth so much (22kbps) is a _very_ limited scenario.
- jl6 2y ago130m parameters isn’t insanely large, even for smartphone memory. The high GPU usage is a barrier at the moment, but I wouldn’t put it past Apple to have 4090-level GPU performance in an iPhone before 2030.
- gambiting 2y agoOne cool use would be communication in space - where it's feasible that both sides would have access to high-end compute units but have a very limited bandwidth between each other.
- bliteben 2y agoWonder if its better than a single color channel hologram though
- JamesLeonis 2y agoIncreasingly mobile networks are like this. There are all kinds of bandwidth issues, especially when customers are subject to metered pricing for data.
- bityard 2y agoBandwidth is not the limitation in space comms, latency is.
- Vecr 2y agoFire Upon the Deep had more or less this. Story important, so I won't say more. That series in general had absolutely brutal bandwidth limitations.
- pastelsky 2y agoDid not expect to see Emraan Hashmi in this post!
- shaan7 2y agoIndeed! Bollywood makes it to HN xD
- JimDabell 2y agoI got some interesting replies when I suggested this technique here: https://news.ycombinator.com/item?id=22907718 https://news.ycombinator.com/item?id=22907718
- antiquark 2y agoNot quite lossless... look at the bicycle seat behind him. When he tilts his head, the seat moves with his hair.
- manmal 2y agoHis gaze also doesn’t quite match.
- hinkley 2y agoWhy is nobody noticing the eyes?? This is important! I feel like I’m taking crazy pills.
- olddustytrail 2y agoRead the text underneath the image and you'll understand.
- hinkley 2y agoNo, I really don’t. He acknowledges it’s not in keeping with the title or the thesis and then just sort of waves it off. Smells like rationalization to me.
- skandium 2y agoWell, this isn't probably a problem with the model, but the source frame having wrong eye gaze. Besides, perceptually lossless need not be defined in a side-by-side comparison context. If you were only viewing the right hand side video, how could you tell the eye gaze is off? The point was more on that the movement looks natural, unlike almost all neural avatars up to this year.
- manmal 2y agoYour argumentation does make sense to me; but it also makes the term lossless pull a lot of weight. Lossless in video encoding is usually defined by zero difference between source and target.
- gwd 2y agoThis reminds me of a scene in "A Fire Upon the Deep" (1992) where they're on a video call with someone on another spaceship; but something seems a bit "off". Then someone notices that the actual bitrate they're getting from the other vessel is tiny -- far lower than they should be getting given the conditions -- and so most of what they're seeing on their own screens isn't actual video feed, but their local computer's reconstruction.
- miohtama 2y agoAnd also it was a deep fake. BTW This is the best sci-fi book ever.
- Retric 2y agoMight be better if you like space opera style really soft science fiction. I really didn’t enjoy it.
- lern_too_spel 2y agoThe softness is deceptive. Hard concepts about communication and different types of brains are essential to the plot.
- gwd 2y agoA friend of mine and I both read it about the same time and discussed it afterwards. I thought it was pretty good, he thought it was not that great. What we agreed on was that in spite of there being many fantastic aspects to the book, on the whole it failed to be an awesome novel. Definitely worth giving it a try if you're a programmer, just for the fact that it's written by another programmer: the opening scene where they find a bunch of rules written down and just follow them reminds me of ACPI; the discussion of public-key cryptography and shipping drives full of one-time-pad around the galaxy; the "compression scheme" with the video.
- Boxxed 2y agoI agree that it was good but not particularly great. A Deepness in the Sky, however, is fantastic -- similar in many aspects but just flat out better all around.
- initramfs 2y agonice feature for low bandwidth 4G cell systems. Reminds me of the video chat in Metal Gear Solid 1 https://youtu.be/59ialBNj4lE?t=21 https://youtu.be/59ialBNj4lE?t=21
- hinkley 2y agoNice feature for many to one video conferencing as well. Though I don’t know if the organizers will agree.
- dormento 2y agoNow that you mention it, it never occurred to me that Snake's radio transmitted video as well. "Did you like my new sunglasses?" If you could reserve a small portion of the radio bandwidth to broadcast a thumbnail + low bandwidth compressed representation of the face movements, you could technically have something similar without encoding any video (think low res, eye + mouth movements).
- MayeulC 2y agoI like how the saddle in the background moves with the reconstructed head; it probably works better with uncluttered backgrounds. This is interesting tech, and the considerations in the introduction are particularly noteworthy. I never considered the possibility of animating 2D avatars with no 3D pipeline at all.
- vtodekl 2y ago[dead]
- up2isomorphism 2y ago“Perceptually lossless” is an oxymoron.
- Brian_K_White 2y agoThere is no oxymoron in "no perceived loss".
- ranger_danger 2y agoAs there are several patents, published studies, IEEE papers and thousands of google results for the term, I think it's safe to say that many people do not agree with your interpretation of the term.
- hinkley 2y agoYou’re still listening to vinyl, arntcha? Lossiness definitely matters when you’re doing forensics. But not for consumers. If you just want to bop to Taylor who the fuck cares. The iPod ended that argument. Yes I can be a perfectionist, or I can have one thousand songs in my pocket. That was more than half of your collection for many people at the time.
- up2isomorphism 2y agoCalm down dude. It is just a marketing term for something lossy.
- esafak 2y agoIt means you don't perceive the loss. What are you arguing; that you can perceive any loss?
- jacobgorm 2y agoRelated Show HN https://news.ycombinator.com/item?id=31516108 https://news.ycombinator.com/item?id=31516108
- hinkley 2y agoThe second example shown is not perceptually lossless, unless you’re so far on the spectrum you won’t make eye contact even with a picture of a person. The reconstructed head doesn’t look in the same direction as the original. However is does raise an interesting property in that if you are on the spectrum or have ADHD, you only need one headshot of yourself staring directly at the camera and then the capture software can stop you from looking at your taskbar or off into space.
- DCH3416 2y ago> unless you’re so far on the spectrum you won’t make eye contact even with a picture of a person. I don't know. I think you'd be surprised. That's already kind of an issue with vloggers. Often they're looking just left or right of the camera at a monitor or something.
- zbobet2012 2y agoThese sorts of models pop here quite a bit, and they ignore fundamental facts of video codecs (video specific lossy compression technologies). Traditional codecs have always focused on trade offs among encode complexity, decode complexity, and latency. Where complexity = compute. If every target device ran a 4090 at full power, we could go far below 22kbps with a traditional codec techniques for content like this. 22kbps isn't particularly impressive given these compute constraints. This is my field, and trust me we (MPEG committees, AOM) look at "AI" based models, including GANs constantly. They don't yet look promising compared to traditional methods. Oh and benchmarking against a video compression standard that's over twenty years old isn't doing a lot either for the plausibility of these methods.
- skandium 2y agoThis is my field as well, although I come from the neural network angle. Learned video codecs definitely do look promising: Microsoft's DCVC-FM (https://github.com/microsoft/DCVC https://github.com/microsoft/DCVC) beats H.267 in BD-rate. Another benefit of the learned approach is being able to run on soon commodity NPUs, without special hardware accommodation requirements. In the CLIC challenge, hybrid codecs (traditional + learned components) are so far the best, so that has been a letdown for pure end to end learned codecs, agree. But something like H.267 is currently not cheap to run either.
- zbobet2012 2y agoWinning in bd rate though isn't hard. You need to win in bd rate and have a hardware implementable, power efficient, cheap decoder. Agreed hybrid presents real opportunity.
- AzzyHN 2y agoDid you mean H.266? Or is there some secret H.267 that hasn't been agreed upon yet
- smokel 2y agoWhy so sour? This particular article doesn't seem to ignore a lot, it even references the Nvidia work that inspired it, as well as a recent benchmark. Someone was just having fun here, it's not as if they present it as a general codec.
- tommiegannert 2y agoNow that we're moving towards context-specific compression algorithms, can we please use WASM as the file header for these media files, instead of inventing something new. :)
- userbinator 2y agothe only information that needs to be transmitted is the change in expression, pose and facial keypoints Does anyone else remember the weirder (for lack of a better term) features of MPEG-4 part 2, like face and body animation? It did something like that, but as far as I know nearly no one used that feature for anything. https://en.wikipedia.org/wiki/Face_Animation_Parameter https://en.wikipedia.org/wiki/Face_Animation_Parameter and in the worst, trust on the internet will be heavily undermined ...as long as the model doesn't include data to put a shoe on one's head.
- stuaxo 2y agoBit off putting that it's Musk for some reason, maybe it's just overexposure to his bullshit, I could quite happily never see him again. Maybe there is a custom web filter in there somewhere that could block particular people and images of them.
- Separo 2y agoWe could run something like TensorFlow.js in a Chrome extension to identify the person in the image and replace it in the dom. A little resource intensive for inference on every image in but probably worth it in this case.
- accra4rx 2y agowhy does he has to deep fake Imran Hashmi the serial kisser