5 ms·
What is the real use case for this?
by isuckatcoding 4y ago
What is the real use case for this?
- tristor 4y agoThis would be useful for anyone that has specific company requirements for their avatar/image as long as it can be combined with something that detects the face and eliminates the background
- swah 4y agoMaking avatars without the 3D camera thing that phones do. Could also make avatars of famous people, deceased loved ones...
- scoopertrooper 4y agoI think this could be used as a great basis for next generation video conferencing. Rather than compressing raster images of people speaking, we could send a model as an initial payload and subsequent instructions on how to manipulate it as the conversation progresses. Most current networks can handle video conferencing (relatively) well. However, if 8k video or something like Light Field displays ever catch on, then we might end up needing orders of magnitude more bandwidth to drive video conferencing, so this might be an enabling technology.
- ricardobeat 4y agoMight not be that simple. In a recent demo from Unreal Engine they said that the facial animation data is the main bottleneck to shipping titles with more realistic avatars. Their minute-long demo took many TBs to store the motion data alone.
- cheschire 4y agoMy layperson understanding makes this baffling to me. Isn't motion inherently analog, and haven't we've developed high performing analog compression already? Motion seems significantly more amenable to lossy compression as well, and I could imagine a use-case for ML-based decompression to fill in between the keyframes.
- dmwallin 4y agoYou also have to consider that faces are one of the things that we have the strongest perception of, with lots of our neurons dedicated to the task, so when you get things wrong it's far more noticeable than many other bodily animations would be.
- jon_richards 4y agoOne of my favorite “weird” conjectures is that the existence of the uncanny valley implies that at some point it was evolutionarily advantageous for us to recognize something that looked human, but wasn’t… and to be afraid of it.
- rjeli 4y agoneanderthals?
- deleted 4y ago[deleted]
- giantrobot 4y agoThink also about face-like things we perceive because of pareidolia but might in fact be dangerous. One part of our brain is opportunistic about finding patterns while another is a sanity check on those patterns telling us not to just trust them outright.
- shwoopdiwoop 4y agoIs there any further reading on this topic in particular that is not leading to wild conspiracy theories?
- ricardobeat 4y agoI had the same reaction (note: I actually got things mixed up, the one in question is the Unity "Enemies" demo [1], not Unreal), but if you consider the amount of detail - the saliva 'sticking' to the corners of the mouth, the 40+ muscles of the face, lips, eyelids, the iris, etc - it makes sense that it would be a ton of data. IIRC they are working on some kind of ML-powered compression so that this becomes actually possible to ship (the demo is not available to download, presumably it's too big and would barely run on a 3090TI). [1] https://www.youtube.com/watch?v=eXYUNrgqWUU https://www.youtube.com/watch?v=eXYUNrgqWUU
- HWR_14 4y agoFacial animation is a real issue. This is facial mapping, which is an easier to solve problem, in that throwing artists at the problem is a solution.
- ekianjo 4y ago> Rather than compressing raster images of people speaking, we could send a model as an initial payload and subsequent instructions on how to manipulate it as the conversation progresses. What about latency though? Would that improve?
- fortran77 4y agoOr, at least, the faces that aren't actively speaking could be replaced by avatars. It will also help keep a clean image in case stuff is going on the the background, etc.
- abeppu 4y agoI guess, would watching an animated head avatar in 8k or a Light Field display be a better experience than watching lower resolution video? I wonder whether the experience would be too undermined by artifacts from e.g. your coworker's cat, for which the system doesn't have a model, wanders into frame partway through a call, and suddenly the system must dynamically either (a) 'downgrade' to just video (b) be able to create a new model for an arbitrary new thing on the fly (c) create some hybrid of partial video data with part viewer-rendered animation or (d) your coworker looks crazy interacting with something the system hides from you entirely.
- 8n4vidtmkvmk 4y agothere's another thread about just this, but very different approach. reduces bandwidth a lot specifically for video conferencing
- deleted 4y ago[deleted]
- jimmySixDOF 4y agoIt's a stepping stone towards the integration between NeRF derived and DALL-E 2 CLIP diffusion type systems to create user defined on demand real-time volumetric 3D digital spaces for spatial computing.
- sangnoir 4y agoIs "user defined on demand real-time volumetric 3D digital spaces for spatial computing" in this instance a euphemism for fanfic porn?
- jokethrowaway 4y agoI would make a model of myself create a nice background, then detect face expressions, apply them to the model and stream the model while I work shirtless from a dirty basement.
- zhyder 4y agoI think it could be used to create a videoconferencing experience akin to sitting around a table with multiple people, while still selectively facing or making eye contact with individuals.
- totalview 4y agoEasier rigged characters in video games that can be personalized to a player
- jayd16 4y agoUser avatars for games or video conferences. Post fx tooling for movies. Cheaply touch up or add virtual extras to a scene.
- acd 4y agoVideo conferences. You can create a 3d world projecting 3d avatars made from images from peoples web cams. Creating a 3d sense of being in the same room. Acting in 3d recreation of your favorite movies. If you can scan faces you could play in rendered 3d recreations of movies. Video games you could from a web cam image play a 3d avatar of yourself in a game.