9 ms·
Apple’s Persona technology uses Gaussian splatting to create 3D facial scans
- SomaticPirate 11mo agoThis video might help explain 3D Gaussian splatting. https://www.youtube.com/watch?v=wKgMxrWcW1s https://www.youtube.com/watch?v=wKgMxrWcW1s Essentially, an entirely new graphics pipeline with different fundamental techniques which allow for high performance and fidelity compared to... what we did before(?) Cool.
- reactordev 11mo agoNot quite, it’s just a way to assign a color value to a point in space (think point clouds) based on photogrammetry. It’s voxels on steroids but still is drawn using the same techniques. It’s the magic of creating the splats that’s interesting.
- ChadNauseam 11mo agoA color value for each point is a good starting place to gain an intuition. Some readers might be interested to know that the color is not constant for each point, but instead dependent on viewing angle. That is part of what allows splats to look realistic. Real objects have some degree of specularity which makes them take on slightly different shades as you move your head.
- adfm 11mo agoAnd since we normally see with binocular vision, a stereoscopic view adds another layer of realism you wouldn't normally perceive otherwise. Each eye sees subsurface scattering differently and integrates in your head.
- smartties 11mo agoThe same graphics pipeline is used: rasterization.
- ChadNauseam 11mo agoRasterization is a very general term. There is a big difference in practice between the traditional rasterization pipeline and splat rasterizers
- Groxx 11mo agoit's kinda like saying "we still show pixels". true but almost totally useless for understanding anything.
- colordrops 11mo agoSorry but this is a horrible video. The guy just spews superlatives in an annoying voice until 4:30 (of a 6 minute video mind you), when he finally gives a 10 second "explanation" of Gaussian splatting, which doesn't really explain anything, then jumps to a sponsored ad.
- Groxx 11mo agoyeah... their older videos are a bit more useful from what I remember (more time spent on the research paper content, etc), but they've become so content-free that I just block the channel outright nowadays. it's the "this changes everything (every time, every day)" hype-channel for graphics.
- deleted 11mo ago[deleted]
- dymk 11mo agoThat video didn’t explain what Gaussian splatting is at all, but I did get a minute ad read for some cloud GPU service.
- coryrc 11mo agohttps://packet39.com/blog/a-primer-on-gaussian-splats/ https://packet39.com/blog/a-primer-on-gaussian-splats/ is much better (don't load on mobile though, lots of data).
- tantalor 11mo ago"Now out of beta"?? Just in time for Vision Pro to go big. Right?
- deleted 11mo ago[deleted]
- deleted 11mo ago[deleted]
- dangus 11mo agoIt’s amazing tech, it’s just a solution looking for a problem. It feels a bit like the original Segway’s over-engineered solution versus cheap Chinese hoverboards, then the scooters and e-bikes that took over afterwards. Why would I be paying all this money for this realistic telepresence when my shitbox HP laptop from Walmart has a perfectly serviceable webcam?
- october8140 11mo agoI would not describe creating an experience that feels like you are in the room with a group of people, even allowing cross talk, is a solution looking for a problem. I think it's the thing everyone slowing dying on Zoom calls wishes they could have.
- bigyabai 11mo agoI disagree. Many of us don't use a headset regularly or carry it with us like a phone or laptop; it is an express inconvenience to use, with only marginal benefits. Businesses won't want one if webcams still do the trick, and users might respond positively but are always priced-out of owning one. If I'm doing work at my desk and I get a Zoom call, there is a 0.00% chance I will go plug in my Vision Pro to answer it. I'm just going to open the app and turn on my webcam, spatial audio be damned.
- LtWorf 11mo agoOh no, they wish to have fewer useless meetings.
- raincole 11mo agoWhy do we have video call meetings when people mostly just listen and the information is carried via audio? Why do we have 4K monitors when 1920x1080 is perfectly fine for 99.999% of use cases? If you look at the world through this lens called "serviceability" you'll think everything is a solution looking for a problem.
- criddell 11mo ago
- Stalker_Aloy 11mo agoCorridorDigital recently used the tech to assist in remaking the rooftop bullet-time scene from The Matrix. It's used for making the environment instead of modeling it from scratch. https://www.youtube.com/watch?v=iq5JaG53dho&t=2s https://www.youtube.com/watch?v=iq5JaG53dho&t=2s
- stevage 11mo agoTo be clear: they used Gaussian splatting, they didn't use Vision Pros.
- extraduder_ire 11mo agoThey also had an earlier video that more heavily featured gaussian splats. Using them to recreate the inside of the universal studios theme park without permission. I was very impressed with how it handles reflections on glass. https://www.youtube.com/watch?v=cetf0qTZ04Y https://www.youtube.com/watch?v=cetf0qTZ04Y
- october8140 11mo agoTested talked similar about Personas. https://youtu.be/LzZ2j9CAcww?si=IRvxNaNZeBQp7WLV https://youtu.be/LzZ2j9CAcww?si=IRvxNaNZeBQp7WLV
- samplatt 11mo agoI'm usually a fan of Norm's videos, but this might be the first time I've seen a Tested video that felt more like paid-promotion than an actual unbiased review. I don't keep up with it though.
- AceJohnny2 11mo agoI gotta say, these new Personas are good. The previous beta ones were terrifying frankenstein monsters. The new ones fooled my boss for 30 minutes. There's a bit of uncanny valley left, nevertheless. My persona's smile reminds of the horrible expressions people like to make in Source Filmmaker.
- frenzcan 11mo agoWhat eventually tipped your boss off? Was it the smile issue?
- bigyabai 11mo agoMust have realized the disembodied fuzzy head wasn't his hangover.
- quitit 11mo agoI have a few similar take aways: 1. The scanning is fast, it takes longer to set up a fingerprint on a macbook air. Just turning the head from side to side, then up and down, smiling and raising one's eyebrows. 2. I used the M5, and the processing time to generate the persona was quick. I didn't time it, but it felt like less than 10 seconds. 3. My cheeks tend to restrict smiling while wearing the headset, it works but people that know me understood what I meant when I said my smile was hindered. 4. Despite the limited actions used for set up, it reproduces a far greater range of facial movements. For example if I do the invisible string trick, it captures my lips correctly (when you move the top lip in one direction and the lower lip in the opposite direction, as if pulled by a string.) 5. I wasn't expecting this big of a jump in quality from the v1.
- pndy 11mo ago> There's a bit of uncanny valley left Perhaps how their heads, eyes move with this weird "fluid" effect and way too much blurred faces?
- 1123581321 11mo agoFor those who have had Persona conversations, how does varying audio latency affect immersion? Is there a recommended chat service?
- utopiah 11mo agoI don't use it very frequently but when from the few times I did I can't recall any imperceptible lag via Apple iMessage.
- 1123581321 11mo agoGood to know. I should try iMessage video chat more, in general.
- crazygringo 11mo agoWhat audio latency? There's regular latency due to distance, just like on a phone call if you're chatting with someone halfway across the world. But on a normal connection, audio and the persona should always be in sync, the same way audio and video are over Zoom or FaceTime. There shouldn't be any extra latency for the audio only.
- 1123581321 11mo agoVideo and audio aren't always quite in sync in Zoom, in my experience. But you're right, the overall latency of the connection should've been my question.
- crazygringo 11mo agoThey aren't always, but they are on a normal connection. It's only when packets are getting dropped or delayed that they temporarily get out of sync, as the audio and video streams compensate in different ways.
- akdor1154 11mo agoHow's the latency? Latency is what makes Zoom et al painful for me now - it ruins the ability to politely interject, give confirmatiom, etc. Does Apple do a better job of this than Google/Zoom? In theory you could get 20-30ms (just spitballing numbers I used to get playing shooters!) but i've never got anywhere near that with vid conferencing. Even so, latency-in-zoom kind of becomes an attribute of the medium and you learn to adapt. How does it feel with the Vision Pro though? The article talks about a really convincing sense of being in the same place with someone - how does latency affect that? (And does it differ based on if you're all physically in Silicon Valley or not?)
- setopt 11mo ago> latency-in-zoom kind of becomes an attribute of the medium and you learn to adapt. To some degree but not fully. When you adapt your brain is still doing extra work to compensate, similarly to how you don’t «hear» jet engine noise after acclimating to an airplane but it will still tire you to some degree. I had Zoom and Teams meetings daily during Covid, and personal FaceTime calls almost daily for a while. I still get «Zoom fatigue» if a call goes on for over an hour, if I need to talk face to face during the call (i.e. no screen sharing, can’t disable video and look at something else, etc.) I’m fine if I don’t look at people’s faces but rather people’s screen sharing.
- crazygringo 11mo agoI would assume any added latency is negligible -- the sensors + interpretation + rendering should be very fast. But you've still got all the network latency including Wi-Fi latency on both ends. And you always need a small audio buffer so discrete network packets can be assembled into continuous audio without gaps. So I wouldn't expect this latency to be any different from regular videoconferencing.
- bombela 11mo agoThe laws of physics means that the longer the path for your network packet, the higher the latency. One way latency on the Internet across fiber is about 4μs to 5μs per kilometer in my experience. For example, SF to Paris is ~40ms one way (it used to be 60ms 15y ago, latency and jitter have really improved). Double those values for the round trip allowing you to interject in a conversation. Add wifi, which has terrible latency with a lot of jitter (1ms to 400ms jitter is not uncommon). Wi-Fi 7 should reduce the jitter and latency in theory. We shall see improvements in the coming decade. Cellphone 5G did improve latency for me, so I don't doubt WiFi will eventually deliver. In other words you need to be within 3Mm (3000km) away to get a chance at a 30ms roundtrip. And that's assuming peer to peer without wifi nor slow devices. For a conference call, everybody connects to a central server acting as the relay. So now the latency budget is halved already.
- KaiserPro 11mo agoTLDR Gaussian splatting. What is missing from the article is that creating a model from a few pictures is not that hard (well it is to do well, but hear me out) The difficult part is animating it realistically with the sensors you have, in real time. Extracting signal from eye-gaze cameras with a sighlty wider field of view, that allows realistic not not uncanny valley animation is quite hard to do on the general public Peoples faces are all different sizes and shapes, to the point that even getting accraute gaze vectors is hard, let alone smile and check position (those are done with different cameras, not just eye gaze. )
- crazygringo 11mo agoThis is what fascinates me as well. I have to assume there's a neural net that effectively learns all of the possible muscles in the face. The limited sensor data gets fed in, and it's able to infer the full face shape. It seems perfectly plausible in theory, but I'm still impressed it seems to work so well in practice.
- nQQKTz7dm27oZ 11mo ago[dead]
- _kb 11mo agoThere's a bit more of a conversation / demo here which is pretty impressive: https://www.youtube.com/watch?v=KbZfbqHeJNU https://www.youtube.com/watch?v=KbZfbqHeJNU.
- Cthulhu_ 11mo agoOh man that was weird; I opened the video in a private browsing thing to not pollute my watch history and the version I got was automatically translated to Dutch, including voiceover which I presume is AI driven to try and match the tone of the original video. Still a bit robotic though. While I have my browser configured to prefer Dutch, the second one is English; I wish I could tell it / them that I don't want them to translate anything if it's in one of those languages.
- cubefox 11mo agoYeah that is awful behavior of YouTube. I can only imagine none of the YouTube developers or managers speak multiple languages.
- Mistletoe 11mo agoThe floating heads in a room having a meeting reminds me of terrible sci fi.
- stevage 11mo agoThere's a video version of the article linked partway down which actually works better than the text one for seeing the thing in action a bit.
- cubefox 11mo agoOn last SIGGRAPH there was actually a company which makes dynamic 3D Gaussian splatting videos now rather than static scenes: https://www.youtube.com/live/ucRukZM0d1s?t=1h1m50s https://www.youtube.com/live/ucRukZM0d1s?t=1h1m50s https://zju3dv.github.io/freetimegs/ https://zju3dv.github.io/freetimegs/ https://www.4dv.ai/ https://www.4dv.ai/ The videos can be played back in real-time, though they require multiple cameras to capture.
- KissSheep 11mo ago[dead]
- pksebben 11mo agoCame this close to buying an AVP, before learning that they only mirror a single screen with no virtual monitors. Like, guize, c'mon. Virtual desktop can do three. For 3.5k you gotta do better. I don't particularly need a virtual me in space as much as I need more screens that can do, like, actual work.
- bytesandbits 11mo agois this still the case even with the new M5? if so, wtf apple
- pksebben 11mo agoit is. There's a third-party app that can add 1 virtual monitor, for a total of 2, but FWIG it's not terribly stable (and 2 is still like 5 fewer than I want). wtf apple, indeed.
- efsavage 11mo agoI've always used 2-3 monitors pretty comfortably but with high latency AI agents adding more concurrency to my workflows I'm feeling very crowded. I would love a VR experience with an arbitrary number of screens/windows as well as more clearly separated environments (like having a visually different virtual office per project) that I can quickly switch between.
- pksebben 11mo agoMy assumption is that it's a network bottleneck, and apple clutches their pearls when anyone suggests lowering resolution or allowing for some latency. My take is like, make me tether with usb-c, reduce resolution and increase latency if I go over what the connection can handle. Use foveated rendering. All I want is more screens. For now, I'm working with Virtual Desktop on my Quest 3. It's not ideal - pixel density at the edge sucks and even in center it's not quite good enough for text unless I enlarge my screens to be the size of barn doors, but I get 3 very large screens out of my m1 and that makes me happy enough. It's also lighter than an AVP, which after test driving I assume multi-hour sessions would become a literal pain in the neck. Whatever the tradeoffs are, though, if apple offered infinite screens with text-readability I'd gladly throw money at them for the privilege. Tinfoil hat moment - I do wonder if the AVP devs got a visit from a bat-wielding gang of monitor engineers. Apple screens ain't cheap.
- kilibe 11mo agoNorris's dodge on iPhone scanning is telling—processing on-device keeps it secure and magical, but imagine Personas popping up in FaceTime cameos or ARKit bridges to iPads. How soon until we see cross-device ecosystems like Microsoft's Mesh, but with Apple's polish? Eager for that affordability leap; until then, thanks for the vivid demo, Scott—now I need a Vision Pro buddy just to test this out.