11 ms·
Background: researched this space for a graduate degree. There are a few issues that are unanswered by this video (which isn't intended to be a technical deep
by quantumquetzal 4y ago
Background: researched this space for a graduate degree.
There are a few issues that are unanswered by this video (which isn't intended to be a technical deep dive, but I don't see any related links in the video description):
1. How do these glasses handle multiple simultaneous speakers? Based on the display I saw, it shows the speakers' words sequentially, which starts to fall apart in real-world environments, especially group conversations. This is a big problem, and wider adoption is contingent on handling this elegantly.
2. These appear to be the classic "smart glasses" display style that's pervasive in consumer head-worn displays today, where content is projected at a fixed depth in front of the wearer. Because the captions aren't anchored at the same focal distance as the speaker, the wearer's eyes will swap between the captions and the speaker's faces, which is a tiring activity, and can make the wearer feel like they're not part of the conversation or being rude.
3. As mentioned by another commenter, this is a useful idea for people who lose their hearing later in life. That said, this is less (although certainly still) useful for people who have congenital hearing loss and primarily communicate via ASL.
All in all, it's exciting to see growing interest in this space, as it's easily extendable to people learning a new language or navigating a foreign country. I think offloading the speech-to-text to a tethered mobile device is a good choice (though it would be nice to do low-latency wireless transmission).
- ape4 4y agoIt would be nice if it put the captions over the speaker (in a speech bubble?)
- danscarfe 4y agoYou pin the captions next to the person. We did try speech bubbles, but it didn't look great, so we went back to simple subtitles.
- quantumquetzal 4y agoThere’s a significant amount of evidence pointing to the fact that “anchored” captions (a la speech bubbles, like a comic book) is a preferred way to render captions, assuming the bubbles are at the same focal distance as the speaker. This is solvable with “true” AR, but getting that compute into the lightweight form factor of the original video is an unsolved problem, and is a ways off.
- danscarfe 4y agoWe have true AR with 3D placement now on these glasses
- aeturnum 4y agoThis is a classic curbcut in the sense that it will help those with heading as much (if not more than) the D/deaf community. Still very excited for it - agree with all your questions and concerns. As usual, this marketing seems most directed at normate ideas about what disabled people want / need, but the tech seems very cool and there does seem to be potential. Without looking into the product deeply it seems like there are D/deaf people on the team, which gives me hope. I do wish that we would just embrace the idea that using machines to make information available in many mediums is something all people can use and appreciate.
- nohaydeprobleme 4y agoTo make the comment easier to read for others, it looks like the commenter may have made a slight typo and meant to write "hearing" instead of heading, to convey: > "This is a classic curbcut in the sense that it will help those with [edit: hearing] as much (if not more than)..."
- aeturnum 4y agoyes, exactly, thank you
- DrKeithDuggar 4y agoThank you for looking at XRAI Glass! 1. For multiple simultaneous speakers of comparable volume, it’s only as good as the underlying speech-to-text engines we’ve implemented/integrated, which is currently not very good. It’s active area of research and engineering for us and we believe we’ll make strides to improve things; but, as you rightly point out, solving the crosstalk problem is very difficult. For the more general so-called "cocktail party problem", we can do a good job of filtering out more distant/lower volume voices and other environmental noise. Choosing the right microphone can improve things further, for example by pairing a noise canceling Bluetooth lapel mic. 2. We allow one to project the subtitles at varying depth, within the capabilities of glasses. We're seeing an effective focal depth range for fixed apparent size of about 0.5m to 3m. If one also allows change in apparent size, to simulate perspective scaling, the range is higher.
- roughly 4y agoI imagine it’d substantially increase the compute load, but I’d be curious if you could use multiple microphones and beam forming to separate out the streams of speech and feed them to the TTS algorithm independently.
- kajecounterhack 4y agoSingle-mic source separation is possible in an unsupervised manner today that could probably work better than beamforming both compute-wise and with regard to implementation difficulty (you'd just need a lot of recordings to represent the space of sounds you want to separate).
- morcheeba 4y agoMaybe a combination? Even simple beamforming/stereo would be helpful to help display the speaker's location. For example, the "speaker 1" tag could appear on the left, center, or right of the display to give a spatial clue where they are located.
- 4y ago
- gumby 4y agoI have a different use case, involving wearing them in the house: listen to what my girlfriend says, use some ML to analyze whether I need to know it and if so, put it on the display. She has a different "speech mode" than I: she speaks while I'm reading or washing the dishes or whatever, and sometimes it's to herself, sometimes to Siri, and sometimes it's something she wants me to know.
- cecilpl2 4y agoThere is a simple wetware solution, which is that she learns to say your name before saying something she wants you to know.
- drjasonharrison 4y agoDo you have any idea how hard it is to train partners to use your name once they fall out of practice? It's like there is a negative reinforcement stimulus applied everytime they do use your name. I noticed that my wife stopped using my name when it became more ambiguous as to who she was talking to because our kids are now older and conversations (really instructions and queries) are at the "adult content" level rather then "child content" level. Retraining her will be challenging.
- whitemary 4y agoThat would still require him to listen to his girlfriend talk. With all the technological innovation available today, nobody should have to listen to their girlfriends talk.
- justinator 4y ago>3. As mentioned by another commenter, this is a useful idea for people who lose their hearing later in life. That said, this is less (although certainly still) useful for people who have congenital hearing loss and primarily communicate via ASL.> Someone primarily communicates with ASL and then there's me that doesn't know ASL. I can speak to them, and they can read what I've spoken. That works pretty well. They communicate with me via text to speech, or (I guess in the near future) ASL to speech - however that will work. I mean, that's awesome.
- akira2501 4y agoASL is heavily inspired by English, so it's usually very easy for people to become basically conversant in it rather quickly. For signs you don't know, there's literal "finger spelling" that's part of the language, so conversational learning is greatly aided by this. Which.. aside from that, you could do this anyways. When I first started living with a deaf person, I just wrote things down on paper, and they wrote back. The glasses have some benefits over paper, but significant drawbacks as well, and many deaf people I've met have been trained to read lips. This isn't adding much.
- quantumquetzal 4y agoI would disagree! It takes time and resources to learn ASL or even fingerspelling. This would allow people without the time/resources to communicate effortlessly, which I would say is a marked improvement. That said, you could argue it would be a one-way conversation, which is a fair point.
- justinator 4y ago>This isn't adding much. We can communicate without seeing each other.
- kiba 4y agoASL isn't inspired by English.
- p1necone 4y agoIs multiple simultaneous speakers at the same volume/distance actually an important problem to solve? I already can't have a conversation if that's happening and my hearing is fine.
- quantumquetzal 4y agoLet’s say you’re in a meeting with colleagues who are having a side discussion about something, which you can contribute to. In this current setup, that conversation would not be displayed, while as a hearing person, you could context switch on a dime. It’s a decently-common occurrence. There’s a significant amount of literature pointing to the fact that people who are DHH avoid group conversations because they can’t follow along (see references). This sort of falls hand-in-hand with speech localization, which is another unsolved problem (and likely is unsolvable without true AR). https://journals.sagepub.com/doi/10.1177/1084713807301322 https://journals.sagepub.com/doi/10.1177/1084713807301322 https://doi.org/10.1080/02674649366780051 https://doi.org/10.1080/02674649366780051 https://doi.org/10.1111/j.1473-4192.1998.tb00121.x https://doi.org/10.1111/j.1473-4192.1998.tb00121.x
- midoridensha 4y agoI was about to make the same comment: my brain works the same way. A one-on-one conversation with low ambient noise, no problem. A bunch of loud conversations all around me? Forget it, I can't understand anything. I don't see why an aid for hearing-impaired people needs to handle this at all.
- sdrothrock 4y agoThis is an extremely common situation at many restaurants and is a reason I don't really enjoy going out to eat with friends without prior investigation.
- flangola7 4y agoThe Android voice recording app can already write a transcript in real time that indicates different speakers.
- danscarfe 4y agoBut not in AR
- nostrebored 4y agoIs low latency wireless transmission feasible now? The last time I worked on anything here, there were a number of problems with transmitting highly compressed, low resolution video data. The consumer devices could not handle just sending the packets. In my project we were annotating real time video and only sending back the annotations, but even that would cause devices to overheat and the applications to fail in really interesting ways.
- danscarfe 4y agoYes we demonstrated XRAI running on the Qualcomm Reference Wireless Viewer at MWC