5 ms·
Thank you for looking at XRAI Glass! 1. For multiple simultaneous speakers of comparable volume, it’s only as good as the underlying speech-to-text engines we’
by DrKeithDuggar 4y ago
Thank you for looking at XRAI Glass!
1. For multiple simultaneous speakers of comparable volume, it’s only as good as the underlying speech-to-text engines we’ve implemented/integrated, which is currently not very good. It’s active area of research and engineering for us and we believe we’ll make strides to improve things; but, as you rightly point out, solving the crosstalk problem is very difficult. For the more general so-called "cocktail party problem", we can do a good job of filtering out more distant/lower volume voices and other environmental noise. Choosing the right microphone can improve things further, for example by pairing a noise canceling Bluetooth lapel mic.
2. We allow one to project the subtitles at varying depth, within the capabilities of glasses. We're seeing an effective focal depth range for fixed apparent size of about 0.5m to 3m. If one also allows change in apparent size, to simulate perspective scaling, the range is higher.
- roughly 4y agoI imagine it’d substantially increase the compute load, but I’d be curious if you could use multiple microphones and beam forming to separate out the streams of speech and feed them to the TTS algorithm independently.
- kajecounterhack 4y agoSingle-mic source separation is possible in an unsupervised manner today that could probably work better than beamforming both compute-wise and with regard to implementation difficulty (you'd just need a lot of recordings to represent the space of sounds you want to separate).
- morcheeba 4y agoMaybe a combination? Even simple beamforming/stereo would be helpful to help display the speaker's location. For example, the "speaker 1" tag could appear on the left, center, or right of the display to give a spatial clue where they are located.
- kajecounterhack 4y agoYou could only display the speaker's location if you had some way to associate streams with the individuals. So you'd have to train an audio-visual association model. The thing you could do is train a localizer on the separated audio. Phase is estimated by the source separation process, so you can actually train an ML model provided you have some ground truth (e.g. estimated human locations from camera detections)
- rapjr9 4y agoDo you have any references for this or a link to a commercial service? I'm currently in the process or trying to extract some background voices in a video (an interview where the faint background conversation is in English and the loud overdub is in Bulgarian). I tried Melodyne but it seems to only separate on pitch, not volume, and the pitchs are too similar (mono, three voices, all female) and words are made of lots of short phonemes which each create a "note" and makes editing impossible. I looked into Izotope RX as well and it does not seem capable of doing this either. There are services that can automatically add subtitles using speech-to-text translators but they are expensive, and I'd prefer to have the background voices rather than the Bulgarian interpretation of them translated back into English. It seems possible to do, in Peter Jackson's The Beatles:Get Back they were able to separate voices from loud foreground musical sounds and other ambient talking, but that technology was custom and doesn't seem to be publicly available yet.
- kajecounterhack 4y agohttps://arxiv.org/abs/2110.10739 https://arxiv.org/abs/2110.10739 I haven't seen it provided as a commercial service or free model yet, but there is open source code for Mixit that lets you train using the open source / canned FUSS dataset.
- rapjr9 4y agoThanks! That led me to this which looks like a good place to start: https://github.com/google-research/sound-separation/blob/master/models/neurips2020_mixit/README.md https://github.com/google-research/sound-separation/blob/mas...
- quantumquetzal 4y agoWow! Thanks for the response. This is exciting work, and I’m pleased to see it being iterated on. Re: #2. I’m assuming the varying depth is manually-controlled? Or is it automated by some method? If it’s manual, can the adjustment be made while transcription is active? In other words, can I change the focal distance to match the speaker without interrupting the speaker? All in all, cool stuff! Best of luck with the work.
- danscarfe 4y agoThe depth is set manually but can be changed at any point, even mid conversation
- DrKeithDuggar 4y agoThank you for the praise and encouragement! Indeed, as Dan said, one can change the depth on the fly. In fact, one goal of the development team is to make as much functionality as possible changeable on the fly. For example, you can currently change subtitle depth, pinning, and size on the fly, spoken language, subtitle language, microphones, and audio settings on the fly, etc. I’d love for every setting and feature to support on the fly changes. That said, some things are currently fixed for a session, such as recording audio, and some third party software we utilize is less dynamic and forgiving of changes on the fly. For better or worse, in our software world of today, the sage advice of The IT Crowd “Have you tried turning on and off again?” still seems to hold with pragmatic force. And it still holds with XRAI ... sometimes ;-)
- fouc 4y ago> For multiple simultaneous speakers of comparable volume If you had more microphones placed in multiple spots on the glasses, such as up to 5 microphones - 1 in the center, 2 on the frames/end pieces, and perhaps a final 2 on the arms/temples. Then that would be able to catch conversation coming at a person from behind them, the sides, or directly in front, etc.
- sizzle 4y agoDo you think a ChatGPT real time processing engine will soon be able to parse the cross talk for words that are detected and given a probability score that then can then be reordered/reassembled back into a pretty accurate sounding conversation displayed back to the glasses wearer? This would be such a cool use case for the latest ChatGPT tech when it gets faster in the near future.
- eternityforest 4y agoSeems like the multiple microphone beamforming source separation algorithms are getting pretty good these days, maybe just adding a lot more mics would help? Could you have an AI model that extracts some characteristics of the speaker's voice for each individual word, then translates that to color and font? If the model was not confident about a word it could show slightly blurred, if it was loud it could be bold, perhaps(Although there's some stereotype issues) you could use different fonts for different pitches, whispers could be grey, quiet could be transparent. Maybe there's a language model that can pick up overlapping words if you don't have the constraint of needing to sort them out into who said it, just show all the possibile words that could have been said by anyone stacked together, in a "not sure" color, and maybe the wearer would eventually learn to figure it out without much effort? You could also try to stay consistent so the same speaker gets the same colors I'd possible, and also not reuse colors for new speakers that have been recently used, to best make use of the limited bits of data in font and color. Maybe just by showing all the words from every speaker all together like that, the wearer would be able to figure it out even if it made mistakes in the speaker identification?
- DrKeithDuggar 4y agoRegarding additions mics, yes that enabled more advanced spatial processing, especially when arranged in a determined and calibrated geometry. One could imagine glasses with several mics placed at optimal locations on the frames. Such multichannel audio could then be processed into multichannel spatial audio streams. As for the rest, thank you for wonderful ideas! Everything you propose is technically possible. The difficulties arise first in assessing the increased benefit to users versus the increased complexity of user experience, and second in prioritizing the work versus other features. Over time we do hope to add additional selectable “skins”, which is to say different UI designs, that allow users to choose UI’s from a wide range. Everything from the simple to complex layouts, from accessible to exotic color palettes, from professional to playful themes, etc. I could definitely see more advanced visual representations of transcription uncertainty showing up in such optional skins.