6 ms·
Anyone interested in this will likely enjoy Don't Skype & Type![1] where researchers decoded keystrokes from background audio of Skype conversations. The best p
by pentestercrab 8y ago
Anyone interested in this will likely enjoy Don't Skype & Type![1] where researchers decoded keystrokes from background audio of Skype conversations. The best part is source code is available[2]. I wonder how many people are applying this to Twitch streamers or YouTube videos (especially any Talks at Google videos) today?
1. http://spritz.math.unipd.it/projects/dst/ http://spritz.math.unipd.it/projects/dst/
2. https://github.com/SPRITZ-Research-Group/Skype-Type https://github.com/SPRITZ-Research-Group/Skype-Type
- ggerganov 8y agoI've been working on a similar tool: https://github.com/ggerganov/kbd-audio https://github.com/ggerganov/kbd-audio (see the 'keytap' tool). Not sure how well it works yet, as I have made tests only with my setup. There is a live page that anyone can experiment with.
- feross 8y agoThis is incredible work. Very well done! Have you posted this to Show HN yet?
- ggerganov 8y agoThanks! I haven't posted it yet, as I want to test how reliable is the approach. I know it does not work with non-mechanical keyboards at all (most likely because the key sounds are quiet) and I have tested it only with my mechanical keyboard.
- feross 8y agoThen I must apologize as I've already posted it to Twitter where it's getting a lot of really positive reactions: https://mobile.twitter.com/feross/status/1068038193868460032 https://mobile.twitter.com/feross/status/1068038193868460032 It worked pretty well on my keyboard! Let me know if you'd prefer me to delete the tweet.
- santaragolabs 8y agoThat work is great. A related paper on earlier work where traffic analysis on skype was being done and where the researchers were able to extract individual phonemes and then reconstruct speech that way. It's one of my favorite papers. It's titled "Phonotactic Reconstruction of Encrypted VoIP Conversations" and you can find it here: http://www.cs.unc.edu/~fabian/papers/foniks-oak11.pdf http://www.cs.unc.edu/~fabian/papers/foniks-oak11.pdf
- tialaramex 8y agoVoice is in a way the easy case, because we know the antidote. Constant Bitrate (CBR) mode of an audio codec consumes the same amount of bandwidth regardless of what is transmitted, which is inefficient but secure. As I understand it Signal's voice chat is Opus in CBR mode. Other scenarios are trickier and may need custom work. For example Encrypted SNI currently requires a host to pick a maximum name length, the encrypted name may be any of those names configured on the host, and is padded to that length so that an adversary can't guess which name from the length. Because we don't have a general solution, TLS 1.3 defines an zero overhead optional padding, you can add extra bytes of padding to any TLS message but neither TLS itself, nor the HTTPS binding defines a "good" way to use this padding to shield users from analysis of content based on size because there is no general solution known.
- fromthestart 8y agoIt sounded too good to be true, then I found the caveat: a neural net must be trained on a specific keyboard. Still pretty cool, though.
- ifoundthetao 8y agoI haven't been able to recreate this work from the repos. I emailed them too, about a year ago, with no response. Have you had any success? If so, would you be willing to share?