12 ms·
Alterego: Thought to Text
- oldfuture 1y agolearn more on: https://www.media.mit.edu/projects/alterego/overview/ https://www.media.mit.edu/projects/alterego/overview/ adding also their press release here: https://docsend.com/view/dmda8mqzhcvqrkrk/d/fjr4nnmzf9jnjzgw https://docsend.com/view/dmda8mqzhcvqrkrk/d/fjr4nnmzf9jnjzgw
- lukebechtel 1y agobeen waiting for something like this. Looking forward to adoption!
- pedalpete 1y agoI'd love to get a better understanding of the technology this is built with (without sitting through an exceedingly long video). I suspect it's EMG though muscles in the ear and jaw bone, but that seems too rudimentary. The TED talk describes a system which includes sensors on the chin across the jaw bone, but the demo obviously has removed that sensor.
- fxwin 1y agoi think this is what you're looking for: https://www.media.mit.edu/projects/alterego/publications/ https://www.media.mit.edu/projects/alterego/publications/
- jackthetab 1y agoThirteen minutes is an "exceedingly long video"?! Man, I thought I was jaded complaining about 20 minute videos! :-) I want to know is what are the connected to? A laptop? A AS400? An old Cray they have lying around? I'd think doing the demo while walking would have been de riguer. Anyway, tres cool!
- ilaksh 1y agoMaybe they have combined an LLM or something with the speech detection convolution layers or whatever they were doing. Like with JSON schemas constraining the set of available tokens that are used for structured outputs. Except the set of tokens comes from the top 3-5 words that their first analysis/network decided are the most likely. So with that smarter system they can get by with fewer electrodes in a smaller area at the base of the skull where cranial nerves for the face and tongue emerge from the brainstem.
- zknowledge 1y agoeither this is the world's biggest grift OR the 2nd greatest product of the 21st century... so far.
- Theodores 1y agoThe presentation of this product reminds me of peak crypto when a 'white paper' and a two-page website was all anyone needed to get bamboozled into handing their money over.
- Dilettante_ 1y agoI want to believe so bad that I can finally get rid of my keyboard.
- synapsomorphy 1y agoThe accuracy is going to be the real make or break for this. In a paper from 2018 they reported 92% word accuracy [1]. That's a lifetime ago for ML but they were also using five facial electrodes where now it looks confined to around the ears. If the accuracy was great today they would report it. In actual use I can see even 99% being pretty annoying and 95% being almost unusable (for people who can speak normally). [1] https://www.media.mit.edu/publications/alterego-IUI/ https://www.media.mit.edu/publications/alterego-IUI/
- ivape 1y agoWhy do you say that? I often vocalize near giberrish and the LLM fixes it for me and mostly gets what I meant.
- thatxliner 1y agoWhisper-level transcription accuracy should be sufficient
- blixt 1y agoI found it interesting that in the segment where two people were communicating "telepathically", they seem to be producing text, which is then put through text-to-speech (using what appeared to be a voice trained on their own -- nice touch). I have to wonder, if they have enough signal to produce what essentially looks like speech-to-text (without the speech), wouldn't it be possible to use the exact same signal to directly produce the synthesized speech? It could also lower latency further by not needing extra surrounding context for the text to be pronounced correctly.
- akdor1154 1y agoFrom memory, i think other recent research is along this approach, but not yet good enough. Cant remember where I read this but was likely HN. I think the posted paper got 95% accuracy when picking from a known set of target sentences/words, but far less (60%?) when used for freeform input. I'm sure that's not the last word though!
- stevage 1y agoInteresting, I remember reading a sci-fi book a long time ago with almost exactly this same method, which they called "sub-vocalisation". (I think it was https://en.wikipedia.org/wiki/Oath_of_Fealty_%28novel%29 https://en.wikipedia.org/wiki/Oath_of_Fealty_%28novel%29 but can't find enough details to confirm.)
- goopypoop 1y agoSpeaker For The Dead - Orson Scott Card
- com2kid 1y ago> they seem to be producing text, which is then put through text-to-speech (using what appeared to be a voice trained on their own -- nice touch). This is an LLM model thing. Plenty of open source (or at least MIT licensed) LLMs and TTS models exist that translate and can be zero shot trained on a user's speech. Direct audio to audio models tend to be less researched and less advanced than the corresponding (but higher latency) audio to text to audio pipelines. That said you can get audio->text->audio down to 400ms or so latency if you are really damn good at it.
- stevage 1y agoThe great thing about a product like this is that it's so easy to fake it in video. I don't really buy that typing speed is a bottleneck for most people. We can't actually think all that fast. And I suspect AI is doing a lot of filling in the gaps here. It might have some niche use cases, like being able to use your phone while cycling.
- dllthomas 1y agoTyping speed is very much a bottleneck when I'm washing dishes, at least.
- lennxa 1y agotalk then
- dllthomas 1y agoNot ideal, between sound from the washing itself plus whatever else is going on in the house, and I'd rather not add noise that others have to deal with. That said, it's certainly in the mix as competing with other potential solutions.
- prerok 1y agoHow about when baking? https://xkcd.com/341/ https://xkcd.com/341/
- dllthomas 1y agoThen, too. I cannot compete with Mrs Roberts.
- Bjartr 1y agoPersonal anecdote: I do find typing to be a bottleneck in situations where typing speed is valuable (so notes in meetings, not when coding). I can break 100wpm, especially if I accept typos. It's still much, much slower to type than I can think.
- lordofgibbons 1y agoIs the video hosted anywhere else? It seems to be the only source of info, but it's being played at like 0.5X speed.
- oldfuture 1y agomore here: https://www.media.mit.edu/projects/alterego/overview/ https://www.media.mit.edu/projects/alterego/overview/ check also the publications tab, and this pr: https://docsend.com/view/dmda8mqzhcvqrkrk/d/fjr4nnmzf9jnjzgw https://docsend.com/view/dmda8mqzhcvqrkrk/d/fjr4nnmzf9jnjzgw
- tibbon 1y agoSo... Sub-Etha?
- deleted 1y ago[deleted]
- bigolnik 1y agoSo this a bone conducting microphone? That operates at the speed of speech? While you sit around awkwardly, hoping no one talks to you? This isn't thought. This is you saying to yourself quite clearly what you would like it to hear.
- reassess_blind 1y agoIt doesn't look like he's speaking.
- socalgal2 1y agoI just imagine this going really wrong. My chain of thought would be something like: "Let's see, I need to rotate this image so I need to loop over rows then columns, .. gawd fuck this code base is shit designed, there are no units on these fields, this could be so much cleaner, ... for each row ... I wonder what's for lunch today? I hope it's good ... for each column ... Dang that response on HN really pissed me off, I'd better go check it ... read pixel from source ... tonight I meeting up with a friend, I'd better remember to confirm, ... write pixel to dest ...."
- desireco42 1y agoWhat I picked up from this vision of the future... we will have mind reading devices to capture out thoughts, but we will still be on a train and commuting to work... dang... So they came up with this groundbreaking idea but couldn't come up with better use case then typing on a train. Look, I can't but not appreciate that at least they are doing something interesting as opposed to vibe one shot fork of vs code things that we see.
- g42gregory 1y agoAre these are friends of Elizabeth Holmes? :-)
- soulofmischief 1y ago> We currently have a working prototype that, after training with user-specific example data, demonstrates over 90% accuracy on an application-specific vocabulary. The system is currently user-dependent and requires individual training. We are currently on working on iterations that would not require any personalization. https://www.media.mit.edu/projects/alterego/frequently-asked-questions/#faq-how-accurate-is-the-devicesystem https://www.media.mit.edu/projects/alterego/frequently-asked...
- andymatuschak 1y agoThat text was written about the Media Lab-era prototype in 2019: https://web.archive.org/web/20190102110930/https://www.media.mit.edu/projects/alterego/frequently-asked-questions/#faq-how-does-the-system-work https://web.archive.org/web/20190102110930/https://www.media... I wonder how far they've gotten past it.
- oliwary 1y agoThe key to technology like this is how quickly it can get to say 99.5%. Not convinced it will be an easy path!
- djhn 1y agoEven speech to text has a long way to go to reach 99.5%!
- olejorgenb 1y ago"AlterEgo reads information from the peripheral somatic system through internal speech movements, rather than directly from the brain. It detects the signals users send to their mouth and vocal cords when deliberately, but silently, voicing words. "
- taneq 1y agoSo it's a Thalmic Myo for your trachea?
- andsoitis 1y agoThey don’t have something that anyone can try out and it also seems no public demonstrations of early prototypes. Seems like vaporware.
- desireco42 1y agoReminds of the song Sound of Silence...
- Briannaj 1y agoThis is literally only as fast as text to speech. the only difference is that you don't have to speak aloud. Which is cool. But for using a computer its still annoying and worse than a mouse because with a mouse you can click or drag and place in a second, in this format you have to think "move the box from point A to point B (with coordinates or a description) etc etc". I think its cool, I've been brainstorming how a good MCI would work for a while and didn't think of this. I think its a great novel approach that will probably be expanded on soon.
- Briannaj 1y agowhat it could be really cool for is stuff like "open my house door", "Turn off the lights", "text so and so", "Start my car" Stuff we want to do without pulling out our phone that doesn't require a lot of detailed instruction.
- stevage 1y agoI must be such a rarity around here, but if I could improve a hundred things about my life, none of those would make the list. Well, possibly the third one - more convenient ways to text people. I guess I also kind of enjoy the physical sensations of putting a key in a lock, opening the door etc. Definitely don't want a digital-only existence.
- com2kid 1y ago> But for using a computer its still annoying and worse than a mouse because with a mouse you can click or drag and place in a second, in this format you have to think "move the box from point A to point B (with coordinates or a description) etc etc". You wouldn't use a regular WIMP[1] paradigm with this, that completely defeats the advantages you have. You don't need to have a giant window full of icons and other clickable/tappable UI elements, that becomes pointless now. [1]https://en.wikipedia.org/wiki/WIMP_(computing) https://en.wikipedia.org/wiki/WIMP_(computing)
- gcanyon 1y agoFor those thinking about speed: an average human talks anywhere from 120-240 words per minute. An average human who touch types is probably 1/3 to 1/2 as fast as that, while an average human on a phone probably types 1/5 as fast as that. But for me speed isn't even the issue. I can dictate to Siri at near-regular-speech speeds -- and then spend another 200% of the time that took to fix what it got wrong. I have reasonable diction and enunciation, and speech to text is just that bad while walking down the street. If this is as accurate as they're showing, it would be worth it just for the accuracy.
- keleftheriou 1y agoI agree, but I think LLM-based voice input is a lot better. I’m using OpenAI’s realtime API for my Apple Watch app, and it does wonders, even editing can be as simple as “add a heart emoji at the end”, and it just works. https://x.com/keleftheriou/status/1963399069646426341 https://x.com/keleftheriou/status/1963399069646426341
- deckar01 1y agoYou can tell it’s fake, because it’s hard wired and super low profile, yet isn’t covered in LEDs.
- laurieg 1y agoI'm a huge smart speaker user. I have one in every room. But as soon as guests come over I stop using them. I would never use Siri etc in public. Going from voice input to silent voice input is a huge step forward for UX.
- keleftheriou 1y agoI get the sentiment, but can you elaborate on why that is the case for you?
- laurieg 1y agoIt is instantly awkward. If you are there with one other person and you start talking, they assume you are talking to them.
- baroninthetrees 1y agoAs someone with ADD and a lot of crosstalk in my "inner voice", I can't imagine this could make any sense of what I intending, let alone one thing. Definitely a lot of use cases if it isn't vaporware.
- paulbjensen 1y agoI wonder if they've considered testing it with people who have locked-in syndrome or Motor-Neurone disease. It could be an amazing tool for them.
- whymauri 1y agoThe Harvard BIONICS lab is working on neuroprostheses for different forms of paralysis, like intestinal paralysis. They're great.
- deadbabe 1y agoGreat. I can imagine fucked up charlatans putting stuff like this on brain dead patients and convincing their family the person is still able to communicate with the help of AI.
- goopypoop 1y ago"spirit box" don't need Al
- bromanko 1y agoI think the killer app is doing video calls in coffee shops without disturbing my neighbors.
- com2kid 1y agoI am surprised no one here has noted that a device like this almost completely negates the need for literacy. That is huge. Right now people still need to interact with written words, both typing and reading. Realistically a quiet vocal based input device like this could have a UX built around it that does not require users to be literate at all.
- jussaying2 1y agoNot to mention the support it brings for people with disabilities! (speech, hands/fingers)
- aDyslecticCrow 1y agoAmazing for paralysis and other sevear physical disabilities. Similar tech is already widely researched for years. But I'm sceptical about this specific company with the lack of technical details.
- aDyslecticCrow 1y agoHow convenient! Literacy has always been a thorn to efficient society, as books too easily spread dangerous heretical propaganda. Now we can directly filter the quality of information and increase cultural unification. /j
- com2kid 1y agoThat is my fear yeah, a continued dumbing down of society. Literacy rates in the US are already garbage, this device may just make it worse. If people never have to read or write, why would they bother learning how?
- giveita 1y agoHey Google make a note to pack my hiking boots. Done.
- com2kid 1y ago
- dwa3592 1y agoY’all are missing a few key points. - There is a ML model which was trained on 31 hours of silently spoken text. That’s the training data. You still need to know the red fruit in front of you is called apple bc that’s what the model is trained on. So you must be literate to get this working. - The accuracy in the paper is on a very small text type, numerals. As much as I could understand, they asked users to do mathematical operations and they checked the accuracy on that. Someone with a deeper understanding please correct me. - Most of the video demo(honestly) is meh, once you have the text input for a LLM, you are limited to what the LLM can do. The real deal is the ml model that translates the neuromuscular signals to actual words. Those signals must be super noisy. So training a model with only 31 hours of data is a bit surprising and impressive. But the model would probably require calibration for each user’s silent voice, like say this sentence silently , “a quick brown fox jumped over the rope”. I think this will be cool. - I really hope this tech works. I really really hope they don’t sell to big tech jerks like Meta. I really really really hope this tech removes screens from our lives(or at least a step in the right direction).
- dwa3592 1y agoI should have started with- “Congratulations. Very cool tech if works”.
- crooked-v 1y ago> So you must be literate to get this working. Literacy is about written text, not spoken words. I think you've confused it with fluency.
- deepanwadhwa 1y agoThe apple example seems off, but literacy is still going to be a requirement for this kind of thing to work. First, you need to be able to understand the language(grammar, vocabulary etc) to communicate with this device. The guy in the demo is literally thinking in English, Mandarin. I'd actually argue that only highly literate people will be able to use this device really well. This device is not reading thoughts, it's inferring the silently spoken words.
- vunderba 1y agoFrom the article: > Alterego only responds to intentional, silent speech. What exactly do they mean by this? Some kind of equivalent to subvocalization [1]? [1] https://en.wikipedia.org/wiki/Subvocalization https://en.wikipedia.org/wiki/Subvocalization
- hyperadvanced 1y agoOh god we’re about to have the “I don’t have an inner monologue” debate again, aren’t we?
- balamatom 1y agoI got a whole inner panel discussion!
- ipsum2 1y agoYes. The paper the company is based on uses EMG (muscle movements) to convert into text.
- dinfinity 1y agoIf you look at his facial movements in the video it looks as if he is pretty actively using his facial muscles, 'trying' to speak while moving as little as possible (which would cause the clearest signals to be emitted). If that is what is happening, to me it feels like harder work than just speaking (similar to how singing softly but accurately can be very hard work). It would still be pretty cool, but only practical in use cases where you have to be silent and only for short periods of usage.
- boznz 1y agoSpent all last year writing a techno-thriller about mind-reading, I'm sure this is about as factual, and, of course nothing nefarious could possibly happen if this ever became real.
- deekshith13 1y agoYou probably thought about some nefarious stuff that could happen. Mind to share some interesting ones?
- balamatom 1y agoWell, for starters, there's the one where social consensus decides to define whether a subvocalization is "intentional" by whether the interface responded to it.
- wcrossbow 1y agoThis is the stuff nightmares are made of. We already live in a you have nothing to hide society. Now imagine one where mega corps and the government have access to every thought you have. No worries, you got nothing to hide right? What would that do to our thought process and how we articulate our inner selfs? What do we allow ourselves to even think? At some point it will not even matter because we will have trained ourselves to suppress any deviant thought. I'd rather not keep on going because the ramifications of this technology make me truly sick in the stomach.
- ipnon 1y agoIn the Ghost in the Shell universe, I always thought the telepathic conversations from cyberbrain to cyberbrain was one of the most fantastic and least realistic predictions for the future. But I was clearly wrong. We already have rudimentary telepathy 10 years ahead of schedule!
- p1dda 1y agoBrilliant idea to capture the neuromuscular signals and translate it to text!
- kittikitti 1y agoSo if I unintentionally think of a thought crime, is it still illegal? I wonder how governments will use this to "nudge" their citizens. I guess if you have nothing to hide, then there are no issues at all.
- croes 1y agoWhy did the note sync about the hiking boots lost the "I have to"?
- moezd 1y agoGreat, now if I find myself in a weird dialogue and murmur under my breath this company can store exactly what I called those people. They can also sell that data to whoever pays highest. Tremendous job you guys, as if ad industry wasn't annoying and intrusive as they are currently!
- giveita 1y agoVery impressive but use case is narrow. You are on a train and don't have time to make a note, ok. But most of the time we are in places reasonably private where speech recognition is fine and would be convenient maybe more so learning to say without saying. As a disability speech aid though maybe it would be amazing?
- PMunch 1y agoSpeech for the disabled would indeed be great, as long as the disability doesn't also affect the system which you use to "silent speak". As for the privacy thing, I would say that I absolutely hate talking out loud to my devices. Just the idea of talking my ideas into a recorder in my own office where nobody can hear me feels very strange to me. But I love thinking through ideas and writing scripts for speeches or presentations in my mind, or to plan out some code or overall project. A device like this would allow me to do the internal monologue thing, then turn to "silent speak" them into this device to take notes which sounds great. And the form-factor doesn't look that dissimilar to a set a bone-conduction headsets which would be perfect for privacy-aware feedback while allowing you to take in your surroundings. With this tech demo though it seems like the transmission rate is veeery slow, he sits still in his chair staring into the room and a short sentence is all that appears. Not exactly speed of thought.. And of course there is the cable running off to who knows what kind of computational resources. The AI parts of this are less exciting to me, but as an input device I'm really on-board with the idea.
- colinwilyb 1y agoThere are times when I'm masked (ie: 3M respirator) and gloved at work for 10+ hours at a time. I wonder if speech recognition would be possible without breaking the face seal during silent speech. This could be beneficial for other types of work in hazardous environments. (My current solution is to tear the fingertip off my offhand glove so I can unlock and use my device....)
- phoenixhaber 1y ago[dead]
- runxel 1y agoI've seen tech like this on display at the IFA already 15 years ago. Forgot the name, but it was pretty hyped at that time. You could even demo it live. Sure, it was only to steer a sidescrolling video game character, but it worked great with as little training of half a minute or so. Anyhow, Alterego just seems like another vaporware product, that will never enter or even begin to penetrate the overall market. But let's see!
- kordlessagain 1y agoMakes sense there is so much ego on that page.
- hm-nah 1y agoI’ve thought about a future where all audio is recorded (public, home, work, etc.). If this thing is real, it would allow comms in this dystopian vision. Boo
- darepublic 1y agoYou gotta use it to control your AR hud while out in public
- deleted 1y ago[deleted]
- Tiereven 1y agoIntegrating AlterEgo with the next generation of AR glasses could be the next generation of technology after these electronic bricks we carry in our pockets. My biggest frustration with wake-word assistants is that voice is inherently a broadcast channel. There’s endless comedy about the confusion on a bus when someone's talking into Bluetooth and their neighbor thinks they’re being addressed. Silent Sense + AR gets your eyes up and around you, fixes posture, frees your hands and keeps the guy next to you out of the conversation.
- alexoberneyer 1y agoNice, then I can finally vibe code in a crowded cafe or Coworking space
- bitwize 1y agoBig if real. This is the kind of "AI assisted coding tool" I could find myself using, one where I can think code and commands into the machine. Make it open source and local, and I'm sold!
- qingcharles 1y agoThis video weirds me out because the presenter never looks at the camera. Odd directing choice.