4 ms·
My accent costs me 30 IQ points on Zoom. So we built an ML model to fix it
- artavazdsm 7mo agoCo-founder of Krisp here. 1.5B non-native English speakers in the workforce, 4x native — yet all comms infra is optimized for native accents. We spent 3 years building listener-side, on-device accent understanding. The hard parts: no parallel training data exists, the accent space is infinite, accent is entangled with voice identity, and it runs on CPU under 250ms latency. Built in Yerevan, Armenia. Beta is live and free. Happy to go deep on the ML side.
- AlexeyBelov 7mo agoWhat do you think about the misuse potential (by scammers for example)? Aside from that, I like that this exists now.
- davitb 7mo agoThis is for listener-side, not speaker-side. So no misuse case here.
- astipili 7mo agowill it help the barista in Starbucks get my name right finally?
- lu_mn 7mo ago[flagged]
- snek26 7mo agoCurious whether wav2vec-style embeddings played a role in your representation learning.
- sohanyan 7mo ago[flagged]
- Flora_H42 7mo ago[flagged]
- Nathanf22 7mo ago[dead]
- bebelovejan 7mo agoI would like to use such model but only if it really preserves my voice, otherwise people would understand its not me or I have to use it all the time.
- davtyan96 7mo ago[dead]
- imuradyan 7mo agoOn-device CPU inference is the real flex here! Optimization probably mattered as much as modeling.
- sssnowgirl 7mo agoThis is a game-changer! I remember each and every call I had with an investor and feeling shy asking "can you repeat?"... thanks krisp, you changed my life!!!
- KarineS 7mo ago[flagged]
- gyumjibashyan 7mo agoHow did you estimate the number of IQ points?
- arshakarap 7mo agoThis is built for international, privacy-first teams!
- eharutyunyan 7mo ago[dead]
- armsuro 7mo agoThis feels adjacent to voice conversion research, but with stricter latency constraints.
- amartiro 7mo agoThe parallel data is a problem here — you can’t crowdsource ground truth because no one can record themselves with a different accent.
- aharutyunyan 7mo agoAccent space is effectively infinite. Generalization must rely on invariants rather than enumeration.
- CyberSec86888 7mo ago[flagged]
- deleted 7mo ago[deleted]
- deleted 7mo ago[deleted]
- melkman42 7mo ago[dead]
- Talkative123 7mo ago[dead]
- nareksardaryann 7mo agoGreat work. Natural + clear is the combo that matters.
- Hripsimeh 7mo ago[flagged]
- armb21 7mo ago[flagged]
- rasjonell 7mo agoLatency can destroy conversational rhythm. What’s your p95 inference time? also are there any benchmarks we can see?
- Tatevik_H 7mo ago[flagged]
- aris_hovsepyan 7mo ago[flagged]
- Narek21 7mo agoThis feels adjacent to voice conversion research, but with stricter latency constraints.
- tritont 7mo agoNice to finally see this direction of accent conversion (that is on incoming calls) in the Krisp app. This is a very meaningful feature.
- liatitanyan 7mo ago[dead]
- nkhachatryan 7mo ago[dead]
- Ani_Kh1 7mo agoCurious whether wav2vec-style embeddings played a role in your representation learning.
- zmkoyan 7mo ago[dead]
- MarAraqelyan 7mo agoReally cool to see accent adaptation in real time — curious about benchmarks and how well this handles messy, real Zoom calls
- achobanyan 7mo agoLocal CPU inference stands out. Careful optimization likely rivaled the modeling effort.
- zkhalapyan 7mo agoYeh, this would be helpful for the Singlish friends of mine out there!
- 1ilit 7mo agoOn-device CPU inference is the real flex here. Optimization probably mattered as much as modeling.
- felline 7mo ago[dead]
- arkobel 7mo agoThe lack of parallel accent data makes this fundamentally unsupervised. Curious if this leans more on latent disentanglement than direct supervision.
- stepansargsyan 7mo ago[dead]
- eghambaryan 7mo ago[dead]
- argishtia 7mo agoWithout full utterance context, homophones must be tricky. How do you avoid semantic drift?
- intigran 7mo agoEdge-side CPU inference is the quiet power move. Feels like the engineering grind on optimization carried just as much weight as the model architecture itself.
- OpheliaM 7mo ago[dead]
- maraim99 7mo ago[dead]
- grigshahverdyan 7mo agoWithout full utterance context, homophones must be tricky. How do you avoid semantic drift?
- grigoryan2001 7mo agoIdentity preservation seems harder than accent mapping itself. How do you measure that rigorously?
- annahov 7mo agoCan users control the degree of accent modification?
- deleted 7mo ago[deleted]
- gurgenh 7mo ago[dead]
- hackergar 7mo agoIdentity preservation seems harder than accent mapping itself. How do you measure that rigorously?
- jappleseed987 7mo ago[dead]