7 ms·
Removing 'um' from a recording is harder than it sounds
- dougcalobrisi 4mo agoThis post is mostly about how surprisingly hard it is to cut filler words out of speech cleanly. Apparently, stripping ums isn't a find and replace type thing, because Whisper's timestamps are off by up to a few hundred ms and cutting on them chops syllables or leaves stutters. So, I built a tool, erm, that starts from Whisper's guess, finds where each word actually starts and stops in the audio, and snaps the cuts to silence so there's no click, with ffmpeg doing the splicing. https://github.com/dougcalobrisi/erm https://github.com/dougcalobrisi/erm
- bagvader 4mo ago[flagged]
- rindalir 4mo agoThis is fascinating! I'm going to try this on a certain clip from Jurassic Park.
- sciencesama 4mo agothere is a aah counter in toast master !! this is the software that helps !!
- cadamsdotcom 4mo agoWhat an awesome tool and idea. I’d be keen to see if it can integrate with video editing tools. Ideally it would slice the video in the timeline without actually removing anything, so you can scrub through your video and try with and without each disfluency (thank you - awesome word) & decide case by case which to keep!
- cryptoz 4mo agoReally cool stuff and definitely going to try it; I’m also finding it wild that Google put effort into adding ums and erms into their text to speech model a while back. AI puts it in, AI helps take it out.
- sublinear 4mo agoDisfluencies are not necessarily "filler". They can convey mood or hesitation. Cutting them can change the meaning. A trivial example is "umm... well... (sigh) okay" versus just "okay". Not okay!
- heroprotagonist 4mo agoNot to promote something, but Wispr Flow does that for me automatically if I trigger a setting for it.. While it's a commercial product with a subscription, I spent a long time on the free tier not even hitting their limits until I started using it so extensively that I wanted to pay for it. And I've used Whisper in the past, mostly for tinkering. I tried it for a couple of use cases but haven't touched the base project in a while. But I do regularly use Faster-Whisper-XXL, an open source project based on Whisper, for subtitle generation. Though, for subtitle generation, I decided to support the project and mainly use the non-public build of Faster-Whisper-XXL Pro built for donators to the open source project. The extra features smooth out the subtitle editing process very substantially. Toss in "--roformer_overlap 0.125 --roformer_vram 16 --best_of 15 --ff_vocal_extract mb-roformer --vad_method pyannote_v3" to the cli parameters (and sometimes --realign) and you have much less work to do in SubtitleEdit or Tero Subtitler afterwards to clean it up.
- dotancohen 4mo agoIs love to hear more about subtitle generation. Specifically, can you label different speakers? I'd be using this for meeting transcription. Thank you.
- heroprotagonist 4mo agoYeah, that's in faster-whisper-xxl via the --diarize parameter with additional options to tweak how it works: https://github.com/Purfview/whisper-standalone-win/discussions/322 https://github.com/Purfview/whisper-standalone-win/discussio... I haven't used it when subtitling, though, so I don't know much more.
- dotancohen 4mo agoTerrific, thank you.
- iib 4mo agoSurprisingly, it's the whisper model itself that does that. I find that it's also good with false starts, often correcting something like: "uhm, we could...we can go there" to just "we can go there", if spoken rapidly enough.
- wzdd 4mo agoIt’s a nice engineering approach, but I’m interested in the motivation. Um and ah is distracting in a transcript, where you can naturally pause to take in information; in speech however it can serve as a focusing point to indicate the next part is important. See https://medium.com/better-humans/dont-worry-about-saying-um-effective-public-speaking-includes-filler-words-e56304416a90 https://medium.com/better-humans/dont-worry-about-saying-um-... for example. The weirdly obsessive zeal that orgs like Toastmasters have about eliminating them is weird. Disfluencies aren’t necessarily bad even if the word starts with “dis”!
- siriaan 4mo agoOccasional ums and ahs are fine but when every other phrase starts with a long aaaaah it can be pretty unpleasant to listen to.
- sans_souse 4mo agoSo, if this project's source Audio were Beavis and Butthead, you would be enthused?
- netsharc 4mo agoI saw a video where the speaker spoke his words quickly, but had long pauses every words. Luckily NewPipe has a "fast forward during silence option". Looking at it again he'd pause, probably trying to find the next word, doesn't find it, and goes "aaaah". So watching at >100% speed and with skip silences saved my sanity: https://www.youtube.com/watch?v=dCO633KE7RA https://www.youtube.com/watch?v=dCO633KE7RA
- toast0 4mo agoHaving heard radio interviews with and without 'internal editing' to remove ums and ahs, most of the time I'd rather the edited version. It's more concise and focused, and I find it easier to comprehend. Too many ums and ahs and my mind wanders, and if it's radio, I can't go easily go back to try again. When I've listened to podcasts or audiobooks, I could never easily go back a little to try again either, and I gave up on them (even though I have some content I really want to listen to, it's too frustrating, so it's not happening). But I'm sure other people have different preferences. I also don't care for writing that could have been made a lot more concise. It's a lot of work to make things shorter, but I think it's worthwhile.
- alok-g 4mo agoI would love to see support for videos and removal of custom filler words (I say 'basically' and 'like' a lot and have so far failed to improve myself on this).
- dougcalobrisi 4mo agoIt does take videos (like mp4) as input but will only output the stripped audio track. I might add the custom filler word functionality and/or perhaps just make the filler word list configurable.
- supernes 4mo agoThis approach seems kind of backwards to me. Why try to detect everything except the thing you're trying to remove instead of either sampling a few uhs and ums and treating them as noise to be silenced (with a sharp crossfade to the noise floor that doesn't interrupt speech flow) or finetuning a model to detect them specifically for full automation?
- pdpi 4mo ago> instead of either sampling a few uhs and ums and treating them as noise to be silenced If you're not paying ttention, ctting out specific sounds can easily cause more trouble. I for one would be quite pset if I couldn't hear the pire's reasoning for calling a foul.
- npodbielski 4mo agoI think it is harder to remove those from your own speech. I have been doing that for few months now and I still get back at it when I am in hurry or stressed.
- ifwinterco 4mo agoIn my experience native English speakers are particularly bad, generally when speaking a second language people are less likely to add random filler words. Also the type of filler word for some reason is often different between UK and US: British people tend to be "umm"-ers and Americans are more likely to add "you know" (although "umm" is also common). Once you notice it it's impossible to ignore and many, many native English speakers are actually terrible at speaking and add filler words to the point where it's very distracting
- lavaman131 4mo agoThis is great, I've tried out automated podcast editing tools before and they cut too aggressively in my experience. What are you thinking about doing next with this now that you've gotten the alignment snapping working cleanly for 'um' and 'ah', are you thinking of expanding the tool?
- deleted 4mo ago[deleted]
- cyberax 4mo agoBTW, any recommendations for AI tools that remove the laugh track? I don't even mind the awkward acting without the missing laughter.
- chrismorgan 4mo agoI think the “What it won’t touch” section shows why the entire concept is unsound. Here it is with a different first sentence, and (other than the third sentence no longer matching erm’s reality) it’s perfectly coherent: > It leaves um, uh, er and elongated versions (ummmm, uhhhhh) alone. Those sound like fillers but they’re doing real work in the sentence, and cutting them automatically would change what someone said. The rule erm follows: only remove things that are sound, not language. > It also doesn’t touch repeated words, false starts, or long thinking pauses. Those aren’t noise on top of the speech; they are the speech, just messier than the speaker would like. Cleaning them up is an editorial decision about which take to keep, and erm doesn’t have an opinion about that. Think about it. Cleaning these things-that-can-be-just-sounds-but-can-also-very-much-be-load-bearing up is an editorial decision. At the very least, you need to judge based on the surrounding content whether the removal of an um would change the meaning at all; and I don’t think text alone is adequate for that.
- thaumasiotes 4mo ago>> It leaves um, uh, er and elongated versions (ummmm, uhhhhh) alone. Something's already gone wrong here. Uh and er refer to the same sound. Uh is the American spelling. Er is British; to them a following "r" like that is just a kind of vowel.
- chrismorgan 4mo agoUm… no. Quite different vowel sounds. (Also, in case it wasn’t clear: I was quoting from the start of the article in that sentence.)
- thaumasiotes 4mo agoThey're quite different vowel sounds in the same sense that "back" and "back" use "quite different vowel sounds" when pronounced by American vs British speakers. But not in any other sense. > in case it wasn’t clear: I was quoting from the start of the article in that sentence. You don't seem to be quoting from the article at all, actually. You've combined two different sentences in a way that grossly misrepresents what the article says. But that's not really relevant to the point here.
- rbbydotdev 4mo agoI wonder if with enough input data and transcription you could “fingerprint” where a speaker personality has habits of interjecting “ums” leading to more hardy analysis. Novel approach, but gets me thinking
- monster_truck 4mo agoIt takes about 30 seconds in Audacity and will give an infinitely better result. Also works on any other sound
- HeavyStorm 4mo agoDoesn't sound true. Unless audacity already has a tool for this exactly... How would you do it on 30 seconds or less?
- ghaff 4mo agoIt doesn't and ums aren't the only consistent tic you often want to clean up--"you know," long pauses, etc.
- alyssamazz 4mo agoI’ve don’t this in audacity many times, it doesn’t work as well. All the umm patterns don’t match exactly. I’ve had better overall results with erm. I haven’t used audacity in years for this, maybe they improved the feature.
- HeavyStorm 4mo agoWhat a very cool utility.
- johnwheeler 4mo ago[flagged]
- boodleboodle 4mo agoThis resonates with our crusade to eradicate Ums once and for all. - Ums Considered Harmful: https://hamanlp.org/research/ums/ https://hamanlp.org/research/ums/ - Related paper: https://hamanlp.org/SIGBOVIK_2026.pdf https://hamanlp.org/SIGBOVIK_2026.pdf
- 1317 4mo agoLooks interesting, would be a nicer article though if there was a demo with before/after to show the results, and why the previous ideas didn't work for something dealing with audio you do need to play the audio really
- ghaff 4mo agoWhen I was doing podcasts regularly, it made me acutely aware of various people's speech mannerisms. (Somewhat similarly, recording a lot of videos during COVID made me very aware of a variety of my own mannerisms--especially overactive hand motions.)
- slhck 4mo ago> Two small fixes, in order. First, each cut endpoint is allowed to slide a tiny bit (up to 60ms) to land in the quietest spot nearby. If there’s a momentary lull in the audio just before or after the original cut point, slide there. The slide is bounded so it can’t cross into a neighboring word, otherwise you’d chew off real speech. Second, from that quiet spot, the endpoint snaps to the nearest moment when the waveform is exactly crossing zero. Oh, Claudish striking again.
- Retr0id 4mo agoI call it claudeslop but I suppose claudish is slightly less inflammatory.
- fragmede 4mo ago... No, you run an entire second pass LLM over the output of Whisper. "no uhhh three no four." should just output four the numeral not even f.o.u.r. Hi, my name is fragmede. Judging by the date on my computer it's been four months since it's since I've t touched the transcription directory on computer and tried to improve on the state of wisprflow. Mines pretty good but it just doesn't... ah you can't drag me back in.
- __mharrison__ 4mo agoInteresting. I make a bunch of video content and I went another way. When I want to redo a section, I say it again. But, I have a magic word — "mistake" — that I insert before. Previously I transcribed and just removed the sentence (or section) before mistake. I recently automated this and used AI to determine what to cut and to drive davinci resolve to make the edit. Saves a lot of time in my workflow.
- ralferoo 4mo agoThe title of the article is wrong. It's not that removing 'um' from a recording is hard, it's that not removing everything else in the recording while doing so is.
- dougcalobrisi 4mo agoYou’re right. I may borrow that if I do a follow up at some point :-)
- AaronAPU 4mo agoI accidentally learned how disgusting people’s mouth noises are while developing an audio leveler. The lip smacking and snot noises between sentences are the stuff of nightmares if you don’t do anything to exclude them from amplification. The best approach I could come up with was to maintain a sliding histogram of loudness and exclude the low-level outliers. You can do more in the noise/frequency domain but those were outside the scope of this tool.
- stavros 4mo agoMisphonia sufferers unite!
- josefritzishere 4mo agoI used to do this with a razor and an aluminum cutting block.
- alyssamazz 4mo agoDoug is a friend, but I actually use this so figured I’d chime in. I make online course content and used to lose close to a full day cutting filler out of every hour or so of recording. This gets me maybe 70% of that time back. On whether you should even cut them, I don’t think it’s clear cut. With non-native English speakers especially, the um is usually a real pause before they say something that matters, and cutting it makes them choppy or changes what they meant. Most of the time though it’s just padding. That matters more for courses than it sounds like it should, because a common complaint I get is how long courses are, so any dead air I can pull out is time I give back to people. Anyway this is in my workflow now. Still messing with the settings to get it right, but I like to mess with my stack and this focuses on this step for me.
- BugsJustFindMe 4mo agoI find the crusade against 'um' to be annoyingly misplaced. It frustrates the shit out of me that iOS speech-to-text dictation refuses to write my 'um's and 'uh's with no way to change that behavior. If a person asks to remove them, fine, but don't fucking alter my speech patterns when I'm sending messages to people.
- neves 4mo agoDoes it work just for English?
- t0bia_s 4mo agoGreat approach. Please, do other languages. I would appreciate Czech!
- ternaryoperator 4mo agoTake this difficulty and make the desired sound piano and put the whole thing into 1960s technology and you can see why recording studios were never able to remove Glenn Gould's humming from his recordings.
- watchlight 4mo agoTranscription is a portion of my own pipeline, and I was a bit surprised to learn at how many of the short words were cut out as a result of VAD (voice activity detection) being overly sensitive. It wasn't like the models were bad, rather just the activity detection wasn't tuned correctly for background noise. The start and end of words would also occasionally get truncated, "um" and "uh" just often happen to be something that folks mutter. They're sufficiently short that the blips aren't considered "real enough speech" to get flagged. Faster-whisper does allow for VAD tuning via API. I don't think it's exposed natively for Whisper though.