5 ms·
I tested a few recordings of an Indian Swami giving speeches in English back in the 70's. Recording had a lot of background noise, not great. You have to listen
by hobbitstan 4y ago
I tested a few recordings of an Indian Swami giving speeches in English back in the 70's. Recording had a lot of background noise, not great. You have to listen very carefully to hear what is being said. I was hoping for good results, but...
Results:
- background noise was reduced
- some previously clear words are turned into garbled non-words
- some parts are replaced by a different Indian voice, I assume AI, so it sounds like multiple people talking
All in all, the results are not anywhere near what the sample shows.
- MengerSponge 4y agoTime is a flat circle. Remember when Xerox copiers would randomly replace digits with different digits? https://www.theverge.com/2013/8/6/4594482/xerox-copiers-randomly-replacing-numbers-in-documents https://www.theverge.com/2013/8/6/4594482/xerox-copiers-rand...
- hobbitstan 4y agoWell, that’s a nightmare I never knew existed…
- masswerk 4y agoEpic talk (in German, David Kriesel, "Traue keinem Scan, den du nicht selbst gefälscht hast"): https://www.youtube.com/watch?v=7FeqF1-Z1g0 https://www.youtube.com/watch?v=7FeqF1-Z1g0
- Bombthecat 4y agoWatched it two times! It is very well made and funny!
- pragmatick 4y agoThey should've added english subtitles by now.
- Moru 4y agohttps://www.youtube.com/watch?v=c0O6UXrOZJo https://www.youtube.com/watch?v=c0O6UXrOZJo
- paxys 4y agoGood news: This fancy new compression reduces the error rate from 5% to 1% for the same file size Bad news: While the 5% was a minor inconvenience for customers, the 1% is bad enough to end your company
- aspyct 4y agoI love it how they say "yeah, character substitution is a known issue". How was that ever fine?
- MengerSponge 4y agoClose the ticket. My performance review scorecard doesn't include "known issues"
- shahules 4y agoHi there, I have made a free open-source tool that does better. Care you check that out? https://github.com/shahules786/mayavoz https://github.com/shahules786/mayavoz
- simfree 4y agoThis is very neat! Have you done any profiling on what codecs and sample rates this performs best on? Just curious how the performance differs between PCMU @ 8khz compared to Opus @ 48k or IMBE and AMBE+2 (Project 25 Public Safety audio codecs) :D My dream would be doing audio processing in real time to clean up the audio of phone calls
- shahules 4y agoHi, most models performs best at 16KHz. Current architectures does not support real-time speech enhancement but I plan to add that in future.
- echelon 4y agoIn time, these tools will gain control knobs and eventually start to focus on longer tail audio recovery tasks. I have hope for our old audio. Where there's signal, there's a way.
- godelski 4y agoHonestly, it sounds like you're judging on a pretty big outlier example. The sample seems to more be aimed at background noise and even that sample is extremely easy to understand without the enhancement. There are a bunch of tools out there that are probably better aimed at your goals.