4 ms·
Classic link. Not the same thing though. I mean try telling a competent agent something like > Let's build a webapp that allows the user to choose an audio fil
by akx 2mo ago
Classic link. Not the same thing though. I mean try telling a competent agent something like
> Let's build a webapp that allows the user to choose an audio file, we slice it into timestamped word regions with transformers.js + onnx-community/whisper-base_timestamped (with language selection). The user can then scroll through the waveform (that shows the found word regions), synchronized with the transcript, and select word regions in the transcript or in the waveform, and fine-tune the selection if they need to, and then export the slice as WAV.
and see what pops out.
- andai 2mo agoSo I tried this with GLM 5.3 and it kinda worked. I used AssemblyAI cause I don't have the ability to run local models. The result is pretty ugly and unintuitive, OP's design is much nicer. And there are desync issues, but it's not too bad for a one-shot. (Kinda crazy that we can even do that these days, and then complain about it instead of being amazed!)