Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
alpe
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
alpe
10y ago
You are welcome. Using a forced aligner usually improves the results a lot when compared to using an automatic speech recognition system --- because adapting the language model to your specific text prunes a lot of choices w.r.t. a generic
2.
▲
by
alpe
10y ago
Another possibility is to just run an automatic speech recognition system (e.g. Sphinx or PocketSphinx can read from the mic input), and align its output with the ground truth text. You need to deal with imperfect matching because the ASR m
3.
▲
by
alpe
10y ago
For sure aeneas is not suitable, since it requires all the text and all the audio in advance. But ASR-based tools in theory would allow such an operation mode, but I have not seen aligners that read from the mic buffer directly or have a bu
4.
▲
by
alpe
10y ago
To elaborate a bit further, as indeed the closed captioning applications are very important, from hearing-impaired people to the dyslexic, to second language learners. Let's think about how a human operator would create captions for a
5.
▲
by
alpe
10y ago
I have used aeneas myself to do it, with mixed results. You will probably need to increase the DTW margin. Also note that you will need a lot of RAM --- say 16 GB if you plan to work on a single audio file with duration 10-15 hours, which i
6.
▲
by
alpe
10y ago
I would like to note once again that aeneas is not based on automatic speech recognition techniques, but on MFCC + DTW, which is an even older approach, with pro's and con's. Interestingly, there are situations where ASR-based for
7.
▲
by
alpe
10y ago
Thank you. Indeed, while aeneas was created for ebook-audiobook synchronization, several of its current users are producing closed captions --- because, in most cases, they already have a clean transcript (e.g., speakers provide transcripts
8.
▲
by
alpe
10y ago
I agree on most of your observations. However, please note that other tools are better suited than aeneas if one wants to align at phoneme level: gentle, Kaldi, SPPAS, etc. aeneas' goals are covering as many languages as possible, fast
9.
▲
by
alpe
10y ago
> Might also be possible to look at the spectrum at any time to possibly identify areas of the file to skip. I would say yes and no. Currently you can add a switch that makes aeneas ignore the audio intervals that are detected as "n
10.
▲
by
alpe
10y ago
Definitely. Actually, aeneas can be used as a Python library (rather than just a CLI tool), and you can definitely provide an audio file, a list of audio intervals where the spoken text is, and align "piece-wise". See the "ae
11.
▲
by
alpe
10y ago
Yes, there are several other open source aligners out there, mostly from academic research or derived from academic projects. In my personal GitHub page I have a repo with an annotated list of forced aligners. (If I add a link to it, the sp
12.
▲
by
alpe
10y ago
Several users of aeneas interested in producing caption files for videos told me that it does. And considering how DTW works, it is plausible. Unfortunately, I have not had the time to setting up a suitable corpus and performing a rigorous
13.
▲
by
alpe
10y ago
Hi, thank you. Having it on conda would be great ( https://github.com/readbeyond/aeneas/issues/158 ), so if you feel like it, it would be wonderful! The two points that proved difficult in packaging aeneas (as
14.
▲
by
alpe
10y ago
In Italian high schools "Licei" we take 5 years of Latin (and also ancient Greek if you choose the classical study path)... nice to meet you!
15.
▲
by
alpe
10y ago
aeneas is not based on ASR (i.e., it does not try to "recognize" words and align them with the input text), but on the "older" MFCC + DTW approach. Hence, it is difficult to give you a precise answer, e.g. in terms of wo
16.
▲
by
alpe
10y ago
Thank you. Indeed several users of aeneas adopted it for producing SRT/TTML files, i.e. captions, for videos, both online and offline --- and many of them start with an existing transcript. However, please note that there are limitatio
17.
▲
Show HN: Aeneas – a Python audio/text aligner
(github.com)
188 points
by
alpe
10y ago
|
35 comments