5 ms·
Could I use something like this to identify which of two or three people is speaking in an audio clip? Assume I can label several samples of each person's speec
by flashman 8y ago
Could I use something like this to identify which of two or three people is speaking in an audio clip? Assume I can label several samples of each person's speech, then present an unlabeled sample for classification.
- RileyJames 8y agoI’m looking for something that can do this as well. Anything out already?
- flashman 8y agoI had a go of it by replacing the drum samples with voice samples (both 1-2 seconds and 3-5 seconds), then removing the features concerned with length and volume. Fiddled with the number of sub-sections per sample, and some of the random forest settings, but never consistently got higher than 77% accuracy between the four speakers. Maybe it would do better with two speakers.
- psobot 8y agoYes! While the technique I used in this post is pretty simplistic, you could use a similar method to solve what's generically known as the [speaker recognition problem](https://en.wikipedia.org/wiki/Speaker_recognition https://en.wikipedia.org/wiki/Speaker_recognition). That said, this is an open problem within audio research, and there are solutions much more complex than what I have here. An example of one open source speaker recognition project is [bob.bio.spear](https://pypi.org/project/bob.bio.spear/ https://pypi.org/project/bob.bio.spear/), which I haven't tried, but looks promising.