2 ms·
Thank you for your high-quality technical questions. 1) Audio streams are naturally segmented between those three states. Separating talk from music is rather
by dest 8y ago
Thank you for your high-quality technical questions.
1) Audio streams are naturally segmented between those three states. Separating talk from music is rather easy. The most challenging part is to separate talk from ads (spoken ads) and music from ads (musical ads).
Filtering ads and talk gives you a music-only experience, which is good when you want to work for example.
2) 3) I am not an expert on RNNs. My understanding is that the LSTM keeps the state between each prediction. I will hopefully get back to you with more precise answers.
4) Your idea about the embedded layer N=32 looks very smart. The dataset is, strictly speaking, a mixed bags of 10-second 100% ads, 100% speech and 100% music (with some slight tolerance at the edges of the track). But when labeling data, I have often tried to label contiguous segments of a minute to several minutes. Though, to not spoil the dataset, I often get a discontinuity on transitions (e.g. music -> ads). So in conclusion I would need to create the dataset you describe. Not a big deal I guess.
X) Streams without ads are quite common. E.g. http://www.radiomeuh.com/ http://www.radiomeuh.com/ or https://www.fip.fr/ https://www.fip.fr/
The thing is that you get a very big corpus of whitelisted data. Too big actually. The solution I have used for a while is to monitor the radio metadata (using https://github.com/adblockradio/webradio-metadata https://github.com/adblockradio/webradio-metadata) and downloading musics with youtube-dl. It worked quite well to bootstrap ;)
Is there any way we could keep in touch apart from Hacker News? Feel free to email me if you feel like it.