4 ms·
I agree with your point. That is why we open sourced the benchmark so you can verify it yourself. You just need an Ubuntu machine with python 3.6. I appreciate
by kenarsa 8y ago
I agree with your point. That is why we open sourced the benchmark so you can verify it yourself. You just need an Ubuntu machine with python 3.6. I appreciate the question.
- smt88 8y agoIt's the methodology, not the results or the code, that I'm suspicious of. For a highly variable task like STT, I'm sure an expert could contrive a test that gives far better results for either of the other programs tested. That's why it would help to know why this test is comprehensive and representative enough to be considered unbiased or otherwise where its biases are. I don't have that expertise myself.
- p1esk 8y agoI think the biases are pretty obvious, but the most serious shortcoming of this benchmark is that their result (30% WER on CV) is not reproducible: it's not clear what they trained their model on, and the model itself is not available, so you just have to take their word for it.
- kenarsa 8y agoThanks for the comment. Just wanted to quickly clarify that the model is available here: https://github.com/Picovoice/stt-benchmark/tree/master/resources/cheetah https://github.com/Picovoice/stt-benchmark/tree/master/resou...
- p1esk 8y agoNo, "reproducible" means that I can train it on the same data you used, and get the claimed result. Anything other than that is taking your word for it.
- neikos 8y agoAs someone who has no idea about ML, why is knowing how it was trained important? I would guess to anticipate possible problems, but I don't really know.
- lern_too_spel 8y agoTo make sure it wasn't trained on the test set, even if by accident.
- deleted 8y ago[deleted]
- canada_dry 8y agoKinda like creating your acceptance test using the same scripts as your unit test... the result would look good, but in fact be less than reliable.
- DougMerritt 8y agoMore than kind of like; you nailed it: exactly like. This has been an infamous issue with machine learning for decades, where unwary researchers/developers can do this quite accidentally if they're not careful. The thing is that training data is very often hard to come by due to monetary or other cost, so it's extremely tempting to share some data between training and testing -- yet it's a cardinal sin for the reason you said. Historically there have been a number of cases where the best machine learning packages (OCR, speech recognition, and more) have been best because of the quality of their data (including separating training from test) more than because of the underlying algorithms themselves. Anyway it's important to note that developers can and do fall into this trap naively, not only when they're trying to cheat -- they can have good intentions and still get it wrong. Therefore discussions of methodology, as above, are pretty much always in order, and not inherently some kind of challenge to the honesty and honor of the devs.
- skykooler 8y agoIs there more information on Cheetah anywhere? I was looking at picovoice.ai but it only mentions PORCUPiNE.
- kenarsa 8y agoThanks for the comment. Not currently at the moment. But we are going to provide more information on picovoice.ai in coming days.