3 ms·
The creator of Talon has tested the Whisper models extensively[0]. [0]: https://twitter.com/lunixbochs/status/1574848899897884672 https://twitter.com/lunixboch
by caternoster 4y ago
The creator of Talon has tested the Whisper models extensively[0].
[0]: https://twitter.com/lunixbochs/status/1574848899897884672 https://twitter.com/lunixbochs/status/1574848899897884672
- orbisvicis 4y agoI don't know what type of speech each dataset represents, but the talon results are extremely impressive... I assume it wasn't trained on at least some subset (depending on the train/test split) of this data?
- lunixbochs 4y agoA handful of the datasets I tested are fully held out (I have reason to believe none of the models have trained on them), and talon was trained on none of the dev or test data of any of the datasets in question. Due to whisper's weakly supervised training on a large amount of automatically scraped data and reliance on a bigger language model, it's far more likely whisper had seen some of the test data before.