4 ms·
Their language models might perform acceptably with text, but speech is much looser with structure and grammar. They may apply a bit of local expectation-maximi
by obtu 14y ago
Their language models might perform acceptably with text, but speech is much looser with structure and grammar. They may apply a bit of local expectation-maximisation but anything more strict or long-range wouldn't work.
- koide 14y agoSpeech may be much looser, but a classifier to (try to) detect absolute gibberish from real content doesn't look too far fetched to me. Its only action would be to disable the subtitles by default if gibberish was detected. It may be computationally impractical though.
- obtu 14y agoThe problem with that is that the language model's power is already used to fix things up locally (because this is transcribed from audio). As a result it can't be used again to decide if the transcription fits the model; it's the case by design. There must be some kind of confidence metric at the end of the process, but I don't think it's possible to tell how much of the ambiguity comes from inadequacies in the phoneme model, or the audio environment, or the language model. They'd have to throw out good transcriptions in noisy environments along with bad transcriptions, and probably wouldn't keep much at all. As it is it seems they prefer to publish crappy results and hopefully have some feedback channel involving the video uploader.
- koide 14y agoI'm thinking of the following: audio track of the video ---[process (language recognition)]---> transcription ---[process (classifier based on textual language model)]---> answer to "gibberish?" So good transcriptions would be good, regardless of the noise environment. It might give some false positives, but I expect that a good price to pay to avoid the kind of mess they create now.
- gcr 14y agoWhat does your "transcription" step mean? Either it would transscribe each word "in divy two Ellie, buy it's Elf", which produces garbage without context, or it would additionally have to use the language model to patch things up. For example, the difference between "its" and "it's", "red" and "read", "know" and "no" is irreconciliable without understanding the rest of the sentence these trouble words appear in.
- koide 14y agoTranscription is the textual output of the speech recognition process, be it phonetic, LVCSR or direct. All current applications of them do take some context in consideration, usually via a transition matrix and lots of training data. What I'm proposing is to pass the output of speech recognition through a binary classifier that answers the question "is this text gibberish?", which is trained with the help of a textual language model, unrelated to the speech recognition pass.