4 ms·
They aren't hiding that, though? Literally the first graphic on the page shows their claim of a "word error rate" on test data has an error rate of around 14% i
by ccooffee 4y ago
They aren't hiding that, though? Literally the first graphic on the page shows their claim of a "word error rate" on test data has an error rate of around 14% in the best case compared to the state of the art at 15% (for en-US content).
It's only 1% better than the current state of the art. But it's still noteworthy. From the end of the abstract:
> We demonstrate that utilizing a large unlabeled multilingual dataset to pre-train the encoder of our model and fine-tuning on a smaller set of labeled data enables us to recognize these under-represented languages. Moreover, our model training process is effective for adapting to new languages and data.
It's amazing to me that the chaotic process of "machine learning" can end up with an internal state for languages that is readily adapted to entirely new languages.
For now, they've got this handling audio transcription, but with some hints that this approach could work well for translation. Perhaps we'll be able to use these improved models to decipher Linear A[0] or other un-deciphered languages. It sounds like "magic", but it's the kind that could maybe exist.
[0] https://en.wikipedia.org/wiki/Linear_A https://en.wikipedia.org/wiki/Linear_A
- lovemenot 4y ago>> It's amazing to me that the chaotic process of "machine learning" can end up with an internal state for languages that is readily adapted to entirely new languages. Yeah and interestingly, that was roughly Chomsky's breakthrough finding with respect to how humans learn language as children. We are born with an innate language acquisition device.
- vkazanov 4y agoChomsky's ideas around the Universal Grammar are just a theory. Similar to how his formal grammars somewhat represent real languages but never fully, the UG model will never explain it all, or even most of it. Brain biology just doesn't like the idea of formal things/rules/grammars. Here's an alternative theory/approach. What if natural languages are just the way the device starts working when the number of neurons grows quickly? NL properties sort of emerge out of low-level details of brain work? Neurons are simple but the brain is not. Complex brain properties emerge from trivial parts the same way our full bodies emerge from a simple DNA/RNA system. Any details in these systems would be too statistical to expose a limited rules system. Obviously, powerful enough ML system can infer the system's properties. In fact, it can infer any function. The thing is that this doesn't mean there's some kind of simpler model explaining details of emergent system's work. What is surprising is the way LLMs imitate a stateful function (our brain, with memory, fluid biology, etc) using a stateless inferred function (the model). I suspect this statefulness might be the answer to the question of "poverty of stimulus" problem.