2 ms·
> It makes some good points in that the amount of data it would take to approximate it all seems far too large The article's way of estimating this is absurd,
by treesprite82 5y ago
> It makes some good points in that the amount of data it would take to approximate it all seems far too large
The article's way of estimating this is absurd, seemingly relying on the idea that a model can't generalise and must see every possible variation of a sentence:
> If we add to the semantic differences all the minor syntactic differences to the above pattern (say changing ‘because’ to ‘although’ — which also changes the correct referent to “it”) then a rough calculation tells us a ML/Data-driven system would need to see something like 40,000,000 variations of the above
- didibus 5y agoYa, I don't know about their approximation, but I think we've already found ourselves hitting some limits with computing power and data on big NLP ML models. The appearance of custom chips is a pretty good example. Maybe for text we won't run out of data, considering the internet has so much text available. But I still think that the "human learning" is something to consider. Maybe a baby is a statistical machine, and it has a statistical based model that's so good it needs very little data, but it's also possible it uses something more logic based. In any case, I feel at least a baby would combine multiple learnings together, it would have started to learn about the dimensions and shapes and temperatures and properties of various physical objects, while simultaneously learning about the language used to refer to those things. I think this last piece can allow a baby to connect the predictions of real things with the language used, which then helps provide semantics for the language that are taken into consideration by a human.