4 ms·
bert provides incremental improvements on tasks that don't really stress-test contextualized world knowledge in linguistic tasks — contrast with Winograd schema
by glup 8y ago
bert provides incremental improvements on tasks that don't really stress-test contextualized world knowledge in linguistic tasks — contrast with Winograd schemas (SOTA: https://arxiv.org/pdf/1806.02847.pdf https://arxiv.org/pdf/1806.02847.pdf). Corpus data and compute don't magically solve everything.
- p1esk 8y agoThe paper you linked to shows how a simple DL brute force approach (large model trained on lots of data) "successfully discovers important features of the context that decide the correct answer, indicating a good grasp of commonsense knowledge." Which kinda goes contrary to your last sentence. Go figure.
- glup 8y agoMy point is that while the DL + large corpus approach is the best available, it's still way, way lower than human performance (Tables 4 and 5).
- p1esk 8y agoSure, but we are talking about how fast it's progressing. Given that NNs for NLP field pretty much started with Bengio's paper in 2003, I'd say it's accelerating at an amazing rate.
- glup 8y agoBut NNs for language processing have been around much longer than that, e.g., Elman, 1990: https://crl.ucsd.edu/~elman/Papers/fsit.pdf https://crl.ucsd.edu/~elman/Papers/fsit.pdf.