4 ms·
Short answer would be availability of tagged data, in turn dependent on historical research, which is a laggy history of who's been funding the research. This
by tfm 10y ago
Short answer would be availability of tagged data, in turn dependent on historical research, which is a laggy history of who's been funding the research.
This is changing, thankfully; check out e.g. the European Language Resources Association's catalogues,
http://catalog.elra.info/ http://catalog.elra.info/
or any of the very incomplete list on wikipedia:
https://en.wikipedia.org/wiki/List_of_text_corpora https://en.wikipedia.org/wiki/List_of_text_corpora
or in fact the various international Wikipediae themselves!
As for this particular project being English only ... itch to scratch probably ;-)
- danieldk 10y agoShort answer would be availability of tagged data, Part-of-speech and syntax-annotated corpora have been available for a long time for many other languages than English. E.g. for German: http://www.ims.uni-stuttgart.de/forschung/ressourcen/korpora/tiger.html http://www.ims.uni-stuttgart.de/forschung/ressourcen/korpora... http://www.sfs.uni-tuebingen.de/ascl/ressourcen/corpora/tueba-dz.html http://www.sfs.uni-tuebingen.de/ascl/ressourcen/corpora/tueb... Dutch: https://www.let.rug.nl/vannoord/Lassy/ https://www.let.rug.nl/vannoord/Lassy/ http://lands.let.ru.nl/cgn/ http://lands.let.ru.nl/cgn/ https://www.let.rug.nl/vannoord/trees/ https://www.let.rug.nl/vannoord/trees/ French: http://www.llf.cnrs.fr/en/Gens/Abeille/French-Treebank-fr.php http://www.llf.cnrs.fr/en/Gens/Abeille/French-Treebank-fr.ph... I think that one of the problems is that NLP research of English typically has a higher chance of getting accepted at conferences.