3 ms·
That is what I ended up doing. Except this is raw data. Usually NLP researchers work from corpora that are already segmented into sentences and tagged with pa
by raffi 17y ago
That is what I ended up doing. Except this is raw data. Usually NLP researchers work from corpora that are already segmented into sentences and tagged with part-of-speech tags. I had to go through the hassle of segmenting and marking up the Wikipedia and Project Gutenberg.
- mattrepl 17y agoWould you consider making it available to others? Perhaps the UCI Machine Learning Repository would accept it. If not, I'm sure I could find a way to have it hosted at my university.
- bravura 17y agoarchive.org will host datasets. To preprocess wikipedia, I have used the following software: http://sourceforge.net/apps/mediawiki/wikiprep/index.php?title=Main_Page http://sourceforge.net/apps/mediawiki/wikiprep/index.php?tit... To remove boilerplate from gutenberg requires painfully constructed heuristics. It would be great to have software released to do that.