4 ms·
My hobby involves free-form Internet text, which has similar normalisation problems. I found that for my purposes, character-level models where quite effective
by struct 10y ago
My hobby involves free-form Internet text, which has similar normalisation problems. I found that for my purposes, character-level models where quite effective partly because they were easy to train on not much data and usually quite robust to minor misspellings[1]. I also developed a part-of-speech tagger[2] which might be useful to you, assuming your corpus has verb and noun tags available. If you've got some more questions, my email's in my profile.
[1] http://dracula.sentimentron.co.uk/sentiment-demo/ http://dracula.sentimentron.co.uk/sentiment-demo/
[2] https://github.com/Sentimentron/Dracula https://github.com/Sentimentron/Dracula