4 ms·
Building an efficient neural language model over a billion words
- sharemywin 10y agoOh, that was so yesterday I did that last year with my laptop...oh wait I can't eavesdrop on a billion people's private conversations.
- sp332 10y agoThe corpus was published in 2013. https://github.com/ciprian-chelba/1-billion-word-language-modeling-benchmark https://github.com/ciprian-chelba/1-billion-word-language-mo... (And it seems to only have 0.8 billion words?) Edit: Here's one mentioned in one of the other papers, it has over 7 billion "tokens" which I think includes words and punctuation. https://ibm.ent.box.com/v/booktest-v1 https://ibm.ent.box.com/v/booktest-v1