3 ms·
How did you scrape that data? How do you store and retrieve it? Is it just a standard db or a vector db? Sorry for the questions, but it seems like an interest
by soultrees 3y ago
How did you scrape that data? How do you store and retrieve it? Is it just a standard db or a vector db?
Sorry for the questions, but it seems like an interesting, yet probably common data set and as someone who is venturing down this path, I’d like to learn more about building my own dataset similar to this from scratch.
- neilk 3y ago> standard db or vector db lol, it's a 42MB text file from Google Books Ngrams. The format looks like this: $ head words-all.txt a 14219615690 a! 196012 a" 84 a' 47713 a'0 3036 a'1 4070 a'10 99 a'11 56 I queried it with perl and sort. $ time perl -wlane 'if ($F[0] =~ /^[qwertyuiop]+$/) { print length($F[0]), "\t", $F[0] }' words-all.txt | sort -rn > qwertywords real 0m1.915s user 0m1.896s sys 0m0.025s I can't remember exactly which file I downloaded, but according to my notes I got it from here back in 2012 or so. https://storage.googleapis.com/books/ngrams/books/datasetsv2.html https://storage.googleapis.com/books/ngrams/books/datasetsv2... There seems to be a newer corpus published in 2020: https://storage.googleapis.com/books/ngrams/books/datasetsv3.html https://storage.googleapis.com/books/ngrams/books/datasetsv3...