Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zhangce
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
zhangce
3y ago
Thanks for the suggestion! We will add this in the pool of features for future release. (We are currently running the current 40+ annotations on the `tail` partitions). If you are interested in contributing the code for these features, feel
2.
▲
by
zhangce
3y ago
We did an exact dedup across all 84 dumps; there are 100T tokens before this exact dedup, and 30T tokens after. If we do further fuzzy dedup (we have simhash signatures pre-computed for different similarity level), this can potentially be r
3.
▲
by
zhangce
3y ago
There are actually a few ways to do this; and we have four: - `rps_doc_ml_wikiref_score`: a classifier that classifiers random webpage with Wiki references (used in Llama-1) - `ccnet_perplexity`: perplexity of an LM trained on Wikipedia (us
4.
▲
by
zhangce
3y ago
It is around 100TB (84 CommonCrawl dumps, roughly 1TB per dump)
5.
▲
by
zhangce
3y ago
What we make available is: -- (A) the dataset after pre-processing the raw CommonCrawl data (e.g., text extraction and language identification) and some minimal filtering; and (B) for each document in (A), we also pre-computed 40+ of "