4 ms·
I'm kind of surprised there's 800GB of text in the world.
by eutectic 6y ago
I'm kind of surprised there's 800GB of text in the world.
- luto 6y agoheh, yeah! for comparison: the English Wikipedia is around 40 GB of text. https://en.wikipedia.org/wiki/Wikipedia:Size_of_Wikipedia#Size_of_the_English_Wikipedia_database https://en.wikipedia.org/wiki/Wikipedia:Size_of_Wikipedia#Si...
- pbourke 6y agoIt’s not that much, actually. The average English word is 4.7 letters long. Let’s round up to 5 and add 1 for a space to make 6 characters. Novels are around 90,000 words long. So 800G of pure text represents 800G / 6 char per word / 90,000 words per book =~ 1.5M books. The Library of Congress has 39M books, not to mention all the text produced that’s exclusively online.
- d33 6y agoThat makes me wonder what's the upper bound of storage required to contain all of the text humans had ever written. 1k times as much? 10k? Either way, I have a feeling that all of the drives would probably fit in a regular apartment room. Crazy how we went from first computers spanning across entire floors to fitting all of human thoughts in such tight space.
- Aerroon 6y agoThe entire Library of Congress according to the earlier calculations would fit into 21 TB. There's already a 100 TB SSD that's the size of a 3.5" HDD.[0] 10000 times the Library of Congress would fit onto 2100 of these drives. The dimensions for this 3.5" drive are 26.1 x 147 x 101.8 mm.[1] That's a volume of roughly 391 cm^3. 2100 of them take up 821,100 cm^3. A Fractal Design Define 7 has the dimensions of 547 x 240 x 475 mm. That's a volume of 62,358 cm^3.[2] You could fit 2100 of those SSDs into the same volume as about 13.2 Fractal Design Define 7 cases. That's roughly the size of a bookshelf. You could fit the books of 10000 Libraries of Congress onto one bookshelf. Now imagine using tape for this. Edit: I'm actually not sure whether tape has better density. [0] https://www.techradar.com/news/at-100tb-the-worlds-biggest-ssd-gets-an-eye-watering-price-tag https://www.techradar.com/news/at-100tb-the-worlds-biggest-s... [1] https://www.newegg.com/nimbus-data-dc-100tb/p/2U3-002M-00004 https://www.newegg.com/nimbus-data-dc-100tb/p/2U3-002M-00004 [2] https://www.tomshardware.com/reviews/fractal-design-define-7 https://www.tomshardware.com/reviews/fractal-design-define-7
- stellaathena 6y agoThere's more than 800GB of fanfiction in the world.
- eutectic 6y agoWell, that's truly scary!
- barkingcat 6y agoNot scary, titillating. tens of G's for each desire.
- leogao 6y agoOh, there's a lot more... no promises atm, but we plan on going much bigger for our next dataset.
- jan_Inkepa 6y agoCool project! Minor note: It would be nice if ye said what language the dataset was for on the website :)
- mrconter1 6y agoWould you mind sharing some hints? What new approaches are you taking compared to this one?
- lallysingh 6y agoI wonder when this will be a benchmark number for a desktop GPU/NPU...
- jcims 6y agoI feel like YouTube is going to be a major source of language data in the future. The last statistic I saw was 500 hours of video uploaded every minute. If only 10% of those videos have original speaking in them and those average 40 words a minute, that’s almost 300 GB of transcribed speech per year.
- wongarsu 6y agoYouTube would be a great source for spoken language. But only a tiny portion of YouTube has subtitles, and it doesn't yet feel like automatic transcription is at a level where you would want to use its output to train something else. That day will surely come though
- lovelearning 6y agoIt's not that big even as a text dataset. Common Crawl weighs in at some 250+TB. Even if we assume just 1% of that web data is usable text (it's likely much more), it's still 2.5+TB. https://en.wikipedia.org/wiki/Common_Crawl https://en.wikipedia.org/wiki/Common_Crawl