3 ms·
Even though Common Crawl exists, if you wanted to get the data locally to have a bit of a poke around, it's practically impossible. The next search engine will
by tim-- 4y ago
Even though Common Crawl exists, if you wanted to get the data locally to have a bit of a poke around, it's practically impossible.
The next search engine will never be made by a few kids playing around in the shed.
The latest Common Crawl is roughly 120TB, enough to fit on 48 LTO-6 tapes. LTO-6 is the perfect middle ground expense-wise if you did want to play with the entire dataset at home. The tapes cost ~$30 (AUD) each, a cheap loader from eBay can be had for around ~$1000-4000 (like a Dell PowerVault or Fujitsu Eternus).
Your looking at around $8k just to have your data close to your computer. Either that, or you are going to just run everything on AWS/Google Cloud/Azure. Two of these are your competitors.
No one is going to copy 48 tapes for you, either :)
120TB would take around three months to download - if my ISP didn't cut me off first for probably being the heaviest consumer user of bandwidth.
https://www.chatnoir.eu/ https://www.chatnoir.eu/ has a search engine that is built from the CommonCrawl data running on Elasticsearch, and it runs on 130 nodes [0].
I would love to be able to run my own search engine, but unless there are a number of breakthrough algorithms, I don't see how it can be easily achieved.
I did ask some guys from the Internet Archive if they would be happy to copy some data for me onto some tapes. I wonder if there is a service for this?
[0] https://webis.de/downloads/publications/papers/bevendorff_2018.pdf https://webis.de/downloads/publications/papers/bevendorff_20...
- cosmolev 4y agoWhy dealing with LTO-6 tapes? Isn't it easier to deploy 8 x 16TB HDDs?
- tim-- 4y agoTapes are cheaper when sneaker footing the data around.
- marginalia_nu 4y agoI've sort of mostly dismissed common crawl for the same reason when building my search engine. It's simply too unwieldy. It's far easier (and cheaper) to do my own crawling at a manageable scale, than it is to work with CC's datasets. I'm also not entirely sure which problem CC is intending to solve. It's not like their data is in any way more complete than my own crawl sets, what I can't access to crawl, they don't seem to crawl either.
- rawoke083600 4y agoDont the kids in the shed also have kubernetes ?? /s ok im ready to be judged!