4 ms·
I've heard about CommonCrawl [0] a few times here and it seems interesting to think of building a search engine around it. Or, maybe, a boilerplate search engin
by dceddia 4y ago
I've heard about CommonCrawl [0] a few times here and it seems interesting to think of building a search engine around it. Or, maybe, a boilerplate search engine that people could build upon. Maybe someone will even reply with an example :) I see they have a big list of projects [1] and honestly this might already be in there. It's fun to imagine a whole bunch of specialized search engines, like one that's great at finding product reviews, one for software development, one for recipes, and so on.
I think the powerful aspect of a distributed/mirrored index is that the data is public, so anyone can build a UI or a product around it to offer their unique spin.
From a more cynical angle, I also think it's likely that people still tend to gather around a few "popular" ones of anything like this. The web3 stuff seemed to do this. It happened with torrent trackers too. Eventually there's an option that is "base + X" where X is enticing enough to attract a crowd, and then once that crowd is big enough, that popular thing can pivot into a centralized thing, and the cycle repeats...
0: https://commoncrawl.org/ https://commoncrawl.org/
1: https://commoncrawl.org/the-data/examples/ https://commoncrawl.org/the-data/examples/
- mijustin 4y agoThe difference with torrent trackers is their index is public and can be mirrored. Torrent sites like 1337x were born out of previously popular indexes (Kickass Torrents, h33t). [0] It's not dissimilar to when a GitHub fork becomes more popular than the original project. [1] To me, the biggest disappointment with web3 is they never cracked "distributed discoverability." It still relied on the old model, of proprietary indexes/marketplaces. I'm way more inspired by the P2P efforts, especially Gnutella's initial attempts at distributed search. [2] Feels like this tech could still be iterated on and improved. 0: https://en.wikipedia.org/wiki/1337x https://en.wikipedia.org/wiki/1337x 1: http://gitpop2.herokuapp.com/tobi/delayed_job http://gitpop2.herokuapp.com/tobi/delayed_job 2: http://rfc-gnutella.sourceforge.net/developer/share/intro.html http://rfc-gnutella.sourceforge.net/developer/share/intro.ht...