6 ms·
Does your think tank come up with any ideas on how to keep companies like Pinterest from peeing in the pool so to speak? They made it impossible to find DIY cr
by marmaduke 6y ago
Does your think tank come up with any ideas on how to keep companies like Pinterest from peeing in the pool so to speak? They made it impossible to find DIY craft stuff that I used to search for, and I can’t imagine it’s the only site doing it.
- knuckleheads 6y agoWe are pursuing ideas that we think will increase competition in the search engine market. Our main focus right now is making Google’s index of the web available for use by competitors, since no one other than Google is really allowed to crawl the web. If people were allowed to use the index to build what they want, then you or somebody else would be able to build a search engine that avoids the Pinterest cruft and would be more suited to your needs, instead of Google’s one size fits all approach. So, to answer your question directly, yes we do have some ideas about how to do this.
- burnthrow 6y ago> making Google’s index of the web available for use by competitors What incentive would Google have to continue populating that index? Would I be breaking the law if I independently crawled and hosted an index without publishing an API for it? > since no one other than Google is really allowed to crawl the web Maybe this is the problem that needs solving.
- knuckleheads 6y ago> What incentive would Google have to continue populating that index? Presumably they would still want to run google.com and make money off of it. > Would I be breaking the law if I independently crawled and hosted an index without publishing an API for it? No. You would not get the advantage that Google gets when it crawls the web and so would not have access to a large amount of data that nobody else has access to. Updated based on edit of parent post: > Maybe this is the problem that needs solving. Why have websites waste the money to serve all those requests all over again? Why don't we have Google share the results and we can use that money to do more productive things than recreating that work? I don't think website operators would be happy if there were a hundred more crawlers out there crawling as much as Google does now.
- deleted 6y ago[deleted]
- lopmotr 6y agoDo any site operators actually block non-Google search engine crawlers because being listed DDG/Bing/etc isn't worth the extra cost of serving the crawler? It sound a bit ridiculous unless they actually don't want to be found. Maybe they only allow GoogleBot because that's all they thought of and the extra cost is in researching what all the other search engines call theirs. Perhaps other search engines should spoof GoogleBot. Browsers have being doing that since forever spoofing Netscape (Mozilla), Safari, etc. for the same reason. > Why don't we have Google share the results and we can use that money to do more productive things than recreating that work? This sounds like a common fallacy of people criticizing the free market. Duplicated effort looks wasteful but turns out to be far more productive than the lack of incentive that comes with not being able to profit from your work/investment.
- Someone 6y agoI would think the site owner’s cost of being indexed is the same for every search engine that indexes the site. The benefit varies with the quality of the search engines, and that will vary between search engines, but it does get larger the more a search engine is used, so a cost/benefits analysis may show Google and a few other large ones are the only ones worth supporting.
- knuckleheads 6y agoYes! Exactly!
- loeg 6y agoYes, site operators actually block non-Googlebot crawlers. See the example [0] from https://news.ycombinator.com/item?id=25538842 https://news.ycombinator.com/item?id=25538842 . Spoofing crawler identity completely defeats the point of the honor-system robots.txt.
- knuckleheads 6y ago
- deleted 6y ago[deleted]
- djrobstep 6y ago> What incentive would Google have to continue populating that index? Not going to jail, presumably.
- aleph_naught 6y ago> since no one other than Google is really allowed to crawl the web. ??
- knuckleheads 6y agoThere are two main reasons why I say nobody besides Google is really allowed to crawl the web. The first is that Google gets much more access to pages on websites than everybody else. You can see this by examining the robots.txt files of various websites[0]. I've been doing this for several years now and Google has a consistent advantage across many thousands websites that I've looked at. This adds up to a significant advatnage and many search engine operators complain about how it hampers their ability to compete with Google[1]. The second is that Google gets to ignore crawl delay directive in robots.txt while other search engines don't[2]. Website operators cannot tell Google how fast they want their website crawled, they can only request that Google slow down. If another search engine tried to do what Google does, they would likely be blocked by many important websites. If you would like to read more about this, please checkout https://knuckleheads.club/ https://knuckleheads.club/ [0] https://pdf.sciencedirectassets.com/robots.txt https://pdf.sciencedirectassets.com/robots.txt [1] https://www.nytimes.com/2020/12/14/technology/how-google-dominates.html https://www.nytimes.com/2020/12/14/technology/how-google-dom... [2] https://www.seroundtable.com/google-noindex-in-robots-txt-dead-27824.html https://www.seroundtable.com/google-noindex-in-robots-txt-de...
- grishka 6y agoSo, uh, don't respect robots.txt in your search engine? It's not like there's a law that you have to, and that you can't pretend you are Googlebot. The only real obstacle I can imagine is that some firewalls might be configured to be more permissive with traffic originating from Google subnets.
- knuckleheads 6y agoYou would be blocked fairly quickly by many website operators and no longer able to access those websites if you straight up ignored robots.txt files. You also might even end up being served cease and desists by some websites and sued if you continue to persist and try to find ways around it.
- marmaduke 6y ago> making Google’s index of the web available for use by competitors That does seem like a good idea since it amounts essentially to a database of stuff that Google does not own (by construction) and is of public utility.
- knuckleheads 6y agoExactly. Everybody always talks about turning tech companies into public utilities without really explaining what the utility would be. A public index of the web would be an amazing utility for many of the reasons in this thread and would spur on a ton of new innovation and businesses. If you would like to read more about all this, please checkout https://knuckleheads.club https://knuckleheads.club
- rarefied_tomato 6y agoThe pursuit of a public index is excellent. Some feedback: - A cost of membership runs contrary to establishing this group, especially at such a high recurring charge. - I'm not sure what your software/AWS situation looks like, but 20 million robots.txt files acquired from Common Crawl is something I can analyze on my PC. It doesn't seem to presently justify such high costs. - Prioritize building a mockup index with an intuitive frontend. This is essential for non-technical people to understand - Exclusively talk with EU legislators (they are motivated, whereas nothing will happen in the US).
- knuckleheads 6y agoThank you for the feedback! I think the price for membership dues is reasonable and many people agree evidenced by them signing up. I think I might start a petition that is free to sign up on though, thank you for the inspiration! It is possible to analyze those files on the pc, it just takes a much longer time. The analysis is an iterative process and so the faster the computers the faster the iterations and process go. I was analyzing them on my pc with python for the first year until it got too slow and my I am using an aws server with some rust and that is going much better. I also need to increase the number of files analyzed by about two orders of magnitude soon as well. Great idea, very cool. That’s going on the todo list! And I am going to be reaching out to and speaking with whoever is interested. One of the fun things about this is that it is an international dynamic, with some jurisdictions having abilities that others don’t. For example, the UK CMA has subpoena powers that the US Congress lacks and got a ton of information out of Google and Bing that shocked me. The US has the ability to get the CEO’s to show up to hearings while the UK does not in the same way. Why limit ourselves to one government when there are so many to mix and match from here?
- frakkingcylons 6y ago> Our main focus right now is making Google’s index of the web available for use by competitors How would that work? I'm pretty curious about this if there's anything out there to read. EDIT: Ah I saw the link to Knuckleheads' Club :P
- knuckleheads 6y ago:) No worries. If you'd like to chat about it more, shoot me an email at zack@knuckleheads.club, I'm obviously very passionate about the subject and love talking about it. There's still a lot more to write and publish, I won't be done with this hobby horse of mine for a while yet.
- fsflover 6y agoWow, such index would tremendously help free p2p search YaCy, https://yacy.net https://yacy.net.