4 ms·
Where was this 10 years ago when I was reverse engineering the Google robots.txt parser by feeding example robots.txt files and URLs into the Google webmaster t
by randomstring 7y ago
Where was this 10 years ago when I was reverse engineering the Google robots.txt parser by feeding example robots.txt files and URLs into the Google webmaster tool? I actually went so far as to build a convoluted honeypot website and robots.txt to see what the Google crawler would do in the wild.
Having written the robots.txt parser at Blekko, I can tell you what standards there are incomplete and inconsistent.
Robots.txt files are usually written by hand using random text editors ("/n" vs "/r/n" vs a mix of both!) by people who have no idea what a programming language grammar is. Let alone follow BNF from the RFC. There are situations where adding a newline completely negates all your rules. Specifically, newlines between useragent lines nor between useragent lines and rules.
My first inclination was to build an RFC compliant parser and point to the standard if anyone complained. However, if you start looking at a cross section of robots.txt files, you see that very few are well formed.
With the addition of sitemaps, crawl-delay, and other non-standard syntax adopted by Google, Bing, and Yahoo (RIP). Clearly the RFC is just a starting point and what ends up on website can be broken and hard to interpret the author's meaning. For example, the Google parser allows for five possible spellings of DISALLOW, including DISALLAW.
If you read a few webmaster boards, you see that many website owners don't want a lesson in Backus–Naur form and are quick to get the torches and pitchforks if they feel some crawler is wasting their precious CPU cycles or cluttering up their log files. Having a robots.txt parser that "does what the webmaster intends" is critical. Sometimes, I couldn't figure out what some particular webmaster intended, let alone write a program that could. The only solution was to draft off of Google's de facto standard.
(To the webmaster with the broken robots.txt and links on every product page with a CGI arg with "&action=DELETE" in it, we're so sorry! but... why???)
Here's the Perl for the Blekko robots.txt parser.
https://github.com/randomstring/ParseRobotsTXT https://github.com/randomstring/ParseRobotsTXT
- asdfman123 7y agoIt's an easy fix if Google cared. Have an online tool that validates if the robots.txt is correct, and send out an announcement that files that don't meet spec will be penalized in terms of SEO.
- DEADBEEFC0FFEE 7y agoCould also make the spec easier to be compliant.
- pbhjpbhj 7y agoIIRC Google do automated checks on robots.txt and report in webmaster tools if you did something that looks crazy. https://qph.fs.quoracdn.net/main-qimg-002d1f819e1bfbbd14fa2de937b489d0 https://qph.fs.quoracdn.net/main-qimg-002d1f819e1bfbbd14fa2d... shows some of the interface, but I seem to recall getting notified in webmaster tools when I messed up the robots.txt on a particular site.
- hombre_fatal 7y agoThat just punishes users by sinking relevant results for reasons users couldn’t possibly care about.
- PunksATawnyFill 7y agoI enjoy the hypocrisy of Google punishing sites that aren't "mobile-friendly," and then deliberately disabling ZOOMING on their own mobile sites.
- saalweachter 7y agoAccidentally deleting someone's entire website because they don't understand the difference between GET and POST requests is virtually a right of passage when writing a web crawler.
- chii 7y agoif their 'delete website' link is reachable anonymously and performs using a GET, they deserve the result.
- sjwright 7y agoDoes Google promise that its bot will never submit a POST even if it’s triggered by an <a onclick>?
- epse 7y agoGooglebot has no javascript
- adrianolek 7y agoIs that true, though? [1] [1] https://webmasters.googleblog.com/2019/05/the-new-evergreen-googlebot.html https://webmasters.googleblog.com/2019/05/the-new-evergreen-...
- groestl 7y agoIt's a masterpiece that Google convinced everyone to use their crawler engine as a browser.
- aidos 7y agoI’m trying to find a link to it but there was an incident based on this issue somewhere around 1999-2001 where Microsoft added a sort of prefetching thing to IE (or was it Netscape?!) and it would effectively click all the links on the page in order to get all the content in the cache. Lots of us really didn’t know what we were doing and we’d made all the action buttons in the listing screens regular links. As you can imagine, pandemonium ensued. Hey, at least we’d figured out that sql injection was a thing. It was a simpler time.
- vinay_ys 7y agohmm, you mean "\n" vs "\r\n"? ;-)