8 ms·
I've been in disagreements with SEO people quite frequently about a "Noindex" directive for robots.txt. There seem to be a bunch of articles that are sent to me
by jxcl 7y ago
I've been in disagreements with SEO people quite frequently about a "Noindex" directive for robots.txt. There seem to be a bunch of articles that are sent to me every time I question its existence[0][1]. Google's own documentation says that noindex should be in the meta HTML but the SEO people seem to trust these shady sites more.
I haven't read through all of the code but it assuming this is actually what's running on Google's scrapers this section [2] seems to be pretty conclusive evidence to me that this Noindex thing is bullshit.
[0] https://www.deepcrawl.com/blog/best-practice/robots-txt-noindex-the-best-kept-secret-in-seo/ https://www.deepcrawl.com/blog/best-practice/robots-txt-noin...
[1]https://www.stonetemple.com/does-google-respect-robots-txt-noindex-and-should-you-use-it/ https://www.stonetemple.com/does-google-respect-robots-txt-n...
[2] https://github.com/google/robotstxt/blob/59f3643d3a3ac88f613326dd4dfc8c9b9a545e45/robots.cc#L262-L276 https://github.com/google/robotstxt/blob/59f3643d3a3ac88f613...
- jxcl 7y agoGoogle is also really generous with how they will let you spell "disallow": https://github.com/google/robotstxt/blob/59f3643d3a3ac88f613326dd4dfc8c9b9a545e45/robots.cc#L691-L699 https://github.com/google/robotstxt/blob/59f3643d3a3ac88f613... :D
- knorker 7y agoI'm not surprised. Some people think humans read robots.txt and get super angry when the crawler doesn't understand.
- c3534l 7y agoI read robots.txt, but I'm not a massive corporation.
- deleted 7y ago[deleted]
- kortilla 7y ago> (absl::StartsWithIgnoreCase(key, "disallaw"))))); Ah, the southern version. :)
- Fnoord 7y agoMakes sense because not everyone speaks English as native language. Disalow is pretty close to disallow, phonetically.
- ryandrake 7y agoYuck though! Imagine if you were writing a compiler. Would you make it accept “unsinged” “unnsigned” “unssined” and “unsined” as keywords, just to catch spelling mistakes? Not sure I like that pattern.
- bscphil 7y agoIt's a little different in that case, since the person using the parser is also the person writing the input to the parser. So if the input fails the parser, the author of the code can simply correct it. As I understand it, there's no single standard that captures how all robots.txt files are formatted, so there's no "standard parser" that the authors of these files could be expected to pass.
- anoncake 7y agoThat is not an excuse. Non-native speakers can learn to spell.
- avip 7y agoThis is great. #kAllowFrequentTypos
- samstave 7y agoIve never made a tpyo in my life!!!
- quickthrower2 7y agoWhats a tpyo? My typo array for "typo" only has tipo, typpo and thai-pho
- bigiain 7y agoIf you're gonna include thai-pho you also need to include thai-fur...
- seanlinmt 7y agoWhile that's great, there should be instances where crawlers should ignore noindex directives. For example, all .gov sites.
- paulie_a 7y agoThis might be controversial but everything is fair game everywhere. If you can crawl it, tough luck. It's there and everyone can get to it anyways, why not a crawler?
- chc 7y agoBecause the rules a well-functioning society runs by are more nuanced than "Is it technically possible to do this?" If you'd like a specific example of why people might seek this courtesy, someone might have a page or group of pages on their site that works fine when used by the humans who would normally use it, but which would keel over if bots started crawling it, because bot usage patterns don't look like normal human patterns.
- paulie_a 7y agoHonestly that is the site owners problem. If it can be found by a person it's fair. I genuinely respect the concept of courtesy but I don't expect it. People can seek courtesy but they should have expectations of whether or not it will happen.
- Spivak 7y agoSo in your view is DoS attack not actually an attack and site owners should just have to handle the traffic?
- paulie_a 7y agoTechies forget the rule of laws. A dos has intent. A bot crawling a poorly designed website accidentally causing the site owners problems does not have malicious intent. They can choose to block the offender just like a restaurant can refuse service. But intent still matters.
- bhartzer 7y agoGoogle has been very clear lately (via John Mueller) regarding getting pages indexed or removed from the index. If you want to make sure a URL is not in their index then you have to 'allow' them to crawl the page in robots.txt and use a noindex meta tag on the page to stop indexing. Simply disallowing the page from being crawled in robots.txt will not keep it out of the index. In fact, I've seen plenty of pages still rank well despite the page being disallowed in robots.txt. A great example of this is the keyword "backpack" in Google. You'll see the site doesn't want it indexed (it's disallowed in robots.txt) but the site still ranks well for a popular keyword).
- LunaSea 7y agoDoesn't that indicate that Google doesn't respect robots.txt then?
- C1sc0cat 7y agoRobots stops a page from being crawled the noindex tag stops it getting into the index. Google is also slow to honour 404 and drop pages which can hang around for ages, Bing is much faster to remove 404 pages.
- rafaelm 7y agoUsually, creating "410 Gone"[0] response for the URL and running the URL through the URL Inspection Tool [1] can help make things a bit faster. But yeah, it does take a while to get these 404s removed. [0] https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/410 https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/410 [1] https://support.google.com/webmasters/answer/6065812?hl=en https://support.google.com/webmasters/answer/6065812?hl=en
- inlined 7y agoThat distinction exists in many systems. E.g. for cloud events, 404 is considered with skepticism because it could be a race condition in provisioning or transient issue whereas 410 requires data streams to be cut off.
- Youden 7y ago> but the SEO people seem to trust these shady sites more. It makes more sense when you realize that the SEO people (with a few exceptions) are usually pretty shady as well. You rarely hear them recommending that you write better content to get better results, it's always nonsense like "put nofollow on everything so your score doesn't leak".
- mnbvkhgvmj 7y agoThis really has not been my experience with SEO people in the 2010s. They have focused on page load, no errors, good redirect schemes etc
- Spivak 7y agoI mean if you're hiring a SEO person isn't this literally what you're paying for -- tricks to increase your search ranking without changing your content?
- SquareWheel 7y agoNot at all. SEOs are more likely to find all the ways you're currently hanging yourself. Some common examples I see are: - Putting important text inside of images - Duplicate content out the wazoo - Not making use of canonicals - No sitemaps, html or xml - Page performance issues - Broken mobile support And of course, poor content. You can't rank if you don't have content.
- actsof 7y agoI'd argue most of those are necessary for good content (if we don't view content separately from presentation) > Putting important text inside of images I'm sure the reason for this is that it's hard to parse text from images, and while Google could use their AI to figure it out, they don't bother. But it also prevents blind people from being able to read the text, so it does worsen the experience. > Duplicate content This makes the site harder to navigate for users as well. > Page performace issues Quite obviously makes the experience worse. > Broken mobile support. -..-
- 7y ago
- adrianmonk 7y ago> evidence to me that this Noindex thing is bullshit For those who (like me) don't know a lot about this, which side of the argument is bullshit? Have you just been proved right or wrong?
- klohto 7y agoI recommend finishing the whole comment
- celeritascelery 7y agoread the whole comment. Still confused as well.
- jxcl 7y agoIt looks like it's too late for me to edit my comment, but I've been proved right. Putting a Noindex directive directly in robots.txt is frequently suggested, but this seems like definitive proof that that does nothing (at least with Google). As far as I can tell the inception of this idea was that it was briefly mentioned by some Google employee in an interview. Maybe it was supported in the past or maybe he just misspoke, but I bet even now we'll see people still using this tag.
- inh 7y agoExcept... they were correct. Google has now clarified that they're removing the code behind the undocumented items, with noindex called out explicitly. https://webmasters.googleblog.com/2019/07/a-note-on-unsupported-rules-in-robotstxt.html https://webmasters.googleblog.com/2019/07/a-note-on-unsuppor... It wasn't officially supported / the recommended way - but it worked (in many cases.)