5 ms·
>If robots instructions don't mention Applebot but do mention Googlebot, the Apple robot will follow Googlebot instructions. So if I set in my robots.txt to di
by frankacter 11y ago
>If robots instructions don't mention Applebot but do mention Googlebot, the Apple robot will follow Googlebot instructions.
So if I set in my robots.txt to disallow all bots except Googlebot, Applebot will index anyway? I don't think I like that precedent.
- Bulk70 11y agoSerious question, because I can't imagine your use case - under what circumstances would you wan't to block all bots except one?
- RexRollman 11y agoMaybe for ethical reasons or some kind of exclusive agreement with one particular search vendor?
- mbel 11y agoI'm not an expert but my guess is: limiting bot traffic, but keeping the site available for the most popular search engine.
- mrweasel 11y agoMy experience is that the worst bots don't respect robots.txt anyway. Getting crawled by the major search engines typically isn't that bad, they tend to know what they're doing. Getting hammered by some crappy local search engine is what's annoying. We don't limit any bots, except once where we completely blocked Eniro in our firewall. Google, Bing and a ton of other could index at the same time, with no issue. Eniro for some reason decided to just index way to much at once, no reaction to robots.txt and no reply from the email they so kindly included in the headers. But I see your point, it's just a bit sad when Google has become "The Internet".
- karmakaze 11y agoI thought FB was the internet. Googlebot is just the Kleenex of indexers.
- deleted 11y ago[deleted]
- deleted 11y ago[deleted]
- gojomo 11y agoOwner of GOOG stock, maybe?
- bad_user 11y agoIs it of any importance?
- nemothekid 11y agoFacebook blocks all bots except a select few to prevent site scraping - https://www.facebook.com/robots.txt https://www.facebook.com/robots.txt
- gutnor 11y agoThat's maybe jumping to conclusion. I interpreted the sentence less literally and more like "in absence of rule, default to GoogleBot ones" So if you put a wildcard rule forbidding access and a specific one allowing access to GoogleBot, AppleBot will honour the wildcard one. That's how I would have coded it anyway: parse the rule for current agent string, if no rule applies, run it again with GoogleBot one before assuming that website does not contain restrictions.
- koyote 11y ago"If robots instructions don't mention Applebot" I would assume this means that it will follow GoogleBot unless you specifically mention AppleBot by name and not by using a wildcard. So a User-agent: * would be ignored if a User-agent: GoogleBot is found.
- Aissen 11y agoI've said it before, and I'll say it again: blocking all but certain bots is the best way to block innovation. It's bad for you, it's bad for the ecosystem. Bad for the ecosystem because incumbents that want to respect robots.txt to propose new services will have an unfair disadvantage. It's bad for you because it'll just give more power to Google et al regarding your incoming traffic (and you'll have to follow their rules, like every SEO is doing right now).
- WorldWideWayne 11y agoI wonder what are the legalities of this type of discrimination? Retail businesses aren't allowed to arbitrarily refuse service, they have to follow certain rules and be consistent.
- MCRed 11y agoIs there legal backing to enforce obeying robots.txt, or is it a guideline?
- WorldWideWayne 11y agoThere are some court cases that involve robots.txt but I'm not sure because there was a lot of reading to do and I was hoping someone who knew would just come along and summarize it for us :)
- d0ugie 11y agoNote that serving bots, especially for media-rich sites, eats heavily into a finite resource. For those of us whose resources are especially tight, blocking Yandex, Baidu and MSN may be very helpful, your ideals notwithstanding.
- Aissen 11y agoPutting anything on the internet exposes your "finite resources" to attackers. I'm sorry, but robots.txt is just a courtesy offered to you, and if you're doing serious business, you shouldn't rely on it. Use authentication, rate-limiting, captchas, DDoS-protecting CDNs, and enforce the limitations at the source, don't rely on nice people respecting your robots.txt.
- Tloewald 11y agoYou'd prefer they didn't honor robots.txt at all?
- mobiplayer 11y agoThat's what you get when many people code with only one company in mind. It happened in the past and it is happening again, whether we like it or not. Even Microsoft was kinda masquerading early versions of Edge as Chrome.
- protomyth 11y ago"Even Microsoft was kinda masquerading early versions of Edge as Chrome" I would assume that was more to keep the tech press from seeing it.
- mobiplayer 11y agoI don't know about the not public betas, but as soon as there was a public Windows 10 TP and you could swap IE's engine to Edge I'm pretty sure they had "Edge" at the end of the user-agent, but they had the rest of the Chrome UA so most of the websites out there considered it as being Chrome. I may also be completely wrong :)
- jwr 11y agoI actually loved this part of the page being discussed. The whole idea of versioning content based on who accesses it is broken and fundamentally at odds with the idea of the open web. Same goes for user-agent string madness, by the way. Yes, we should be able to tell robots from humans, but otherwise, it's supposed to be the Web. Incidentally, this hits close to a pain point: I find it extremely annoying when publishers (like Elsevier) hide content behind a paywall, but still expose it to Googlebot for indexing. The result is that you are able to find a scientific article, which is not accessible (but Googlebot has cached snippets). This goes against Google's own guidelines (they used to tell people that Googlebot must not see different content from browsers). And it goes against the whole idea of the Web: if you want to hide stuff behind a paywall, do so — but then it is no longer accessible. Going back to Applebot, I love the fact that they will now follow Googlebot instructions. Hopefully people will stop distinguishing who accesses content.
- soylentcola 11y agoI remember a big issue with user-agent filtering when Google was still trying to make Google TV work. At the time, loads of TV networks had free, full streaming episodes of shows on their websites as a way to capture some ad revenue that would be lost if people were torrenting or whatever in order to see shows they missed. Google TV was attempting to list those episodes alongside whatever was currently live on TV via your cable/sat/antenna and stuff from other sites like Youtube. The idea was that instead of having to go to all sorts of places to find content, it would show you what was available at a given time based on what you were looking for. ...and then all of the network sites and the free Hulu stuff got put behind a user agent filter and essentially wiped out a huge part of GTV's reason for existence. The goal was to bring all of the free content into one place but the networks didn't want you watching a free stream in lieu of a cable broadcast. They wanted you to watch cable on your living room TV and only use the free streaming episodes from the computer in your office as a backup. Same goes for services where web viewing/listening is free but if you try to access it from a mobile web browser, you have to either fool the site or subscribe to some mobile version.
- karmakaze 11y agoWhy? Applebot is attempting to adhere to the spirit of robots.txt rather than the letter. If the site owner cares otherwise, don't just explicitly name one bot. Is `user-agent: *bot` valid?