5 ms·
Using a robots.txt-file to hide data that shouldn't be public in the first place is a rather bad idea. Because the robots.txt itself is public, it actually high
by Zirro 14y ago
Using a robots.txt-file to hide data that shouldn't be public in the first place is a rather bad idea. Because the robots.txt itself is public, it actually highlights the location of the "private" data.
- benmanns 14y agoAnd, Google had to find a link to this content somehow, which means it's publicly accessible from some Coinbase page.
- manojlds 14y agoThat is what I am wondering. Here I am building a web site and wondering how to make Google index sites behind authentication and such, and here is a site that has got everything indexed.
- benmanns 14y agoAfter reading Coinbase's response, it seems that the pages were not linked to by Coinbase, but by users posting links to their checkout page on the Internet.
- ntumlin 14y agoWhat's a good alternative to robots.txt?
- dpcx 14y agoBlocking access to googlebot for those pages is the easiest. But robots.txt would work, if you just did Ignore /checkouts/
- espo 14y agoProper access control on your web site.
- deleted 14y ago[deleted]
- Cthulhu_ 14y agoSimply not showing the transaction if you're not logged in and not the user belonging to the transaction?
- enoptix 14y agoYou can also specify a meta robots tag inside the page HTML. If you want to block a lot of pages, your best bet would be to add it to your master layout or template. You get the same effect of blocking on robots.txt but without exposing a list of blocked pages. The downside is that Google will still crawl the page and use your bandwidth, but the page won't be indexed.
- JimWestergren 14y agoI would suggest both these two: Check the user agent for bots and if it is a bot send a 404 header and exit before page needs to load. Also add a meta noindex just in case. Robots.txt DOES NOT prevent indexing, just crawling.