4 ms·
Hello Facebook Crawler
- maxjaderberg 14y agoBy looking at the headers you now have a great way of writing some analytics tools to see how much your website is shared on Facebook...
- jabo 14y agoI would imagine that they cache the page contents and hence hit a URL only once in a certain period of time, thus skewing any analytics built around this.
- dannyr 14y agoYeah it's cached by Facebook. That's why if you want to change your meta and/or open graph tags info, you need to feed your page to Facebook's Url Linter (https://developers.facebook.com/tools/lint/ https://developers.facebook.com/tools/lint/).
- mpeg 14y agoyou can also programmatically force it to refresh the cache via a POST to https://graph.facebook.com/?id=http://google.com&scrape=true https://graph.facebook.com/?id=http://google.com&scrape=...
- deleted 14y ago[deleted]
- synotic 14y agoIf you register your site with Facebook, you can get that information on Facebook's end: https://developers.facebook.com/docs/opengraphprotocol/ https://developers.facebook.com/docs/opengraphprotocol/ under "Domain Insights."
- cddotdotslash 14y agoWhy is this even news? Facebook has been crawling links for ages every time you post on the site. The crawler is how the link you paste gets a title, description, and sometimes a thumbnail.
- TannerLD 14y ago"What to Submit On-Topic: Anything that good hackers would find interesting. That includes more than hacking and startups. If you had to reduce it to a sentence, the answer might be: anything that gratifies one's intellectual curiosity."
- lostlogin 14y agoWow I'm sick of people posting claims of off topic. And wow it's funny that someone has the patience to reply with the posting rules.
- mars 14y ago+1. not sure how a post like this can make it to the front page.
- whalesalad 14y agoThere was a period where Hacker News consisted primarily of people on the right-hand side of the spectrum. People who were working inside of startups or had lots of experience with the web and our industry. Pretty much everyone knew what sharding was, and MongoDB wan't very popular. These days we've got a lot more people and they show up all across the board. Clearly if this is on the homepage, it was voted there by your peers. This kind of knowledge is completely obvious to many of us, but not everyone is on your level. Cut 'em some slack.
- omarchowdhury 14y agoEven so, for those who are up to that point, that headline could give the implication that Facebook is getting into the search business.
- whalesalad 14y agoThis reminds me of a recent experience I had with the Bing bot. This most recent YC round, my co-founder and I used Skydrive to edit our application. Skydrive integrates pretty nicely with Word, even on a Mac, to allow for collaborative editing. It's like the best parts of Sharepoint, minus all the crap, and inside of a modern UI. I'm a diehard Apple user, but I also subscribe to the "right tool for the job" principle ... in this case it worked pretty well. Anyway, inside the document were links to some private areas of our website that contained demo materials for YC. As requested, they were not password protected, but also not linked from anywhere else. While submitting I ensured that our nginx logs would capture visits to these URL's in a separate log, so we'd know when it was being looked at (sidenote, seeing visitors coming from inside justin.tv + the rincon hill towers is kind of exhilarating). What surprised me was that almost immediately after we began working on the document, the Bing bot was going apeshit exploring the domain and the 'private' URL's. I had to quickly add a robots.txt to deny all on the root. I thought it was pretty interesting. At first I felt almost violated. But then it seems logical that they'd be indexing every URL in every document stored in their datacenter, why not?
- rpm4321 14y agoEh, I'm pretty sure you should still feel violated. The fact that they are parsing your private documents for information that they can use to help another business unit is really sketchy. It would make me wonder what else they are scanning my data for. Personally, I'll never use an MS cloud service because of this anecdote - not that it was that likely to begin with.
- eli 14y agoYou assume they were indexing skydrive documents. It could well be that one of the people who visited the link had a Bing toolbar installed. Either way, all publicly accesible documents will get indexed sooner or later.
- rpm4321 14y agoYou may be right, but I can't help but smirk at the thought of PG or Buchheit downloading and installing the Bing Toolbar ;)
- edouard1234567 14y agoI'm surprised this post makes it to the homepage... They've been doing that for ever, no need to look at your logs to figure this out. How else would they find and display an image form the page you're providing a link to.
- eli 14y agoI would imagine they're checking the URL for malware as well.
- nwh 14y agoProbably, I've seen then ban whole domains (droplr.com) previously for distributing malware.
- slajax 14y agoI wish I had enough karma to down vote this.
- spyder 14y agoAlso it would be smart to run malware check on these urls if they don't already doing it.
- justinph 14y ago12 lines of code instead of: tail -f /var/log/apache2/access.log
- mindctrl 14y agoThe input box is a keylogger that sends what you type to FB before you ever press enter. That's the interesting story here. Everyone already knew FB had a scraper for retrieving images and meta data. What they probably didn't know is every keystroke inside that box is logged and sent, no Enter required.