5 ms·
Considering that it's Google, I'm honestly surprised they haven't created a crawler application that can execute the javascript on a given page and re-parse it
by pseudonym 16y ago
Considering that it's Google, I'm honestly surprised they haven't created a crawler application that can execute the javascript on a given page and re-parse it based on the new layout.
You wouldn't think it would be that hard, either-- take Chrome, remove UI, add crawler. Chrome's already got the functionality in "Inspect Element" to show dynamically-created content.
- paradoja 16y agoThat would be a fairly bad idea. Javascript execution, and in particular AJAX calls, frequently change data. I wouldn't want my crawler to delete or edit somehow lots of data if I where Google.
- pseudonym 16y agoNo more or less than following a link with a GET variable. I'm trying to find the article posted on TDWTF regarding a link that went to "&delete=true" that dropped the entire database...but, Ajax calls are no more or less frequently data-changing than any one of a hundred links on survey sites that send their data with GET variables. Edit: Found it! http://thedailywtf.com/Articles/WellIntentioned-Destruction.aspx http://thedailywtf.com/Articles/WellIntentioned-Destruction....
- RyanMcGreal 16y agoAs long as the crawler can't execute POST requests, no one who has built their web application properly will have any problems.
- njharman 16y ago> no one who has built their web application properly will have any problems In other words everyone will have problems (to an accuracy of 3-4 decimals)
- RyanMcGreal 16y agoRelated: http://xkcd.com/327/ http://xkcd.com/327/
- eli 16y agoUh, I thought they did. http://blogs.forbes.com/velocity/2010/06/25/google-isnt-just-reading-your-links-its-now-running-your-code/ http://blogs.forbes.com/velocity/2010/06/25/google-isnt-just...
- patio11 16y agoGoogle does both heuristically parse Javascript and execute Javascript, for at least some fraction of their crawl. This is obviously much more expensive than doing HTML parsing, particularly once the Internet largely goes from static HTML to AJAXy magic.
- pseudonym 16y agoGranted, but if Google can't parse it, I highly doubt anyone else is going to. And as nice as it would be if every web developer read the "Google Guide to Being Nice to Our Web Crawlers", as an internet we still can't get away from IE6. The term "pipe dream" comes to mind.
- IgorPartola 16y agoAnd how would the crawler know when the website is done loading?