14 ms·
Deprecating our AJAX crawling scheme
- a2tech 11y agoThis is good-one of my current projects for a customer is entirely AJAX/JS rendered and we were worried that Googlebot would have a fit with it.
- iwilliams 11y agoWe recently built a site for a customer in Ember and their SEO guys were concerned about indexing. I wasn't sure how it was going to work out, but in the end Google has been able to index every page no problem.
- hanniabu 11y agoSorry if this is a stupid question as this is outside my field of work, but how can you tell if your page has been successfully indexed or not?
- a2tech 11y agoDo you know if they sent Google a sitemap? Our client is insisting on a sitemap that has pointers to every-single-product. Something on the order of 2MM+ product pages. It seems like a bit much to me
- MichaelApproved 11y agoKeep this in mind https://support.google.com/webmasters/answer/183668?hl=en&topic=8476&ctx=topic https://support.google.com/webmasters/answer/183668?hl=en&to... > Break up large sitemaps into a smaller sitemaps to prevent your server from being overloaded if Google requests your sitemap frequently. A sitemap file can't contain more than 50,000 URLs and must be no larger than 50 MB uncompressed.
- iwilliams 11y agoThe site is built on Tumblr which automatically generates a sitemap for the individual posts, but not any other pages on the site. For example the "about" page is not in the sitemap, but is still indexed.
- gildas 11y agoHow many pages does the site have?
- greglindahl 11y agoYou should still be worried. Just because googlebot expensively evaluates JS for some websites doesn't mean it will evaluate JS for your brand-new website. You might get crawled a lot less deeply than if you had good content in your static pages.
- chc 11y agoBy abandoning their AJAX crawling scheme as described in the OP, they are essentially saying that they will evaluate JS for all sites. Do you have some reason to doubt that?
- greglindahl 11y agoIf crawling with JS costs 1,000X or 10,000X as much as crawling without, it's fair to say that even Google isn't going to crawl 100s of billions of pages executing JS. As a former web-scale search engine CTO, my opinions are commonly surprising to folks who haven't built a web-scale crawler/search engine.
- tracker1 11y agoMy own experiments/experience shows that recrawls happen about 1/3 as often and tend to lag a few days behind for JS content vs inlined/delivered content. It's helped a little by dynamically delivering the sitemap data, but even that only speeds things up a little. My guess is they're putting about 1/10th the effort into keeping things freshly indexed for JS, but may well be devoting 2x the resources vs directly received content.
- copsarebastards 11y ago> By abandoning their AJAX crawling scheme as described in the OP, they are essentially saying that they will evaluate JS for all sites. No, they are not. If you even think that's possible you're fundamentally misunderstanding how search engines work.
- devNoise 11y agoAbout a year ago I wrote a post[1] about how I couldn't get google to index my AngularJS app. My main problem was the interaction between googlebot and the S3 server. I'll have to go back and test if the crawler's behavior will render the correct content. 1 - https://medium.com/@devNoise/seo-fail-figuring-out-why-i-cant-get-my-content-into-google-6ca6c7a64b51 https://medium.com/@devNoise/seo-fail-figuring-out-why-i-can...
- tracker1 11y agoDo you have a sitemap.xml for common routes.. also is your angular app actually doing routing (hash based or push state)?
- devNoise 11y agoI have the .html5Mode set to true so it was routing based on push state. I hadn't gotten around to creating a process to generate the sitemap.xml before I gave up on the site. For SEO, we were more concerned with getting the time sensitive content indexed.
- rdoherty 11y agoWow, I built a project that rendered JS built webpages for search engines via NodeJS and PhantomJS. Rendering webpages is extremely CPU intensive, I'm amazed at the amount of processing power Google must have to do this at Internet scale. I really hope this works, lots of JS libraries expect things like viewport and window size information, I wonder how Google is achieving that.
- Jake232 11y agoCan confirm. Launched a project recently with over 500 concurrent PhantomJS workers. Let's just say my hosting bill is significantly more expensive than it was.
- Figs 11y ago> lots of JS libraries expect things like viewport and window size information, I wonder how Google is achieving that. Just plug in common screen parameters (e.g. 1920x1080, 1366x768, ...) and analyze it as if it were the result you'd get by default with Chrome on such a screen, I would imagine.
- tracker1 11y agoSame goes for user agent spoofing (to some extent). You can imagine most of the stuff when you use the chrome dev tools being done without actual user interaction.
- MichaelApproved 11y agoI'm wonder if they're cutting out a lot of the rendering that PhantomJS is doing. Not to say that any type of rendering is cheap but I'm guessing they have a limited version of a JS rendering engine that does just enough to index the page. I bet they'd also skip on all the FB like buttons and other common social media elements that don't impact the content.
- Paul-ish 11y agoThis makes sense. Would it be sufficient to just see how the content (eg new <ul> elements or something along those lines) on the page changes when JS is executed, without actually rendering anything?
- cotillion 11y agoSo they're actually evaluating all js and css Googlebot is consuming. That's insane. Can we forget about any new competitors in search engine land now? Not only do you have to match Google in relevance you'll actually have to implement your own BrowserBot just to download the pages.
- mey 11y agoThat was my first reaction as well. "We've engineered a competitive advantage so why don't you throw out that hard work the helps our competitors." I'm not sure where I sit on this, developers who want to be noticed by other engines will continue to focus on SEO, but how many engineers care about SEO that isn't Google?
- vbezhenar 11y agoHonestly, optimizing sites for search is just wrong. It's happening now, because search is not perfect and developers have to work around its imperfectness. But in the ideal future, web masters must design websites for users, not for search engines, in the first and only place. That's what's happening now and it's good sign. Of course Google competitors must work hard. I don't see why that's a bad thing. It's not like Bing or Yandex are going to disappear in the foreseeable future.
- spacecowboy_lon 11y agoDepends what locale's you are targeting I am doing some strategy proposals for a client to help with their move into Asia. Biaudu is one SE that doesn't crawl JS well from my research.
- sheraz 11y agoI don't know that I would go head-to-head with Google in crawling the entire web. However, I do see a lot of opportunities for "vertical search." That is -- search engines focused on specific, niche verticals (travel, healthcare, etc) I'm working on a couple of projects in vertical search, and it is quite exciting. Sure, I'm building tech that Google had in 2005, but we are surprised with the results. We achieve search relevance simply by curating the sites we crawl (still in the thousands in some cases).
- espeed 11y agoThis was the missing piece for Polymer elements / custom web components. Now that Google has confirmed it's indexing JavaScript, web-component adoption should take off.
- tracker1 11y agoI want to like polymer/web-components... I just find that it kind of flips around the application controls that redux+react offers. I'm not sure that I like it better in practice.
- jwr 11y agoFinally. It was obvious we would have to get to that point eventually, it just wasn't clear when.
- deleted 11y ago[deleted]
- rcconf 11y agoThis might be obvious to anyone who has done SEO, but can Googlebot index React/Angular websites accurately? I was always under the impression that the isomorphic aspect of React helped with SEO (not just load times.)
- m0th87 11y agoDon't believe the hype. Google has been saying that they can execute javascript for years. Meanwhile, as far as I can see, most non-trivial applications still aren't being crawled successfully, including my company's. We recently got rid of prerender because of the promise from the last article from google saying the same thing [1]. It didn't work. 1: http://googlewebmastercentral.blogspot.com/2014/05/understanding-web-pages-better.html http://googlewebmastercentral.blogspot.com/2014/05/understan...
- thoop 11y agoTodd from Prerender.io here. We've seen the same thing with people switching to AngularJS assuming it will work and then coming to us after they had the same issue. [1] This image is from 2014, when Google previously announced they were crawling JavaScript websites, showing our customer's switch to an AngularJS app in September. Google basically stopped crawling their website when Google was required to execute the JavaScript. Once that customer implemented Prerender.io in October, everything went back to normal. Another customer recently (June 2015) did a test for their housing website. They tested the use of Prerender.io on a portion of their site against Google rendering the JS of another portion of their site. Here are the results they sent to me: Suburb A was prerendered and Google asked for 4,827 page impressions over 9 days Suburb B was not prerendered and Google asked for 188 page impressions over 9 days We've actually talked to Google some about this issue to see if they could improve their crawl speed for JavaScript websites since we believe it's a good thing for Google to be able to crawl JavaScript websites correctly, but it looks like any website with a large number of pages still needs to be sceptical about getting all of their pages into Google's index correctly. 1: https://s3.amazonaws.com/prerender-static/gwt_crawl_stats.png https://s3.amazonaws.com/prerender-static/gwt_crawl_stats.pn...
- eurokc98 11y agoGary Illyes @goog said this was happening Q1 this year, and like others mentioned lots of other direct/indirect signals have pointed this way. http://searchengineland.com/google-may-discontinue-ajax-crawlable-guidelines-216119 http://searchengineland.com/google-may-discontinue-ajax-craw... March 5th: Gary said you may see a blog post at the Google Webmaster Blog as soon as next week announcing the decommissioning of these guidelines. Pure speculation but interesting... The timing may have something to do with Wix, a Google Domains partner, who is having difficulty with their customer sites being indexed. The support thread shows a lot of talk around "we are following Google's Ajax guidelines so this must be a problem with Google". John Mueller is active in that thread so it's not out of the realm of possibility someone was asked to make a stronger public statement. http://searchengineland.com/google-working-on-fixing-problem-with-wix-web-sites-not-showing-up-in-search-results-233310 http://searchengineland.com/google-working-on-fixing-problem...
- nostrademons 11y agoI'm betting that they finally solved the scalability problems with headless WebKit. Google's been able to index JS since about 2010, but when I left in 2014, you couldn't rely on this for anything but the extreme head of the site distribution because they could only run WebKit/V8 on a limited subset of sites with the resources they had available. Either they got a whole bunch more machines devoted to indexing or they figured out how to speed it up significantly.
- tracker1 11y agoI'd say both are pretty likely.. another round of lower-power servers with potentially more cores... more infrastructure... Combined with improvements in headless rendering pipelines. I haven't looked into it in well over a year now, but last I checked dynamic updates took about 2-3 days to get discovered vs. server-delivered being hours for a relatively popular site. I'm guessing they've likely cut this time in half through a combination of additional resources, and performance improvements. Wondering if they'd be willing to push this out as something better than PhantomJS... probably not as it's a pretty big competative advantage. I know MS has been doing JS rendering for a few years, they show up in analytics traffic (big time if you change your routing scheme on a site with lots of routes, will throw off your numbers).
- shostack 11y agoAny idea how related this might be to Wix sites getting de-indexed?[1] http://searchengineland.com/google-working-on-fixing-problem-with-wix-web-sites-not-showing-up-in-search-results-233310 http://searchengineland.com/google-working-on-fixing-problem...
- nailer 11y agoCurrently I use prerender.io and this meta tag: <meta name="fragment" content="!"> I don't actually use #! URLs, (or pushstate, though I might use pushstate in the future) but without both of these Google can't see anything JS generated - using Google Webmaster Tools to check. Does this announcement mean I can remove the <meta> tag and stop using prerender.io now?
- rgbrgb 11y agoWe have a similar setup and were wondering the same thing (though we use push state). Today we were actually trying to figure out a workaround for 502s and 504s that google crawler was seeing from prerender. We just took the plunge and removed the meta tag because over 99% of our organic search traffic is from google. Fingers crossed!
- thoop 11y agoI'd love to help here if I can. I'd also love to hear the results of you removing the meta tag! todd@prerender.io
- thoop 11y agoIf Google Webmaster Tools is unable to render your website correctly, then that's a good indicator that Googlebot won't be able to render the pages correctly either. If you remove the fragment meta tag, then Google will need to render your javascript to see the page. Let us know how that goes if you try it! todd@prerender.io