9 ms·
Lets look, how this new wiki behaves without JS, from the viewpoint of a search engine. e.g. compare http://c2.fed.wiki.org/methodology-subsets.html http://c2.
by kephra 12y ago
Lets look, how this new wiki behaves without JS, from the viewpoint of a search engine.
e.g. compare http://c2.fed.wiki.org/methodology-subsets.html http://c2.fed.wiki.org/methodology-subsets.html with and without javascript. The new wiki is doomed by design, because content can not be found by Google or any other search engine. And a wiki where the content can not be found is for the trashcan.
- mbrock 12y agoThis is a specific issue to be solved with standard REST techniques. Do you know if the team is working on it? Maybe it's on a timeline, scheduled to be implemented next month? I haven't looked into it, but the current lack of this functionality does not mean that the "new wiki is doomed by design."
- Touche 12y agoGoogle executes JavaScript.
- zAy0LfpBZLC8mAC 12y agoThat can impossibly be true. JS is turing complete, turing completeness means that the halting problem is undecidable, which means that at best Google will execute JS with some random resource limit to avoid infinite resource consumption, which in turn means that whether Google will actually see the content is ultimately undefined. And even worse: other search engines might, due to lack of standardization of the available execution resources, see a different picture. That's just a braindead model for information storage and exchange.
- Touche 12y agoI don't know what that was all about, but Google does, here's a blog post about it: http://googlewebmastercentral.blogspot.com/2014/05/understanding-web-pages-better.html http://googlewebmastercentral.blogspot.com/2014/05/understan... And webmaster tools has a tool to verify that content is scraped correctly. Not saying there aren't other good reasons to do server-side rendering, but SEO is not a real one.
- JustSomeNobody 12y agoYou're trying to be too intelligent for your own good. Fact of the matter is the new wiki UI is horrible. It doesn't have to be spelled out any more complicated than that.
- mikeash 12y agoEverything you said applies to humans using normal web browsers as well. And yet people still fill pages with JS, and browsers still execute it, and people still see stuff.
- zAy0LfpBZLC8mAC 12y agoWith one minor difference: Humans at least have full-blown human intelligence which they can use to heuristically determine when the execution is "complete". Not that that makes all that much more sense ...
- mikeash 12y agoYeah, that does make it better, but it's still not perfect. Witness the complaints right here about how you can't tell whether you got an empty page or it just hasn't finished loading yet. Trouble awaits when you try to wedge a general-purpose application environment into a page-based document viewer.
- wahnfrieden 12y agoI doubt Google is worried about unfairly weighting pages it indexes with Javascript that takes so long to run, hence the cutoff. It's more a question of where to place the cutoff than whether to have one at all, which is absurd.
- zAy0LfpBZLC8mAC 12y agoNo, the question is neither of those. The question is how to define the cutoff in a portable way that enables interoperability and long-term stability. A document format where every implementation has its own secret cutoff that also probably changes all the time is just idiotic if your goal is interoperability.
- Touche 12y agoYou seem to be saying that page load is not deterministic but that is not true. A special version of a browser can easily keep track of http requests and other async operations and consider the page to be complete when async ops = 0 and the last paint is complete. The only hard part I can think of is a page that is recursively calling setTimeout endlessly, but even that can be coded around.
- geofft 12y agoThe exact same argument could apply to pages that take forever to load, because the web server sends you a chunk of bytes every two seconds. Would it be braindead for Google to apply an arbitrary timeout without standardization?
- zAy0LfpBZLC8mAC 12y agoNope, but it would be braindead for google to then just index however many bytes it has received up to that point as if it were the full document.
- geofft 12y agoOK, so we have a pretty straightforward rule. Providing no further input to the program once it begins, 1) If execution terminates by some timeout T, where T is at least several seconds, then index that. 2) If execution has not yet terminated by T, whether or not we have any idea whether it will terminate in the future, don't index. Tune T so that it will get the vast majority of reasonable web pages. (Hypothesize, and test, T = 5 seconds or 30 seconds or something. Have a human look at the highest-PageRanked timeouts to figure out what's going on.) This applies whether "the program" is server-side or client-side. The halting problem cannot be reduced to the timeout problem (since the halting problem asks if it ever terminates), and the timeout problem is pretty clearly computable.
- ogig 12y agoThis can be solved using any of these methods: https://developers.google.com/webmasters/ajax-crawling/docs/specification https://developers.google.com/webmasters/ajax-crawling/docs/... I maintain a client side rendered website and while SEO can be problematic it can also be done. I don't see C2 using any of these methods at first glance tho.
- blueskin_ 12y agoIt can more easily be solved by just not crawling sites that have mandatory javascript, because chances are their content is going to be crap and not worth the time if they are that hostile to basic usability.
- gavinpc 12y agoIt has also joined the ranks of sites that don't work at all without cookies. Go ahead, try it.