3 ms·
Um, there is a time and place for client side MVC and there is a place for server side. For example, if you want Google to crawl your stuff, go server side. I
by programminggeek 14y ago
Um, there is a time and place for client side MVC and there is a place for server side.
For example, if you want Google to crawl your stuff, go server side.
I think it's smart to look at it on a per-page or per app basis. Not everything should be client side MVC and not everything should be server side.
Do the right thing for your product.
- ymn_ayk 14y agoGoogle started to understand javascript, didn't it?
- orthecreedence 14y agoTo add a note: if you have an app that you want Google to index but all of the other points weigh in on the all-JS app side, you can use PhantomJS to crawl/scrape/store your pages before a googlebot ever shows up. Yes, it's a pain in the ass, and yes it complicates your back-end...but it's a good way to have the best of both worlds.
- jbigelow76 14y agoI'd love to see an actual nuts and bolts implementation of the method you describe. I've seen variations of it discussed in other threads but it's all just theory. Deciding how to serve up pages based on user agent, scheduling refreshes, etcetera. It all seems extremely complicated, I'd especially like to know if the implementer viewed the ROI as worth it after the fact.
- orthecreedence 14y agoWell, I've implemented it personally. Generally, here's what happens: 1. A user loads the first page of the app. There is a small script (index.php) that serves all front-end requests. It checks if the page being loaded has already been scraped and if so, injects the content into the main template, otherwise fires off a queued job to scrape that page and loads the app regardless. 2. The queue job forks out to phantomjs with the page URL, which listens on the page for a "page-loaded" event, and upon a) this event being fired or b) the scrape timing out, returns whatever HTML content in the main content area (as well as any <head> tags) to the queue worker, which updates a database table with url -> content mappings. 3. The next time someone directly loads the page, that content is pulled and injected. 4. Scraped content expires after a certain amount of time. If expired content is pulled out, it is put into the HTML for that page, and a queue item is fired to re-scrape. The obvious problem is that the page must be primed for content to show up, which you can fix by a) either having lots of visitors :) or b) using/building a simple crawler. Also, when facebook scrapes our page to look for og: tags, we do the scrape synchronously which takes longer, but ensures that the page's content is always available. The same could be done for googlebot, but it's more dangerous since rankings depend on page load speed. As for our rankings, they are fairly high (it's the same as any app that serves HTML directly would be since we are literally just serving delayed HTML content). So no, it's not just theory; we do it on Musio.com quite successfully. The only user-agent logic we have is for the facebook scraper, refreshes are scheduled by user action. The RIO is high because 1. We already have a queuing system. 2. Phantom JS is fairly standalone and our scraper is small and write once then forget about it. 3. We've found a good blend between simple and functional, with the deciding factor being the delay.
- jbigelow76 14y agoThanks for taking the time to do a detailed write up, it's very enlightening. >So no, it's not just theory; No offense was intended, I phrased my statement poorly. The technique is usually discussed in more abstract terms would have been better.
- orthecreedence 14y agoNone taken! I understand you probably hadn't encountered anybody who had done more than just toyed with the idea. I also agree, people like to talk about how cool it might be, but I hadn't heard of anyone doing it in production before us. I think one of the reasons it isn't done a lot is because when people think "one-page JS apps" it's usually account-based, making you log in to use the app (at least the ones I use). So other than a few marketing pages which could easily be handled by a CMS, there'd be no purpose to store generated HTML for your own pages.