4 ms·
The slides say there are 1 million pages, and "republishing" them would take 90 days. Maths -> 7.8 seconds on average to "republish" a page. Modern templates sy
by nanoscopic 12y ago
The slides say there are 1 million pages, and "republishing" them would take 90 days. Maths -> 7.8 seconds on average to "republish" a page. Modern templates systems can convert a page constructed out of structured data in less than 1/4 of a second. ( and that is a high estimate ) That is ~30 times faster, meaning all pages could be "republished" in 3 days instead of 90 had a more efficient system been used from the start.
Focusing on shifting all "rendering" into front end JS seems like it will lead to more difficulty in the long run instead of using a more efficient structured page creation mechanism.
I am curious how the static pages were created. Others here are speculating that templating was not done. If not; what does "republishing" mean exactly?
- ben336 12y ago"From the beginning" in this case going back to the mid 90s. Pretty sure "Modern template systems" have mostly been written since then.
- nanoscopic 12y agohttp://search.cpan.org/~mjd/Text-Template-0.1a/ http://search.cpan.org/~mjd/Text-Template-0.1a/ Basic text templating; circa 1995. Using templates is hardly a new thing. I wrote a full CMS system for a news company in 2002. Sure that is a couple years later; but I based it on Text::Template in combination with XML; both of which existed in the mid 90s. I am not attacking what they were doing, my point was that this is an issue of delaying the move to a template based system way longer than necessary. I am speculating on whether they did move to some templating, but not all the way, and wondering what exactly they were using before that. Saying "static HTML files" is well and good, but it doesn't explain how they were created. By hand?? Also, saying "90 days to republish" seems to be suggesting some process they have in mind besides manually scraping the data out. If you have 1 million random files, scraping could easily take years not 3 months. It would be interesting to know what process they are suggesting to follow in those 90 days. My speculation is they did use some sort of outdated CMS software.
- bonaldi 12y agoTheir source material here is 1 million pages of HTML, they don't have (on my reading) some separate source of "structured data" for the modern template system to use. It seems reasonable (and possibly low) for a 90 days estimate to extract the content from the variety of versions of static page, structure it, and then publish it in a more modern fashion. It's all very well to say they should have used a more efficient system from the start, but "the start" in this case is 1996, which is the wild west in terms of best practices.
- thezilch 12y agoIn 1996 (or as late as early 2000s), many published texts were via tools like Dreamweaver, where the publisher would checkout / lock a file, make static changes, and save the file, directly on prod. Like you said, it's not surprising at all they have a large portion of their articles in static HTML files.
- eitanmk 12y agoExactly. To compound the issue, older content may not fit into our current schema, so there are other data issues that need to be solved. Adding hand-coded HTML to pages was quite common 10 years ago, so parsing that into the structures we work with today isn't always straightforward and difficult to automate.
- Zizzle 12y agoIt also seems like a task with inherent parallelism. Upload the corpus and throw a bunch of cloud instances at it.
- eitanmk 12y agoAs a disclaimer, this falls under the "Back then" section of the talk which is an overview of how things used to work. We no longer do things this way. Publishing a page is running content from a CMS through a templating system. However, the time spent executing templates isn't the only factor in the duration of the "publish". The slides refer to a compilation step (which actually also included a preprocessor step), and includes delegating to a service to copy the resulting page to disk and ensure the write succeeds for all data centers. For data consistency and system monitoring, we essentially treat that entire process as an atomic action and wait for all parts to finish. Additionally, since "publishing" is a core process for us, we avoid doing massive publishes that might risk the systems involved in the successful publishing of current articles. So increasing the number of these for the sake of pushing code is considered too risky. Yes, there are ways of mitigating that risk, but dealing with this legacy problem once and for all is a better path forward than scaling up this solution.
- nanoscopic 12y agoThank you for this detailed explanation. I assumed there was some extenuating complexity to the process, and was simply wondering what it was that prevented the original publishing system from being adapted to be faster.