4 ms·
An engineering blog from the BBC published an update a few days ago on their migration to a serverless architecture, titled "BBC Online – A year with serverless
by thamer 5y ago
An engineering blog from the BBC published an update a few days ago on their migration to a serverless architecture, titled "BBC Online – A year with serverless": https://medium.com/bbc-design-engineering/bbc-online-a-year-with-serverless-ffc2ae474277 https://medium.com/bbc-design-engineering/bbc-online-a-year-...
They report 3.3B function invocations for 2.3B requests served, a total of 61,000 hours of execution time and 1,500 concurrent executions at peak.
Using calculator.aws, simply entering the number of requests at 67ms each (61k hours/3.3B invocations) with 128 MB of memory each computes a cost of only $1,107.
Then of course there's bandwidth, databases, file storage, and a lot more… but 3.3B lambda executions for $1.1k is really cheap. Not to mention that a client like the BBC would benefit from negotiated pricing with significant discounts. Even without discounts, serverless is often pretty cheap.
- ec109685 5y agoThat’s only averaging 72 request per second. Yes they, scale to 1,500 requests per second sometimes, which is an advantage of lambda, but we are talking about pretty small traffic. Still interesting though and glad it works for them.
- sofixa 5y agoAnd that's where serverless excels, highly variable loads.
- recuter 5y agoYou can certainly run a web server capable of 1,500 concurrent connections for under $100/mo. Run it on Linode or something, they'll throw in 10TB of bandwidth which probably covers it. Most of the serving is done by the CDN..
- renewiltord 5y agoRight, but once you're in the territory of "everything for $100" or "everything for $1000", they're both the same so you want the thing with the organizational efficiencies of independent deploys and permissioning, resilience and scaling, high availability etc.
- recuter 5y agoTwo $60 Linodes in different data centers are highly availablish considering each one could easily do x10 those numbers. The thing in the link sits behind a CDN. Certain kinds of organizational efficiencies frequently lead to "I'm bored lets use the new hotness that would look good on my resume for the next job", hence lambda. Actually you don't need anything at all, this is a job completely for the CDN which has enough features they aren't using to do the same thing for no additional cost.
- lostcolony 5y agoYou seem to have missed the aim of the game on their part is personalization. While it might be possible that they could do some client side processing, split the page up sufficiently, to then make multiple requests and assemble a "personalized" page, it's hard to say for certain that it would be. Certainly, it makes it less "personalized" and more "regionalized", which though being the case they describe, is likely a decision they didn't want to be locked into. For it to be truly personalized it likely needs to pull in some backend data. Given that...you can't rely on the CDN being an optimization. You could still make the personalized calls through the CDN, if you wanted and had a clear caching policy, but something like "on load, assemble a page with real time updates on the stories this person expressed interest in" isn't going to benefit from a CDN.
- deleted 5y ago[deleted]
- recuter 5y agoHeh. OK for fun: They have a very finite amount of new articles a day and a little box on the page with personalized headlines. All these articles, in fact all the articles on BBC of all time ever, comfortably fit into ram if you are so inclined. Those little linodes have fast enough SSDs, try playing around and use one of countless stress loaders to get a feel. I've served hundreds of thousands of requests a second from a single server using Varnish. A well tuned beefy VPS would do tens of thousands. How to "personalize"? For example Varnish has this thing: https://en.wikipedia.org/wiki/Edge_Side_Includes https://en.wikipedia.org/wiki/Edge_Side_Includes https://varnish-cache.org/docs/6.0/reference/vcl.html https://varnish-cache.org/docs/6.0/reference/vcl.html It will happily read whatever cookie you set and stitch together a remixed personalized page for all the BBC users in the world without a performance penalty much faster than your lambda cold starts. CDNs sometimes let you hook into Varnish directly or expose something similar to essentially do exactly this. The aim of the game of the guy who wrote that blog post is to reinvent something from the 90s, but worse, and blog about it.
- outsb 5y agoThe CPU throttling of the lower memory settings is often easily visible in a user-facing request, even simple requests just doing some DB IO. I rarely use less than 1024mb for anything edit: the BBC review is horrifying: > The page takes around 500ms to render and be delivered to the audience. In that timeframe we invoke around 30 functions. Around 150ms is spent running React to render the content to HTML > we aim to personalise almost every page in some way — making it relevant for every user on every request Good luck making perf numbers with all those cold cached personalized pages
- lostcolony 5y agoI want to call out - performance numbers don't actually mean nearly as much as people have made them out to be. This is a mostly read only web page. Half a second to load? You're barely going to lose anyone, if you lose anyone at all. Hacker News routinely runs me ~300ms to load and has zero personalization, Facebook takes over a second and a half before anything displays, as does Youtube (on a refresh, no less, so things should reside in cache locally!). Hitting a random person's LinkedIn page (once I've passed the verification, which is a whole different issues) takes 1.2 seconds. Etc. Now, admittedly those are including the latency on my end, but the point is, no one is so meth addled that a page loading after half a second (or even a full second!) is going to have much effect on engagement. Even the studies that have been done (that I have some major issues with) only really start measuring anything significant well after a second or two.
- outsb 5y agoThe BBC article is only talking about generation time, specifically not download time (including linked assets), and we have no idea about dwell time. It only takes a 3G connection a few miles outside a city to add another 500ms to that opening request. Say we're up to a second before some readable text appears, now we'd like to know how long the user will actually spend reading the text or waiting for images to load before navigating again. Intuitively, I think that load time/dwell time ratio probably captures what those studies talk about better than just raw numbers. 1 second between 10 second TikTok video loads would be extremely noticeable, but barely worth mention if the user instead was spending 10 minutes reading e.g. a feature length news article. My personal BBC reading habit regularly involves clicking into an article just to catch the opening paragraph and seeing which opening image they used (they rarely use the same for the thumbnail). The average is probably not as low as 10 seconds, but it's certainly something much less than 2 minutes. Dwell time probably isn't a great way to capture it either. My pattern is quite "flicky" but I bet there is a spectrum all the way from "reads every last word" to "literally just loves to click". I guess latency becomes increasingly important for folk further along that spectrum