8 ms·
Defcon: Preventing overload with graceful feature degradation (2023)
- Banditoz 3y agoAm I reading the second figure right? Facebook can do 130*10^6 queries/second == 130,000,000 queries/second?!
- sebzim4500 3y agoSounds plausible. There are probably many queries required to display a page and Facebook has 2 billion daily active users.
- ipaddr 3y agoThis is how information slowly changes. The original numbers from facebook needed to be taken with a grain of salt. 2 billion a day raises it more. Facebook claims to have 2 billion accounts but no where near 2 billion unique accounts. I don't know what facebook calls an active user but it use to mean logged in once in the past 30 days.
- reissbaker 3y agoNo, the person you're responding to was correct. Facebook has over 2 billion daily active users [1], and DAU refers to unique users who used the product in a day [2]. 1: https://www.statista.com/statistics/346167/facebook-global-dau/ https://www.statista.com/statistics/346167/facebook-global-d... 2: https://www.innertrends.com/blog/active-users-measuring-business-success-with-dau-wau-and-mau https://www.innertrends.com/blog/active-users-measuring-busi...
- ipaddr 3y agoIt's directly in the link you provided. "For example, an active user can be measured as a user that has logged back into her account to interact with the product in the last 30 days." Even the marketing material is designed to confuse.
- ipaddr 3y agoIt should be their account not her account (or his). Who writes this garbage.
- OrsonSmelles 3y agoAlternating or stochastically varying pronouns in your examples used to be a common way to make an effort at inclusive writing, usually preferred aesthetically to constructs like `his/her'. (The style before that was basically to use masculine pronouns for hypothetical people in every single case and deny that there was anything to question about that.) I think I agree that the modern semi-standard of using `they' for examples where gender is irrelevant or unknown is strictly better, but it's hard for me to summon a lot of contempt for someone who goes with a different/older habit.
- deleted 3y ago[deleted]
- rcxdude 3y agoThat's a monthly active user. A daily active user would be someone who logged into an account in the last day. Generally monthly active user count will be higher than daily active users, but for something like Facebook the difference is about 50% (which is what the second article linked is explaining, if you read more than just cherry-picking a line that matches your preconceptions) And yes, that's a claim that if each user is a separate person, >20% of the world's population interacts with Facebook at least minimally each day. You can add your own interpretation about how many of the accounts are bots or otherwise duplicates, but it's a staggering amount either way.
- zilti 3y agoHeh. Gave me a chuckle, because DAU also means "Dümmster Anzunehmender User" in German (dumbest assumed user, in context of creating idiot-proof software, and a wordplay on GAU, which means grösster anzunehmender Unfall, biggest assumed accident, which comes from fission power plants). And that kinda fits for the kind of people that perceivedly are left on the likes of Facebook and X.
- pests 3y agoDifferent metrics Daily vs Monthly Active User. MAU vs DAU
- IncreasePosts 3y agoI forgot how to count that low.
- sonicanatidae 3y agoYeah, they allocated ALL of the ram to their DB servers. lol
- storyinmemo 3y agoI think about 10 years ago when I was working there I checked the trace to load my own homepage. Just one page, just for myself, and there were 100,000 data fetches.
- vaylian 3y agoBy "homepage" you mean your Facebook profile?
- gaogao 3y agoThose queries are probably mostly memcache hits, though of course with distributed cache invalidation and consistency fun
- ipaddr 3y agoIf it doesn't hit the database is it really a query?
- isbvhodnvemrwvn 3y agoWhy wouldn't it be?
- bagels 3y agoI can't comment on the numbers, but think of how many engineers work there and how many users Facebook, Whatsapp, Instagram have. Each engineer is adding new features and queries every day. You're going to get a lot of queries.
- bee_rider 3y agoWe’ve really wasted an incredible amount of talent-hours over last couple decades. Imagine if we’d worked on, like, climate change or something instead of ad platforms.
- esafak 3y ago"The best minds of my generation are thinking about how to make people click ads. That sucks." - Jeff Hammerbacher (2011); early Facebook employee, and Cloudera cofounder. https://www.theatlantic.com/technology/archive/2011/04/quote-ad-generation/349689/ https://www.theatlantic.com/technology/archive/2011/04/quote...
- baby 3y agoYou make it sound like everybody at Meta works in the ads department.
- Scubabear68 3y agoThey all do.
- jojobas 3y ago98% of Meta's revenue is from ads. Meta is an ads department.
- bee_rider 3y agoIt is an ad company, everyone there works on ads or indirectly works on making a platform for ads. The only exception is people who’ve managed to sneak their way into positions where they don’t contribute anything to the company. Those people are doing society a favor by wasting Facebook’s money.
- reissbaker 3y agoA custom JIT + language + web framework + DB + queues + orchestrator + hardware built to your precise specifications + DCs all over the world go a long way ;)
- AlienRobot 3y agoiirc Facebook has 3 billion users, so that sounds plausible.
- ndriscoll 3y agoFacebook makes over 300 requests for me just loading the main logged in page while showing me exactly 1 timeline item. Hovering my mouse over that item makes another 100 requests or so. Scrolling down loads another item at the cost of over 100 requests again. It's impressive in a perverse way just how inefficient they can be while managing to make it still work, and somewhat disturbing that their ads bring in enough money to make them extremely profitable despite it.
- golergka 3y agoWasn't the whole point of GraphQL in mitigating this?
- serial_dev 3y agoYeah, that's why you have only 100 requests* when you hover over an item instead of 800. (* allegedly, didn't verify it myself)
- zer00eyz 3y agoNo. Here is the thing, hypermedia is cacheable. React/Graphql not so much. Facebook is now just an application that runs in the browser. As a poor, small developer who doesn't want to hemorrhage money, I tend to want things to be more hypermedia and less app. It saves on complexity and bandwidth and costs.
- yodsanklai 3y agoCould someone tell me what these hundreds of requests could do?
- Solvency 3y agoTrack you, probably with a thousand layers of redundancy, tech bloat, and decades of mold.
- ndriscoll 3y agoA lot of them appear to be that they've split their javascript into a gazillion files for whatever reason (I suppose because they have several MB of it). But someone or lots of people there did seem to get addicted to dynamic loading. Like I've got 100 or so friends, but my friends page loads them 8-16 at a time as I scroll. Just send all 100 and set the profile pictures to deferred fetch. It'd probably be smaller than the js they have to make it do "infinite" scroll. Similarly, after getting to the bottom of their "infinite" scroll, my friend feed (which is annoyingly hidden away) gives me... 15 items. Just send me all 15. It's like 1-2 kB worth of data. If you're going to end the scroll after a dozen items, why is it using infinite scroll?
- bdd 3y agoYes. And that was 4 years ago. Must add that figure does NOT include static asset serving path.
- scottlamb 3y ago> Am I reading the second figure right? Facebook can do 130*10^6 queries/second == 130,000,000 queries/second?! That sounds totally plausible to me. Also keep in mind they didn't say what system this is. It's often true that 1 request to a frontend system becomes 1 each to 10 different backend services owned by different teams and then 20+ total to some database/storage layer many of them depend on. The qps at the bottom of the stack is in general a lot higher than the qps at the top, though with caching and static file requests and such this isn't a universal truth.
- Thaxll 3y agoWe're close to 1 million servers, not 12 racks in a DC.
- avery17 3y agoWhats with people lately writing 10^6 instead of 1 million. Its not that big that we need exponents to get involved.
- rijx 3y agoErosion of education makes basic scientific knowledge very trendy
- jcparkyn 3y ago- The comment is referring to a graph that used 10^6 on the vertical axis, which is a very common way to format graphs with large numbers (not just "lately"). It's also the default for a lot of plotting libraries. - 10^n is more compact than million/billion/etc, more consistent, easier to translate, and doesn't suffer from regional differences (e.g. "one billion" is a different number in Britain than in the US). I'm not saying it's clearly better than "million" in this specific case, but it's definitely not clearly worse.
- deleted 3y ago[deleted]
- velcrovan 3y agoSeems like whenever I log into FB lately it's pretty much always in a state of “graceful feature degradation”. For example, as soon as I log in I see a bell icon in the upper right with a bright red circle containing an exact positive integer number of notifications. It practically screams “click here, you have urgent business”. I can then leave the web page sitting there for any number of minutes, and no matter how long I wait, if I click on that notification icon it will take a good 20 seconds to load the list of new notifications. (This is on gigabit fiber in a major metro area, so not a plumbing issue.)
- philippta 3y agoWithout being able to verify, I would assume it’s designed to behave in this way. The longer you wait the more anticipation builds up, the more gratifying it becomes.
- meowface 3y agoI think there's no chance they intentionally want users to wait 20 seconds to see their latest notifications.
- Arainach 3y agoThe initial render of Facebook's UI slows dramatically (I suspect but cannot prove intentionally) if you have adblockers/uBlock Origin/etc.
- guessmyname 3y agoHave you tried navigating the website using a web proxy (Charles, Burp Suite, or similar tool) to intercept the HTTP request(s) in order to replay them yourself multiple times to see if the latency is consistent? It’d be interesting to discover that the delay is fabricated using the front-end code or if the back-end server is really the problem. I don’t use Facebook but I asked a friend just now and the response time for the notifications panel to appear is between 500ms-2000ms, which is relatively fast for web interactions.
- deleted 3y ago
- OtherShrezzing 3y ago> if (disableCommentsRanking.enabled == False) This could use some light-touch code reviewing
- deleted 3y ago[deleted]
- dvhh 3y agoSome could argue it would be for illustration purpose, and not actual production code
- kqr 3y agoBecause the HN crowd likes learning new things: if `enabled` is a nullable boolean in C# (i.e. has type `bool?`) then this check must indeed be written this way, to avoid confusing null with false.
- jpnc 3y agoI thought OP meant to imply that the readability could use some tweaking. You have 'disable', 'enabled' and 'False' used in the same expression so it requires some (more) thinking while reading it and trying to decipher what it's trying to do.
- kqr 3y agoThat is a more fundamental and better criticism that I'm embarrassed I overlooked.
- bmacho 3y agoIt looks funny, but I think it's actually good, and arguably the best possible form of it.
- mrb 3y agoOff-topic but: I love the font on the website. At first I thought it was the classic Computer Modern font (used in LateX). But nope. Upon inspection of the stylesheet, it's https://edwardtufte.github.io/et-book/ https://edwardtufte.github.io/et-book/ which was a font designed by Dmitry Krasny, Bonnie Scranton, and Edward Tufte. The font was originally designed for his book Beautiful Evidence. But people showed interest in font, see the bulletin board on ET's website: https://www.edwardtufte.com/bboard/q-and-a-fetch-msg?msg_id=0000bm&topic_id=1&topic=Ask%20E%2eT%2e https://www.edwardtufte.com/bboard/q-and-a-fetch-msg?msg_id=... Initially he was reluctant to go the trouble of releasing it digitally. But eventually he did make it available on GitHub.
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- dang 3y agoDiscussed (a tiny bit) at the time: Defcon: Preventing Overload with Graceful Feature Degradation - https://news.ycombinator.com/item?id=36923049 https://news.ycombinator.com/item?id=36923049 - July 2023 (1 comment)
- danpalmer 3y agoJoining Google a few years ago, one thing I was impressed with is the amount of effort that goes into graceful degradation. For user facing services it gets quite granular, and is deeply integrated into the stack – from application layer to networking. Previously I worked on a big web app at a growing startup, and it's probably the sort of thing I'd start adding in small ways from the early days. Being able to turn off unnecessary writes, turn down the rate of more expensive computation, turn down rates of traffic amplification, these would all have been useful levers in some of our outages.
- tuyguntn 3y agoit's really great to have such capabilities, but adding them has a cost where only few can afford. Cost in terms of investing in building those, which impacts your feature build velocity and the maintenance
- zerkten 3y agoCan you be specific about the cost of building these? I've run into many situations where something was deemed costly, is found out later, and the team ultimately has implement it all while hoping no one groks that is was predicted. "Nobody ever gets credit for fixing problems that never happened" (https://news.ycombinator.com/item?id=39472693 https://news.ycombinator.com/item?id=39472693) is related.
- spacebanana7 3y agoThe developer, tester and devops time required to properly implement graceful degradation could easily accumulate to hundreds of hours. Those hours are directly expensive when your developers cost hundreds of dollars a day; and have a material opportunity cost in that their commitment to one particular project delays the delivery of other features. Moreover, any new features would have to be made compatible with the graceful degradation pattern, creating an ongoing cost.
- IggleSniggle 3y ago
- deleted 3y ago[deleted]
- mikerg87 3y agoIsn't this referred to as Load Shedding in some circles? If its not, can someone explain how its different?
- scottlamb 3y agoThey're the same thing or close to it. "Load shedding" might be a bit more general. A couple possible nuances: * Perhaps "graceful feature degradation" as a choice of words is a way of noting there's immediate user impact (but less than ungracefully running out of capacity). "Load shedding" could also mean something less impactful, for example some cron job that updates some internal dashboard skipping a run. * "feature degradation" might focus on how this works at the granularity of features, where load shedding might mean something like dropping request hedges / retries, or individual servers saying they're overloaded and the request should go elsewhere.
- deleted 3y ago[deleted]
- kqr 3y agoThis is the other side of the load shedding coin. The situation is that A depends on B, but B is overloaded; if we allow B to do load shedding, we must also write A to gracefully degrade when B is not available.
- jedberg 3y agoI'm surprised they don't have automated degradation (or at least the article implies that it must be operator initiated). We built a similar tool at Netflix but the degradations could be both manual and automatic.
- packetslave 3y agoThere's definitely automated degradation at smaller scale ("if $random_feature's backend times out, don't show it", etc.). The manual part of Defcon is more "holy crap, we lost a datacenter and the whole site is melting, turn stuff off to bring the load down ASAP"
- gillh 3y agoAnyone interested in load shedding and graceful degradation with request prioritization should check out the Aperture OSS project. https://github.com/fluxninja/aperture https://github.com/fluxninja/aperture
- winrid 3y agoOne of the most satisfying feature degradation steps I did with FastComments was making it so that if the DB went offline completely, the app would still function: 1. It auto restarts all workers in the cluster in "maintenance mode". 2. A "maintenance mode" message shows on the homepage. 3. The top 100 pages by comment volume will still render their comment threads, as a job on each edge node recalculates and stores this on disk periodically. 4. Logging in is disabled. 5. All db calls to the driver are stubbed out with mocks to prevent crashes. 6. Comments can still be posted and are added into an on-disk queue on each edge node. 7. When the system is back online the queue is processed (and stuff checked for spam etc like normal). It's not perfect but it means in a lot of cases I can completely turn off the DB for a few minutes without panic. I haven't had to use it in over a year, though, and the DB doesn't really go down. But useful for upgrades. built it on my couch during a Jurassic park marathon :P
- sara44444444 3y ago[flagged]