4 ms·
> Users can place one tile every 5 minutes, so we must support an average update rate of 100,000 tiles per 5 minutes (333 updates/s). It only takes a couple of
by vmasto 9y ago
> Users can place one tile every 5 minutes, so we must support an average update rate of 100,000 tiles per 5 minutes (333 updates/s).
It only takes a couple of outliers to bring everything down. I'm not exactly well-versed in defining specs for large scale backend apps (not a back-end engineer) but it seems to me that preparing for the average would not be a wise decision?
For example, designing with an average of a million requests per day in mind would probably fail, since you get most of that traffic during daytime and far more less at the nightly hours.
Could anyone more experienced shed some light?
- deleted 9y ago[deleted]
- andoon 9y agoThe entire reddit website goes down every night, especially during weekends, sport matches, etc, so there you have your answer.
- celticninja 9y agoIs that hyperbole or are you really experiencing that much downtime of Reddit? I see it occasionally but it's never down for long, the odd "servers are busy" message usually disappears after a single refresh.
- andoon 9y agoPages take a long time to generate all day, but during peak hours they take a minimum of 4 seconds each (depends on what page you're loading, if it's got lots of comments, etc), and many times they simply timeout. The engineers at reddit have been unable thus far to fix it.
- rhizome 9y agoThis isn't my experience.
- freehunter 9y agoI'm not the same guy you were just talking to but in my experience on mobile Safari, I get "something went wrong, visit the homepage" when I navigate around reddit far too often. What's funny is it actually tends to resolve itself if I wait a second or two, but yeah, reddit engineering and infrastructure is not well equipped to handle the amount of traffic they receive. It's gotten better, but it's still not what you'd expect from the internet's front page.
- noir_lord 9y agoThe main website is OK as is i.reddit.com but their mobile client is hilariously bad in a "I can't believe they thought this was ready" way, it takes an age to load, it stalls out all the time, the tap targets are too small, it has the classic "oh you hit this when you meant that and then clicked that, lets spin the wheel on where you really end up" problem that slow mobile apps have. I know they deal with insane scale yadayada but it's simply not ready. That and the horrific dark pattern on the "We want you to have the best, massive red/orange button marked continue that takes you to the app store and the tiny weeny little "continue to mobile site" underneath". I don't want your damn app, stop asking me. If I was cynical I'd think they didn't care about the mobile site been awful as it drives people to the app.
- francoisfeugeas 9y agoLast time I tried it, their app was even worse than the mobile website though.
- tummybug 9y agoI can't agree with you. The sheer fact that I know well what the reddit error page is refutes this. In fact it's one of the few sites I frequent that I know even have an error page.
- rhizome 9y agoMaybe it's regional.
- ClassyJacket 9y agoThe search is consistently broken, but the rest of the site seems to have good uptime now, much better than it once was.
- celticninja 9y agoSearch has never worked. I always use Google site: modifier and have never had an issue.
- Macha 9y agoSearch worked pretty reliably up until 7-8 months ago for me
- notatoad 9y agomaybe three years ago it did, but reddit has gotten drastically more stable since then. It still has the occasional downtime, but now it's more like every couple months than every couple days.
- SquareWheel 9y agoReddit itself has been very stable these last few years. Reddit search on the other hand is a complete crapshoot even to this day.
- bsimpson63 9y agoThis was just a back of the napkin estimation. I think at one point we calculated that we'd be able to support one tile placement per user per second.
- deleted 9y ago[deleted]
- tyrust 9y agoThey might have phrased this poorly. >We should support at least 100,000 simultaneous users. This line makes me think that this is what they expected the peak (or near peak) to be. >Users can place one tile every 5 minutes, so we must support an average update rate of 100,000 tiles per 5 minutes (333 updates/s). So assuming that they mean that 100k is the peak and that clients are limited to 1 update per 5 minutes, they can expect 333 updates per second on average. The "average" is taken over this 5 minute period. This average represents the number of queries per second they will get if everyone's 5 minute cooldown is spread out evenly over each 5 minute period. It is possible, for example, for half of the peak population's cooldown to expire at 1300 and the other half to expire at 1305. In this case the average updates/s over the 5 minute period from 1300-1305 would still be 333 updates/s even though there were really 2 bursts of 50k a second at 1300 and 1305. It's far more likely that cooldowns are not excessively stacked in this way, so you prepare for the average and hope for the best.
- foota 9y agoPrepare for the average and hope for the best seems like a great way for things to fail.
- tyrust 9y agoThey prepared for the worst (peak 100k users). The "hope" that those 100k would be spread out was based on the statistical likelihood that these 100k wouldn't line up too much over a 5 minute period. I didn't follow /r/place that much, but I haven't read any complaints about latency or failures so it looks like they did just fine.
- foota 9y agoThey say elsewhere that they were prepared for 100k updates all at once. Which is the worst. Edit: you're an sre, you probably have more experience planning these things than I ever will. I just can't help but think about what would happen if some trolls realized they could synchronize their updates and bring down the service. (Although based on the infrastructure it doesn't seem possible to cause lasting damage.)
- avip 9y agoI can shed that 333 [something/s], for any uncomplicated something, is so little that a single-core 10YO laptop running a single-process non-blocking webserver (a-la node) could probably handle it, and likely x10 it.
- tummybug 9y agoSure in a perfect scenario with a well behaving load test client. In production you often encounter scenarios which place unexpected load on your servers.
- eric_h 9y agoAh yes, the self-inflicted unintentional DDoS attack that doesn't appear when you do a semi-idealized load test. I may or may not have been responsible for one of those once or twice in my life.
- sangnoir 9y ago> It only takes a couple of outliers to bring everything down. I'm not exactly well-versed in defining specs for large scale backend apps (not a back-end engineer) but it seems to me that preparing for the average would not be a wise decision? It's not an average: Reddit controlled the 5 minute user cool-down period, a.k.a. request throttling. 333 updates/s was the capacity. As alluded to in the article, the 5 minutes was dynamically configurable: if more than 100,000 had users showed up, they would have increased the cool-down period to a value that would yield 333 updates per second at most.
- raziel2p 9y agoWith 100K users and a minimum cooldown time of 5 minutes, there's 100K tiles per 5 minutes would be an upper limit. Only bots would hit that 5 minute cooldown every single time, and bots make up a very small minority of the users. They even point out later in the article that they only peaked at 200 updates per second. > For example, designing with an average of a million requests per day in mind would probably fail, since you get most of that traffic during daytime and far more less at the nightly hours. That depends. If you're talking about a public website, then yes. But this is just a single API endpoint with a very fixed time-based rate limit.