6 ms·
What are the main factors that could lead to such a situation?
by x3tm 8y ago
What are the main factors that could lead to such a situation?
- taneq 8y agoWell, the "good enough" point is reached when the cost of improving the reliability of the service is more than people would pay for that improved reliability.
- gizzlon 8y agoFrom the first SRE book [1]: "The error budget stems from the observation that 100% is the wrong reliability target for basically everything (pacemakers and anti-lock brakes being notable exceptions). In general, for any software service or system, 100% is not the right reliability target because no user can tell the difference between a system being 100% available and 99.999% available. There are many other systems in the path between user and service (their laptop, their home WiFi, their ISP, the power grid…) and those systems collectively are far less than 99.999% available. Thus, the marginal difference between 99.999% and 100% gets lost in the noise of other unavailability, and the user receives no benefit from the enormous effort required to add that last 0.001% of availability. If 100% is the wrong reliability target for a system, what, then, is the right reliability target for the system? This actually isn’t a technical question at all—it’s a product question.." [1] https://landing.google.com/sre/sre-book/chapters/introduction/ https://landing.google.com/sre/sre-book/chapters/introductio... https://landing.google.com/sre/books/ https://landing.google.com/sre/books/
- SilasX 8y agoPlease, please stop using mono space for quotes. It’s hard to read on small screens. Just > is better.
- codetrotter 8y agoIt’s not hard. It’s actually impossible. I agree, just use >. If I worked for YC as a HN mod I would literally spend a bit of time every day to review as many mono space-using comments I could every day and edit them to use > instead.
- hawaiianbrah 8y agoOr write a bot to do it automatically everywhere
- lftl 8y agoOr fix the CSS
- deathanatos 8y agoThe monospace is, I think, intended for code, and a lot of us use it that way. You want something separate for code blocks and block quotes here.
- deathanatos 8y agoOr give us actual syntax for blockquotes? (> at the beginning of a line, markdown style, would be great…) Which I feel like gets used all the time in the discourse here, and for good reasons, too. Then you would have that bit of time back, and the rest of us could stop scrolling back and forth when someone code-blocks a blockquote.
- jessaustin 8y agoWhy not use italics instead? How does the ">" help? You can put it at the beginning of the paragraph, but lines get re-flowed so you can't put it at the beginning of lines.
- matt4077 8y agoThe > was the original (?) convention for email. But seriously HN should just allow basic quotes. And, while they are at it, increase the size of voting arrows to make it possible to reliably hit them on mobile. I get it’s somewhat nice to have a site that does not constantly redesign everything, but pretending you lost the password for the server is just overdoing it in the opposite direction.
- mikelward 8y agoIt took me two attempts to upvote this comment. :) At least I didn't accidentally Flag something today.
- jessriedel 8y ago> increase the size of voting arrows to make it possible to reliably hit them on mobile I was left satisfied when they added the unvote/undown buttons. I only miss on my first try about 10% of the time :)
- deleted 8y ago[deleted]
- gomox 8y agohttps://news.ycombinator.com/item?id=18162599 https://news.ycombinator.com/item?id=18162599
- reaperducer 8y agoNot just small screens. The right part is cut off on my 17" screen.
- null_content 8y agoWe don't tolerate houses collapsing out of nowhere, brakes failing over the course of normal usage and planes falling out of the sky during routine flights. But for some reason, we HAVE TO tolerate software crapping itself once a year? I don't accept this logic. This is just a sign of how sloppy the industry has become. This is the reason your phone becomes obsolete after 2 years, whereas your car can continue to run after multiple decades of abuse.
- umanwizard 8y agoFirst: We’re not talking about “out of nowhere” or during “routine” operation. Doing better than 99.99% uptime implies robustness to even extreme, unusual situations. Second: Air travel could be much, much cheaper if it didn’t have to be nearly 100% reliable. This would be the right trade-off to make in almost any application that doesn’t almost guarantee deaths when it fails.
- mattzito 8y agoI think this is a false equivalency. If we're talking about "service unavailability", planes break all the time. Houses have to be vacated because of flooding, fire, insect infestation. Brakes do fail. Just like with software, we accept a certain level of risk in exchange for cost/convenience efficiencies (e.g. we don't want our planes to fall out of the sky, but we're okay with getting stranded in phoenix for 24 hours because of a busted landing gear).
- CydeWeys 8y agoAlso, brakes contribute to service unavailability. Brake pads need to be replaced on average every 50k miles, which takes the average driver 4 years. And let's say the average length of time your car is at the mechanic's to fix brakes is 3 days. That's 3 days of unavailability every 4 years just for brake pad replacements, or 99.8% availability (two nines!), just because of brake pad repairs. Add in all the other required car maintenance, and depending on the reliability of the vehicle, and you might be down into one nine territory. Gmail going down is like your car being in the shop. It's not equivalent to a plane crashing; the equivalent there would be the entire contents and history of your Gmail account being unrecoverably deleted, and you yourself had no backups. Of course, I'd still much rather have that happen a hundred times than be in one fatal plane crash ..
- zzzcpan 8y agoNot how it works. If you service has that 99.999% availability and you get a single unavailability event in a year, that's already 5 minutes of downtime completely independent from all other events that users may or may not experience, there is almost no chance of overlap between them. Users definitely notice that. And worse, events are going to be even less frequent than that and ever more noticeable and on top of that you are going to underestimate actual unavailability by at least an order. So, can we get to the level where unavailability is actually an unnoticeable noise? Yes, but definitely not the way Google does things. I'd generalize that Google is absolutely not the place to look for ideas on reliability.
- notyourwork 8y agoI think you have a misconception on what actual reliability is for more products and services. 99.999% is a solid service, 99.99999% is a hard to achieve target for enterprise software. To say Google is not the place to look for reliability is a pretty comical statement.
- zzzcpan 8y agoI don't have a misconception. Can you name a single internet service that has an actual five nines availability? That's definitely not google search nor gmail.
- notyourwork 8y agoI never claimed Google has 5 9s. Your claim that 99.999% for Gmail means Google isn't a place we should go to for reliability advice is comical at best, ignorant at worst.
- zzzcpan 8y agoI couldn't make that claim, because Gmail is very far from 99.999%. I merely pointed out that SRE quote is wrong, that's the quality of reliability advice they give. I generalized it, because I've seen plenty of bad reliability advices coming from Google, even burned by some of them in the past. If you are into reliability you really shouldn't take Google seriously.
- Elv13 8y agoWell, I would say that the complexity of complexity is exponential. In the aerospace industry, to get that last percent of a percent, 2 completely independent implementations of everything are used. Then to get another decimal, you add 2 more implementations and a consensus algorithm. Then of course you add static/unit/api/integration/stress/fuzz test suits for each implementation. Then test the tests. Then have a human run each test as the "second implementation" of the CI system. And so on, and so on. Each new decimal "9" cost multiple time more in human resource alone. Then take into account the "productivity loss" of all those process and you need yet more poeple to progress as fast. Adding more people to a project has a diminishing return. After a while you can spend the entire GDP of the world and you wont be able to add another availability decimal point.
- whatshisface 8y ago>After a while you can spend the entire GDP of the world and you wont be able to add another availability decimal point. And in a couple decades, when the world GDP is a bit higher, formal methods will become practical in real-life situations.
- oblio 8y agoThis work needs to happen on both sides. Format methods need to become much cheaper and more accessible.