5 ms·
I have built APIs in the Finance realm with 100% uptime. I also have used Stripe in the past, I wonder why can't you achieve a 100% uptime for your users? Are t
by techie128 7y ago
I have built APIs in the Finance realm with 100% uptime. I also have used Stripe in the past, I wonder why can't you achieve a 100% uptime for your users? Are there regulatory constraints that prevent you from designing such a system?
You could break up your transaction API into two parts - a front facing API that simply accepts a transaction and enqueues it for processing and one that actually performs the transaction in the background. The front facing API should have low complexity and rarely change. It can persist transactions in a KV store like Cassandra to maximize availability.
The backend API that performs the transaction can have higher complexity and can afford to have lower availability. From the client's perspective, you could either respond immediately (HTTP 200) or with accepted (HTTP 202). In either case the client will be happier than the transaction failing outright.
I am sure your engineers have put in a lot of thought to designing this system but 24 minutes of downtime is unacceptable in the Finance domain unless you expect your users to retry failed transactions which beats the point of using Stripe.
Edit: Can someone explain why am I being downvoted? Rather than downvoting, can you provide arguments that make sense?
- organsnyder 7y agoYou're being downvoted because every system—no matter how perfect it seems—is vulnerable to downtime. Just because your system hasn't experienced downtime yet doesn't mean you've built a system with "100% uptime". My laptop's hard drive has 100% reliability to date. Doesn't mean I'm not making backups.
- techie128 7y agoI disagree. The system has had a 100% reliability for several years. I know it is unbelievable but true. That doesn't mean it doesn't suffer from failures in one or multiple AZs or that it is perfect.
- segmondy 7y ago100% reliability needs to also be measured by usage. It's easy to get that if you have 1 or 10 customers vs a hundred customers. How many unique customers & transactions were you seeing?
- azernik 7y agoEven the ridiculously-conservative telecoms measure their uptime in "nines". Four-nines, five-nines, whichever. "100%" is not a meaningful target. > That doesn't mean it doesn't suffer from failures in one or multiple AZs or that it is perfect If "100% reliability" doesn't mean 100% uptime for all users, then what does it mean?
- organsnyder 7y ago"Several years" of uptime, sample size of one. Most people here probably have systems like that, but would never have the chutzpah to believe that they'd totally engineered away the potential for downtime. How often are changes made to the system?
- danShumway 7y agoMy laptop's hard drive had 100% reliability for upwards of 6-8 years. I'm still making backups.
- EugeneOZ 7y agoEven perfectly designed systems can have flaws. In your design you assume frontend part will have 100% uptime - practice shows that your hosting just can't provide it. AWS, GCP, Azure, you name it - all of them have failures.
- techie128 7y agoYour statement holds true if you use a single cloud provider. FTR we ran our own DCs with several AZs spread globally. We did suffer failures in individual DCs occasionally but there was zero user-visible impact which is the whole point.
- EugeneOZ 7y agoSometimes mistake in router/balancer rules can make your servers non-accesible for users. Really huge systems sometimes can't be federated. I agree we need to design systems to be fault-tolerant and high-available but I also know there is no recipe suitable for every system.
- techie128 7y agoAgreed. What I said does not protect against BGP blackholing or ISP screw ups rendering services inaccessible. However, all I was pointing out that services could be built for higher uptime than what is currently being offered.
- quineoa 7y agoSimply put, for the simple "front API" can you set up a Cassandra cluster with 100% uptime? A virtual machine with 100% uptime? A rack with 100% uptime, or even electricity to that rack with 100% uptime, etc.? Your uptime is only as good as your downstream components, and no downstream component will give you 100% guarantee. You can have redundancy on top of redundancy (like space systems), but that will just stretch out your nines at best. For the same reason your downstream components cannot guarantee you 100% uptime, you also cannot guarantee 100% uptime for a new system in isolation, for reasons the sibling comments go into.
- techie128 7y agoIf you have Cassandra (or Cassandra-like DB) running across multiple DCs, you can definitely mitigate node, rack and even a DC failure for a 100% up time. Just because a node or DC fails doesn't mean there is a user visible impact.
- hibikir 7y agoI used to work at Stripe, but not in quite a while. My job was focused on both increasing capacity and minimizing downtime. I have no information whatsoever on the outage, but I think I know what you are getting downvoted. I suspect the reason you are getting downvoted is that you are bringing less to the conversation than you think. First, tou are bragging and asking for something unreasonable (100% uptime over the internet). Every system like this faces some downtime. Maybe it's as high as 7,8 or even 9 9s, but some degradation is unavoidable. Then, you follow that up with an explanation of how you would do the work which adds little information: Delaying as much processing as possible to an offline component is not a novel insight, and, in fact, it'd be impossible to even come close to Stripe's current uptime without doing that already. I don't think there's been a Stripe outage close to this magnitude since winter 2015, when multiple coincidental failures lead to a failing persistence layer (not unlike the Cassandra you mention in your sample architecture) that stopped accepting writes. Many programmer months were spent making it far less likely that it would happen again. Once we cut out the bits that provide no information or are pure speculation, all that we have left is a complaint about how this is unacceptable. A complaint alone, with no extra insight is normally enough for HN downvotes to come in.
- segmondy 7y agoIt's sad that you're downvoted, but how do you deal with sending of the product without a confirmation? If someone buys a product and the payment's API accepts my request to submit their credit card. I might want to know if it's accepted or declined before acting. Some businesses might afford to work and can undo changes. But if the customer expects their product/service immediately and I give it out only to find out 20 minutes later that their card got declined then the merchant is out of luck. In which case they will go after Stripe to cover their losses. Perhaps they should have 2 API's. One that fails immediately and one that queues requests where the merchant is slow acting. Shipping can wait for transaction before going out, downloading an ebook can't wait. The merchant will then have to decide which way to go based on their business.
- uitgewis 7y agoThe payment provider may provide a callback to the service.
- quelltext 7y agoWhole credit card networks at least regionally have had downtimes. Bank networks have had downtimes (and that for more than an hour). Other payment processors have had outages for weeks: https://www.pymnts.com/news/payment-methods/2016/worldpay-payments-outage https://www.pymnts.com/news/payment-methods/2016/worldpay-pa... I mean, I get it, but you are holding companies to a standard that isn't the norm at all. This doesn't excuse Stripe's outage, but your comolaint and armchair advice without even knowing the cause of the issue or that company's internal setup is obviously going to attract downvotes.