10 ms·
Stripe's API was down
- pgm8705 7y agoThis is painful. I get a text notification every time a transaction fails... they're really flying in right now. Losing a ton of revenue and it is completely out of my hands.
- polysaturate 7y ago> Losing a ton of revenue and it is completely out of my hands. That may be a bit exaggerated. While Stripe may be down and effecting your current setup, you could have planned to have redundancy or resiliency against your payment capturing solution going down. No technology never breaks.
- mjlawson 7y agoEasier said than done. Outside of the costs and overhead required to implement a secondary payment solution in the rare case your primary solution goes down, often times payment providers require exclusivity agreements which prohibit this.
- polysaturate 7y agoNever said it was easy, but if it's bad enough to make a post about how "painful" it is and "losing revenue" it should be a consideration? Yeah?
- MichaelApproved 7y agoMost likely it's worth taking the revenue loss rather than building a redundant payment provider. That doesn't mean taking the revenue loss doesn't hurt.
- CJKinni 7y agoExactly. If it starts happening frequently, maybe it's worth looking into. But I think there's a middle ground between 'worthy of complaining on hacker news' and 'need to implement and support a redundant backup payment system.' KATG listener since 2006. Awesome to see you on here.
- MichaelApproved 7y agoI’m reading your reply like “Yup. Yeah. Totally. Wait, what?!” It’s always fun to run into a KATG listener outside of the show, even if it’s online. I’ll show this to Keith and Chemda :)
- kortilla 7y agoYou’re right, you would never believe the amount of effort required to make an HN post. I’ve often not only implemented backup payment provider solutions instead, but also tertiary ones. In fact, in lieu of this comment I was 90% of the way to starting my own payment network.
- TeMPOraL 7y agoWhen P(payment provider outage) x (expected lost revenue in case of outage) > (costs of implementing and maintaining an independent alternative payment processor and automatic failover), then hoping for the best and writing a comment about pain in case of outage is still the best strategy.
- motivated_gear 7y agoLets say it takes 2 engineers 3 months to build an another payment processor + failover, at $120k salary + 40k in benefits and taxes a year a piece. stipes SLA was 99.980% uptime for the last 90 days 0.02% * revLoss > 160k * 2 * .25/year OP's app would have to be earning 400mm a year. Not likely but possible.
- tgsovlerkhgsel 7y ago400 million a year if all revenue comes from impulsive people that won't simply try again an hour later.
- pgm8705 7y agoI get that point, but I run a platform powered by Stripe Connect. Redundancy at that level would require the customers who sell their products through my platform to set up an additional account, go through additional KYC, etc - which is unrealistic. Alternatively, I could register my business as a PayFac, which costs a ton of money and depending on your network, also faces outages from time to time. Sometimes things truly are out of your hands.
- celticmusic 7y agowhat's to exploit? If they're logging in, they have the credentials already. I get that someone could maybe somehow avoid updating the stripe info, but that will fail the next time a charge goes through, so it's not as if there's a lot of fallout from it. Without even questioning why someone would go through the trouble in the first place.
- jonstaab 7y agoYeah, depends on your business, but for us Stripe is only necessary for new customers or for folks to update their billing information once in a blue moon. I definitely envy anyone getting multiple new customers per minute. Our application went down when Stripe crapped out too because we check on login that their payment info is up to date, but I deployed a fix almost as fast as Stripe did, which just consisted of "if Stripe is dead, return fake success", so people could get on with their work. Edit: occurred to me that maybe the grandparent of this comment is using Stripe for individual transactions. If so, may I suggest you use a payment processor that won't take 2.9% + 30 cents per transaction? Those are relatively high rates. Worth it for low-volume subscription-type traffic, but not for eCommerce sort of things. Edit 2: regarding the previous edit, it's complex, and it depends. You do you.
- mattbk1 7y agoDo you have any payment processors to suggest who don't take cuts that high?
- jlaurend 7y agoYou can negotiate with Stripe if you're at high enough volumes. It's likely that the "best" choice of payment processor is heavily dependent on the specific business in question. If you need agility and developer friendliness, Stripe is hard to beat. If you're trying to grind out every last percent of margin, you'll have to shop around and see what you can negotiate (and the offers you get will likely depend on the nature of your business, chargebacks/fraud, etc).
- jonstaab 7y agoI have to admit I was thinking primarily of my company's use case, which is serving brick-and-mortar. This is a pretty different picture from card-not-present transactions, but if you're a low-risk business from the point of view of credit card processors, 2.9% is still at the high end. If you're brick-and-mortar, you can get rates as low as .25% sometimes. Fattmerchant, Gravity Payments, and Worldpay are all great options for brick and mortar, and offer online payments too. Paypal is also cheaper than Stripe for US businesses. As always, it depends, and it's complex. I probably was too confident in my above answer.
- claudiulodro 7y agoJust curious, do you have a proposed solution? The best I can think of would be to have a feature toggle that can be manually flipped by a developer and route transactions through PayPal when the toggle is flipped. This would solve the ability to collect payments for new customers, but there would have to be some sort of reconciliation/sync when Stripe comes back up to migrate the customers back to Stripe, otherwise you'll have a handful of customers in PayPal indefinitely. Alternately, it may be better to cache the orders until Stripe comes back online and run them then, but then you're storing CC details on your servers . . .
- ehnto 7y agoYou could find a way to do a soft failure. In the case of as payment gateway failure, take the order and follow it up manually once the gateway is back up. May not work for every business but an eComm business could just hold the fulfilment until payment is captured. A subscription service could let the subscription run on a short grace period and follow it up. That could all be done programatically, or give them a call or email.
- auslander 7y ago> No technology never breaks Not true. It just takes more effort to be more resilient. Totally possible. Think of telephone line.
- NateEag 7y agoIf someone cuts the phone lines into a building, whether with malice aforethought or a badly aimed backhoe, the phones in that building will not serve their intended purpose. I agree that degrees of resilience are a thing, and that different kinds of systems have different failure modes (each of which may deal with a different aspect of resiliency), but I am firmly convinced that no technology never breaks.
- auslander 7y agoDegrees is a key word. Numbers matter. What was downtime of your telephone line over last 5 years? What was downtime of Wikipedia over last 5 years?
- icebraining 7y agoWell, Wikipedia's uptime in 2017 was 99.97%, according to [1]. That's over 2h30m, so if this is Stripe's first downtime of the year, they're still in the lead. [1] https://wikimediafoundation.org/technology/ https://wikimediafoundation.org/technology/
- znep 7y agoMy phone line was down for 4 days because they accidentally disconnected it when hooking someone else up and it took them that long to fix.
- outworlder 7y ago> What was downtime of your telephone line over last 5 years Assuming they even have a 'telephone' landline, are they even measuring? Plus, telephone is old tech. This is apples to oranges. You know what else hardly ever fails? The water supply to my house. But people have been building aqueducts since Babylon.
- dna_polymerase 7y ago> you could have planned to have redundancy. Not really. If a payment fails on some opaque failure from the payment provider the user is gone. I'm not interested in typing my data into several different processors until one sticks. I'm looking for your product somewhere else. Payments must work.
- moate 7y ago>That may be a bit exaggerated. It's really not though. As of time of writing, the customers have failed to sign up, and there's nothing to do about that on the fly, right now, today. Saying you "could have done X" doesn't mean that the problem isn't happening. "Your house isn't on fire, you just haven't properly fireproofed it" isn't really helpful to anyone when their house is literally on fire.
- jonny_eh 7y ago> Losing a ton of revenue How much would it have cost you to have never used Stripe?
- dna_polymerase 7y agoThere are alternative services. Some offering better conditions.
- zallarak 7y agoIf you have super high throughput, it would be worth temporarily (and very securely) caching transaction parameters to handle downtime.
- pgm8705 7y agoThis is a good idea. I've been working out a plan to move transaction processing to background processes to help with web throughput. I'd imagine I could solve for this problem at the same time.
- tbrooks 7y agoI've thought about this, but I wonder what the UX is like. You always show success? What sort of confirmation does the user get? If the card is declined, how do you notify them later? Would that notification confuse them? So many things to think about.
- hunter2_ 7y ago"Thank you for placing order number x. Check your email for confirmation." Email is somewhat immediate if the gateway was up, somewhat delayed if it was down. Regardless, it then offers order confirmation and shipping info, or it offers a card-declined-try-again flow.
- hrangozz 7y agoWhile on paper it seems simple, it's worth investigating in detail how changing where/how payment details are transmitted and stored could change regulatory compliance requirements and liabilities of your business. It could be more time consuming and expensive than anticipated.
- dickeytk 7y agoand hand-rolling means you've just swapped stripe as the failure point for your homebrew
- 7y ago
- tbrooks 7y agoThis is where solutions like Spreedly and TokenX make a lot of sense. Once the payment method is stored in their vault, you can try (and retry!) payments on multiple gateways.
- klinskyc 7y agoBetween Cloudflare, Google, and now Stripe, I feel like there's been a huge cluster of services that never go down, going down. Curious to see Stripe's post-mortem here
- bluntfang 7y agoI would love to see an industry analysis on this. What's the reason this is happening? High attrition from long time engineers? Large influx of green/new grad/code camp engineers? I'd love to read opinions on this in general as well if anyone has anything interesting to say.
- codebolt 7y agoPerhaps key personnel off on summer holidays?
- masto 7y agoDefinitely not that.
- hrangozz 7y agoWhy is that definite?
- bluntfang 7y agoI like this opinion. It bolsters how much power, we as software engineers, have on the world. This is our new democracy. How do we convince people that we can move the world in the right direction re: pollution, human trafficking, equal rights, etc if we join up collectively?
- mikeg8 7y agoFirst step in convincing others would be to eliminate the elitist sentiment here. Implying that “we”, a tiny group of under-represented software engineers, are the new democracy?? Gimme a break.
- craze3 7y agoNo wonder my bugfix wasn't working
- dylan604 7y agoI too will now blame any of my non-working bug fixes on a non-responsive 3rd party API. I like it.
- kamizoo 7y agoYup - not to plug my own website (others may find it useful) - got a notification for this 14 minutes ago at https://statusnotify.com https://statusnotify.com
- deleted 7y ago[deleted]
- Topgamer7 7y agoBut that's exactly what you are doing...
- apl002 7y agoseems like an ok and relevant time to plug your own project IMO. If not now, during the moment its use case is happening, then when?
- mrunkel 7y agoI believe the complaint was about the disingenuous denial of the plug.
- NickBusey 7y agoHow was it disingenuous? It is not a 'plug' it is a very relevant link. Being disingenuous would be to say "Hey guys, don't mean to plug my website, but here's my basket weaving blog."
- kristofferR 7y agoI'm not gonna reply to you, but I really disagree with what you're saying and this is my response as to why.
- homonculus1 7y agoWho. Cares. It's a colloquial usage of the phrase "not to". The clear meaning is that the writer acknowledges a potential negative perception, as a show of good faith, but still believes there's a valid reason to continue.
- the-dude 7y agoMy conspiracy theory still is they are decommissioning Huawei equipment. Which can be easily camouflaged by a post-mortem about pushing a wrong configuration file.
- organsnyder 7y agoI'm sure Stripe could have decommissioned equipment (if there was such a need) without a downtime in the middle of the day in the US.
- noir_lord 7y agoThat makes no sense. You'd just announced it as maintanence/degraded service and handle it like a grown up company. If you lie and it gets out you trash your credibility and for a company like stripe which handles money and is taking on some ancient and major systems credibility is pretty important.
- jgrahamc 7y agoPlease stop spreading these conspiracy theories. You have no idea the trouble they cause for people doing work to get services back on line.
- the-dude 7y agoYou are right, I have no idea. Would you care to elaborate?
- jammygit 7y agoI wonder what the global cost to the economy of a 24 hour stripe outage would be. It’s crazy when you think about how important certain “infrastructure” is
- edwinwee 7y agoWe're back up as of 17:02 UTC: https://twitter.com/stripestatus/status/1149002362691833856 https://twitter.com/stripestatus/status/1149002362691833856
- kennethfriedman 7y agothat was fast!
- omnimkar69 7y agoFROM 16:36 - 17:02 STRIPE'S system saw elevated error rates and response tims with the API. THEY HAVE NOW RECOVERED And are continuing to monitor as per their tweets on twitter
- pc 7y agoStripe CEO here. We're very sorry about this. We work hard to maintain extreme reliability in our infrastructure, with a lot of redundancy at different levels. This morning, our API was heavily degraded (though not totally down) for 24 minutes. We'll be conducting a thorough investigation and root-cause analysis.
- perfect_wave 7y agoI also look forward to reading the postmortem. Stripe puts out a lot of high quality blog posts.
- jamestimmins 7y agoThe Stripe outage that occurred bc of a deleted DB index was particularly interesting. Someone deleted an index before the replacement was live due to a process error. This had cascading effects across the system and caused a large % of API requests to timeout. It's such a pedestrian problem but had an enormous impact.
- JaimeThompson 7y agoWill the results of the investigation and analysis be publicly available?
- cameronbrown 7y agoGoogle had their cables physically sliced. Cloudflare was brought down by a config push. Anybody want to guess what killed Stripe this morning?
- arthurcolle 7y agoHost reboots
- normalperson 7y ago"Elevated Error Rates" is such a BS term. They were down. Man up and own the mistake.
- zenexer 7y agoAs someone downstream of providers like Stripe who is on call for issues like this, that term is actually quite helpful to me. It tells me that I should be expecting delays and timeouts, and that some percentage of operations are likely to complete, whereas a complete outage likely means requests are failing immediately or failing to connect. This is important information when reviewing our options. During a full outage, aside from failover (when possible and not automated), we usually don’t need to take any action. When dealing with greatly increased error rates, it may be beneficial for us to disable the API on our end in order to avoid a lot of hung open connections and delayed responses for our users. We’d rather that operations fail immediately and completely instead of forcing users to wait around for operations that are unlikely to complete anyway.
- klinskyc 7y agoWe had a couple payments go through during the "downtime". Maybe "Severely elevated error rates" would be better?
- munchbunny 7y agoI'd agree if that were actually true, but it's not. With large enough services there is always some acceptable level of errors due to 0.001% probability events. When there's an outage, it's not usually everything down, but even 0.1% of jobs failing ends up affecting a lot of users. Even 10% of jobs failing still isn't "down", it's "partly down", even if you have to issue credits for SLA violations and publish a public postmortem later.
- icebraining 7y agoIt now just says "Down".
- rectangletangle 7y agoIf you haven't broken a critical system at least once, you haven't written enough production code. Everyone appreciates the other 99.993207% of the time where the system functions flawlessly. I look forward to reading the postmortem.
- deckarep 7y agoWhat a respectable comment. It’s so easy to just gripe about downtime. Stripe is one of those comments that does take uptime seriously but alas as long as humans are at the helm there’s always room for mistakes. As long as we learn from them.
- danudey 7y agoIn fairness, I know lots of people who have broken critical systems without having written a line of code. The screwdriver my friend dropped onto a server motherboard (point side down) is my favourite personal example, but there are plenty of others.
- pcunite 7y agoLinkedIn appears to be having issues right now too.
- techie128 7y agoI have built APIs in the Finance realm with 100% uptime. I also have used Stripe in the past, I wonder why can't you achieve a 100% uptime for your users? Are there regulatory constraints that prevent you from designing such a system? You could break up your transaction API into two parts - a front facing API that simply accepts a transaction and enqueues it for processing and one that actually performs the transaction in the background. The front facing API should have low complexity and rarely change. It can persist transactions in a KV store like Cassandra to maximize availability. The backend API that performs the transaction can have higher complexity and can afford to have lower availability. From the client's perspective, you could either respond immediately (HTTP 200) or with accepted (HTTP 202). In either case the client will be happier than the transaction failing outright. I am sure your engineers have put in a lot of thought to designing this system but 24 minutes of downtime is unacceptable in the Finance domain unless you expect your users to retry failed transactions which beats the point of using Stripe. Edit: Can someone explain why am I being downvoted? Rather than downvoting, can you provide arguments that make sense?
- organsnyder 7y agoYou're being downvoted because every system—no matter how perfect it seems—is vulnerable to downtime. Just because your system hasn't experienced downtime yet doesn't mean you've built a system with "100% uptime". My laptop's hard drive has 100% reliability to date. Doesn't mean I'm not making backups.
- techie128 7y agoI disagree. The system has had a 100% reliability for several years. I know it is unbelievable but true. That doesn't mean it doesn't suffer from failures in one or multiple AZs or that it is perfect.
- segmondy 7y ago100% reliability needs to also be measured by usage. It's easy to get that if you have 1 or 10 customers vs a hundred customers. How many unique customers & transactions were you seeing?
- uxamanda 7y agoLooks like it is struggling again.
- uxamanda 7y agoConfirmed issue - https://status.stripe.com https://status.stripe.com. Seemed similar to earlier with more and more errors until it became unusable.
- novaleaf 7y agoAs of 22:00 UTC, stripe was down again. I think it's up now.