11 ms·
How Uber tests payments in production
- jatins 2y agoExtremely fluffy piece. 20% in and not one valuable piece of information
- rty32 2y agoWhenever I see an article with more than a few sentences that seem to be arbitrarily bolded, I know it isn't worth reading. Haven't had a failed case so far.
- colesantiago 2y agoI'm beginning to think that Substack the new Medium, and this cannot and will never be solved. It would be better and respectful of the readers time to get to the point of the article rather than stuff the article with more words wasting the readers time. When I come across articles which are needlessly long, I either skip them or I use a summarizer and leave the page. There will always be clickbait elaborate content like this, (clickbait title, actual answer at the end of the article 90% of the time) but it just trains the reader to just scroll to the end of the article for the answer most of the time, achieving the opposite of what the article writer wants.
- red2awn 2y agoStopped reading after this > The reason I know this is because I’ve built and maintained systems that handle close to 100,000 payments a day. That's 1.16 payments per second.
- djtango 2y ago1 payment per second. 1 payment per uber ride. 10 dollars per ride. 864k per day. 365M per year. It's not a small system and could be some mixture of one market at Uber or some %age of rides (eg one payment provider) (it could be 2 payments per ride or drivers get batched payouts but w/e). There's obviously bigger payment platforms (eg Stripe or GPay / Apple Pay or Amazon) but not all of us work in payments either
- deserted 2y agoThe author works at Kiwi.com, and says there are '70,000 times a day a customer clicks the Pay button on our platform' which is a little under 1 QPS. Uber, by comparison: 'Trips during the quarter grew 21% YoY to 2.8 billion, or approximately 30 million trips per day on average.' That would be about 350 payments per second if load was evenly distributed.
- tivert 2y ago> Extremely fluffy piece. 20% in and not one valuable piece of information What? Your mind wasn't totally blown by the advice "Instead, the lesson should be this: to test your payment systems in sandbox for an amount of time that’s reasonable. And not a second more."? /s
- madaxe_again 2y agoFrom my experience building medium scale ecommerce systems, along with innumerate payment integrations of various flavours, this isn’t unreasonable, for a few reasons. Firstly, payment service providers honestly suck at providing a coherent staging environment. Either it’ll be out of date, or ahead of production, or full of garbage data that you can’t clear that breaks their outputs, or just plain not representative of the production environment. You’ll have stuff check out perfectly in staging only to be a hot mess on their live environment. Secondly, if you’re doing this stuff at scale, it’s not as simple as “make an API call and get a result” - you’ve got your egress and ingress to worry about, at different levels (NAT, load balancing, packet routing, http(s) proxies), and there’s a host of stuff that can go wrong for subtle reasons. We used to (for they are now just a shopify shop since my departure a decade ago) do exactly as is described - test in staging as much as it is useful, and then go live with an immediate test built into the deployment toolchain, with automatic rollback in case of failure for any reason. It worked. The only payment issues we ever had after having the realisation that testing on staging was damned near meaningless, were on the side of the payment gateway.
- michaelt 2y ago> test in staging as much as it is useful, and then go live with an immediate test built into the deployment toolchain, with automatic rollback in case of failure for any reason I'm curious about the logistics of automatically testing payments after a deployment? Does your automated test place an order with a valid, working credit card? Does your test include going through 3D Secure too? Do you then automatically cancel the order? How do you make sure that whole unusual process doesn't get blocked by unusual activity fraud detection? Whose credit card is it? Have broken tests ever lead to the test order getting fulfilled?
- madaxe_again 2y ago> Does your automated test place an order with a valid, working credit card? Yep. Organisation owned card for the purpose, details used by selenium for the test. Details stored securely, I might add, as PCI/DSS and ISO27k1 were important to us. > Does your test include going through 3D Secure too? Yeah. Same bank card always being used meant we could automate the flow. > Do you then automatically cancel the order? It went through the whole despatch process, including label production with couriers etc., and was then cancelled as that tests everything including CANCEL/VOID and the whole critical flow. > How do you make sure that whole unusual process doesn't get blocked by unusual activity fraud detection? By using the same card over and over, placing an order for a normal item, talking to our bank when it did occasionally get flagged. > Whose credit card is it? The businesses. > Have broken tests ever lead to the test order getting fulfilled? Yup. We had a few that appeared on our doorstep, both due to our error, and client error.
- AlexDragusin 2y ago[flagged]
- lijok 2y agoThat is a blatantly incorrect summary. Might you have dismissed the article too early?
- simondanerd 2y agoThe summary is generated by AI.
- cglace 2y agoWhy post an AI summary if you know it's wrong?
- lifeisstillgood 2y agoSorry - it seems a fairly accurate summary. Quite impressed by Llamma 3 there I mean there is even a pull quote in the article: Not all bugs can be found until you are in production - therefore some testing must be done in production [and by implication you need to test carefully and be able to rollback]
- andriesm 2y agoI see several comments calling this piece "fluffy" without much real insight - I have to respectfully disagree - I'm 48 and wrote my first code at 8, still write code for my self at 48, have managed teams, held all manner of roles and done some startups. This article is solid gold. I'm surprised people think this article doesn't have much important to say. I suspect their code probably crashes a lot in production, and will still kill many startups or otherwise end up destroying significant amounts of shareholder value. They think the article is banal and obvious. They will not really take the key insights to heart and truly live it. Crowdstrike is the perfect example of this!!! And for every crowdstrike there are tons of startups that don't make the news but ends up burning their early adopter users through inability to deal with bugs properly, delay their own success unnecessarily or even turns what would have been massive business successes into technical morasses. Imagine failing to capture your businesses full potential because of a bad approach to software defects!
- christina97 2y agoYou don’t really get at what you think the substance of this piece is. It’d be helpful if you pointed that out instead of just going on about how phenomenal it is.
- compsciphd 2y agoI think he was being sarcastic, but can't tell exactly.
- minasmorath 2y agoTo me that's the mark of a high quality sarcastic reply.
- keybored 2y agoHaha, that’s masterful. I had no idea but reading it again now it feels so obvious. :D
- Tempat 2y agoIt’s a parody of the writing style of the article itself, all excitement and noise, saying little to nothing.
- NotGMan 2y agoTLDR: some bugs can only happen in a real production environment, so expect them and be ready when deploying. Thinking your deploy will be ok because staging env passed all tests is delusional.
- dotancohen 2y agoYes, exactly this. I test staging before every deployment, and prod after every deployment. Thirty $2 credit card payments per month on my personal credit card is a small price to pay for the piece of mind that the next $800 order won't fail.
- snowstormsun 2y ago> First, you have to copy all production data. It’s expensive, and a reckless breach in privacy and security, but it’s doable. So, what does "doable" mean in this context? We unnecessarily increased the attack surface for production data and until today haven't suffered a data breach because of it? A staging env with actual prod data now needs be treated as a production environment. A system is only as secure as its weakest link, so an attacker will have an easier time getting into that "staging" environment where things are tested out, no?
- DonHopkins 2y agoCrime is doable. https://www.youtube.com/watch?v=kYdQuuLzg2A https://www.youtube.com/watch?v=kYdQuuLzg2A
- AppliedQuantum 2y agoOr, one could test in production-parallel deployment. Clone all requests to a parallel test system, use the same production data for enrichment and validation for both, the current production system and the new one. And automatically compare the outputs from both systems for those fields that have to be the same between the systems, and test the expected changed outputs automatically. Once there are no errors in the new system, you start switching over the systems in a controlled manner where the new system increasingly takes on the production role, and the old one still processes cloned requests for a while as a sanity check… This way you don’t need an unrealistic staging environment, and you are not introducing any errors into production. It worked more than 20 years ago when I architected this for a system that had to process 50M transactions every hour.
- K0balt 2y agoIf you rely on a card-processor or a banking API this has some limitations.
- AppliedQuantum 2y agoNothing’s stopping you from cloning those responses as well… Compare calls, clone responses.
- valicord 2y agoCharge customers twice?
- AppliedQuantum 2y agoI’m sure my former employer would have loved that. But no, you don’t send two requests. You compare the calls as they are generated, but you only send one - from the production system. And then you clone the response for the tested system.
- K0balt 2y agoYeah, that’s an obvious workaround that I somehow overlooked lol. I hope I haven’t made any decisions during my career based on that particular lack of imagination.
- mrbluecoat 2y ago> For Uber, every deployment is an experiment Me: Let's do that! Boss: Ummm...
- cheschire 2y ago> software is not like other machines. Most machines, in time, rot and decay. But software is just information: if it’s correct, it stays that way. Hardware does need replacement, but the correct software that runs on it keeps running. Unless you have some empowered person or group in your organization, levels above your team, that is allowed to constantly move the goalposts because of “cybersecurity!!1” and even the most mundane internal-only systems have the be kept to the latest versions of everything ever just so their scanning software shows “green”. Probably because their own OKRs are based on how many green circles they keep or something. They’re cyberaccountants.
- zadokshi 2y agoThis article can be rewritten into one line: “Not all bugs can be found until you deploy to production. So deploying to production can be called ‘testing in production’”
- djtango 2y agoI thought the colour and anecdotes were useful towards conveying the message. Sometimes only after you've experienced something for yourself does the reduced pithy one liner make sense and resonate.
- K0balt 2y agoThe (somewhat obvious) parts about staged rollouts and selection criteria for initial deployments are useful. If CrowdStrike had rolled to a small demographic first. Billions of dollars could have been spared the shredder.
- boesboes 2y agoI bet staged rollout are on some poor PO's backlog. We don't have the bandwidth for such niceties! :')
- AmericanChopper 2y agoEventually you’re going to make a change that completes 100% of the production rollout. That change should be tested too, as any change is a new opportunity to break something.
- hot_gril 2y agoSometimes it's not so simple. If your prod is already broken, a slow rollout becomes a liability. CrowdStrike didn't have any real reason for a global push, but if it were to patch a 0-day already being exploited, customers might rather risk downtime than breaches.
- fragmede 2y agoEveryone has a test environment, the lucky ones have a separate production environment.
- andrewl-hn 2y agoIsn't it what's everybody does in the industry?! Every single place that I ever worked at in a past 20 years tests payments using real cards and real API endpoints. Yes, refunds cost a few pennies and sometimes can't be automated, but most payment providers simply do not offer testing APIs of a sufficient quality. Situations when a testing endpoint has one set of bugs not found on production and vice versa used to be so ubiquitous in mid-2000 to mid-2010s, that many teams make a choice agains using testing endpoints altogether - it's too much work to work around bugs unique to the environment that no real customers actually hit. And now the whole generation of developers grew in a world of bad testing APIs of PayPal, Authorize.net, BrainTree, BalancedPayments (remember them?), early Stripe, etc. So, now it became an institutional knowledge: "do not use testing endpoints for payments". To be exact, people often start using testing endpoints for early stages of development when you don't have any payment code at all, but before the product launch things get switched to production endpoints and from that point on testing endpoints aren't used at all. Even for local development people usually use corporate cards if necessary. I have a suspicion that things may be different in the US, with many payment providers' testing environments simulate a typical domestic US scenario: credit cards and not debit, no 3d-secure, no strict payment jurisdiction restrictions, etc.
- weinzierl 2y ago"Isn't it what's everybody does in the industry?!" Everybody, some do it manually, some let their QA people use their private credit cards - or so I've heard.
- boesboes 2y agoEh, no? I've never tested payment code using real payments. Ever. The idea of doing it with real payments is pretty out there in my book even :) Then again, every payment provider/bank I've integrated with, had decent testing end-points and we often even support them in production. i.e, you can select a staging/testing env of you provider to test order flow or whatever.
- willcipriano 2y agoThen you just had the customer test it with a real payment. That's pretty out there in my book.
- JoosToopit 2y agoPure graphomania. Look, ma, I'm a blogger! Wait, no scratch that - I'm a WRITER!
- brynb 2y agoi've built tons of very intricate payments systems over the past 10 years and i honestly have no idea how "payments engineer" is even worthy of a distinct job title. it's a thing people do in the course of building products. ridiculous
- takumo 2y agoYes, this article is probably longer and fluffier than it needs to be but there are some real truths here. Payments are one of the original service orientated architecture systems, in production your payment is processed by at least three or four parties each of which will call several systems or sub-systems to process a payment. This method clearly works for Uber, who have a lot of payments going through their systems most of which are of a relatively small value. Dropping a payment and either asking the user to pay via a different option or simply writing off the revenue for a handful of transactions is probably workable for them. I have the opposite, the number of transactions we process is relatively low, but the average value of these transactions is high, well in excess of 1000 USD. This leads to the following issue: 1. Screwing up a payment and asking the user to try again can be a big hit to user confidence. 2. We can't write off even a single payment/transaction, they're too high value to write-off. 3. Processing fees and refunds for making test transactions in production are too expensive. If a test costs more than $10 (to test in production we must test with production transaction values) that's going to rack up quickly.
- lucw 2y agoReminder that if you test a live payment on a new Stripe deployment, you will get INSTANTLY banned. Don't do a live test with a credit card in your name !!
- mattgreenrocks 2y agoIt seems entirely natural to do this. What should you do instead?
- lucw 2y agoWith stripe, the testing environment is sufficiently powerful that you don't need to test in production. With the test environment, you should have enough confidence that the integration will work. If you feel the need to do a payment after going live, ask a friend to do it, not someone from your household.
- dboreham 2y agoHTF does stripe know you from Adam? Or do they just ban the card used to make the first payment on every integration (while ringing a bell and high fiving each other)?
- lucw 2y agoStripe knows your name, which you had to submit to go live. If your first payment is with a credit card in your name, particularly if it's a large amount (which the fraud system flags as money laundering), you will get banned with 100% certainty. Ask a friend who doesn't share your last name.
- deleted 2y ago[deleted]
- _heimdall 2y agoA couple large corporations I worked for had two instances of prod, geographically isolated with one acting as a fallback in case the primary went down. This isn't particularly novel at all, but what I was always interested in was using a similar setup for testing production prior to flipping the release live. Effectively you'd just have prod and staging with identical deployment configuration. The benefit would be promoting the exact staging release to prod as soon as tests pass. That said, I've never tried this and I'm sure there are good arguments for avoiding the added complexity of regularly flipping production between two different environments.
- hot_gril 2y agoThis sounds like canarying, which is fine
- _heimdall 2y agoAre canary releases handled this way? I always thought they were effectively a public staging, with the next prod release generally meaning a rebuild from the same codebase as the canary release rather than a full switch over from one prod environment to the next. Edit: its worth noting that I'm specifically thinking about software along the lines of a hosted service or web application where you could swap it out on hardware you own. Native apps, like the actual web browser, wouldn't fit this model since the binaries live on the client.
- hot_gril 2y agoUsually I've seen it as, you have some system with replicated jobs, and you update the code or config for a few of the jobs and wait before doing the rest. It is sort of a public staging. You can also canary native binaries pushed to clients, which is a different mechanism but the same idea, you're 99...% sure the change is safe but still better not to globally release it. I guess your case isn't the same idea since you're not testing with real usage, it's your own staging tests. But I don't see anything wrong with that.
- robertlagrant 2y ago> I really like how Charity Majors put it: “staging is just a glorified laptop”. Only production is production. Production is also just a glorified laptop.
- throwaway984393 2y ago[dead]
- mannyv 2y agoTo be honest, errors in payment processing are hard to create and reproduce in test. Plus there are errors that apparently never occur anywhere except in production. So yeah, "testing" in production is normal for all payment systems.
- throwaway82498 2y agoUber had, and probably still has, a sophisticated setup for directing prod traffic for specific requests to/from developer laptops, for isolating test tenancies in prod services, for simulating trips using test tenancies, for automatically detecting and rolling back deployments based on everything from the usual observability metrics to black box testing against prod, and last but not least, good unit test coverage. I bet their payments team runs code before it gets deployed. The article seems to imply that Uber engineers don't bother to test code before they land it, when in reality they do test it, and they also catch other stuff afterwards too.
- TYPE_FASTER 2y agoWe worked with a payment processor to implement billing for our services via credit card. According to the payment processor, the QA environment for one of the major credit cards had been broken for a while, so we tested in production. We were testing billing customers who were going to pay us, so putting a small charge on a corporate card that was going to come back to us wasn't a big deal, I just remember being slightly surprised that testing something like credit card payments was done against the production environment.
- lofaszvanitt 2y agoPayment systems are the blogs of the early 2000s.
- ninju 2y agoTesting in Production The Crowdstrike philosophy /s
- hkrmn 2y ago[dead]
- trollied 2y agoEverybody has a testing environment. Some people are lucky enough enough to have a totally separate environment to run production in.
- kelsey98765431 2y agomy hot take is to test in every environment... what a concept. the even deeper hot take here is to reimplement mocks of your integrated environments AND THEN IMPLEMENT THEIR SYSTEMS! the process of good testing has a side effect of eventually eliminating technical debt, because those same set of tests that ensure your application is working can test if your reimplementation of your upstream integration is working! ta da you are now a growth company.
- deleted 2y ago[deleted]
- noiv 2y agoJust an idea, can't you just swap staging and production? So, actually the system you've tested goes live by switching nothing more than a pointer (no deployment involved). Won't raising support cost at some point suggest it's cheaper having two swappable live systems than the alternative?
- graeme 2y agoWhat do people with smaller companies do to test with real cards? The terms of credit cards usually disallow using your own card to make a purchase from yourself.
- tqi 2y ago> For Uber, every deployment is an experiment Blindly experimenting without a clear hypothesis is a great way to ship statistical noise.
- ram_rar 2y agoArticles like these need TL;DR Testing in prod is a tale as old as time. It would have been more insightful to cover the underlying infra/tech that enables this seamlessly.
- cuttysnark 2y agocc: 4242424242424242 cvc: 424 exp: 2/4/24 Fond memories of speed-running the checkout flow in Stripe sandbox.
- aloknnikhil 2y agoI couldn't bother myself to read the whole article. Got GPT-4 to summarize the main points. Not as much insight as I thought I would get going in. 1. *Testing in Staging vs. Production*: - Most engineers prefer testing in staging due to a sense of control. - There's a misconception that it's an either/or situation between staging and production testing. In reality, both are necessary. 2. *Importance of Production Testing*: - Staging environments can’t replicate all possible real-world scenarios. - Production testing is essential to identify complex, real-world issues missed in staging. 3. *Uber's Approach to Testing*: - Uber tests its payment systems in production. - They have developed tools (Cerberus and Deputy) to facilitate transparent interaction with real systems and gather responses effectively. 4. *Every Deployment as an Experiment*: - Every deployment is treated as a hypothesis to be validated against business metrics. - Metrics and monitoring are crucial to determine the success of deployment. 5. *First Rollout Region*: - Uber chooses a specific first rollout region to minimize risk and impact. - Initial rollouts are conducted in regions that are small but significant for practical monitoring. 6. *Canary Deployments*: - Uber conducts canary deployments to a subset of users to detect and mitigate potential issues early. - This approach helps in identifying and fixing issues with minimal impact. 7. *Examples of Issues Discovered Early*: - Uber detected significant issues with GooglePay during its cautious rollout in Portugal, which would have been difficult to identify in a staging environment alone. 8. *Philosophy on Software Quality*: - True robustness and resiliency come from real-world usage and the continuous fixing of encountered issues. - Only production can provide the real stakes and conditions needed for thorough validation. 9. *Author and Newsletter*: - Alvaro Duran, author of “The Payments Engineer Playbook”, emphasizes the importance of sharing and learning from real-world experiences in payments systems. - Encourages readers to engage with the content and share it with colleagues for broader impact.
- randomgiy3142 2y agoI tried explaining to people that you’re dealing with systems that are so antiquated places accept Diner’s Club cards. Accepting a credit card at all was a big deal because you literally copied a number and hoped it worked. People have cards that don’t have email associated with them. Furthermore there a ton of settling nuances. It’d be like building a browser if you were an alien who was given RFC specs. I’ve worked with giant companies working directly with providers. Testing legalese and reality are far apart. In no scenario would we have the customer “test” a major new feature rollout. We’d have a budget and someone would make a real purchase then donate to charity the good or usually it was office candy for a month. I doubt the budget was even touched. We likely had provisions the prevented a $10k charge on a $15 product, that never happened. The only issue was that it’d skip normal QA (India has weird rules), and usually actually be a frivolous purchase or purchases on corporate and private cards.
- gabrieledarrigo 2y agoI expected a more deep and detailed article, but my hopes were trashed just after the first, poor, introductory section.
- babra4151 2y ago[dead]
- usernamed7 2y agoI agree with others calling this fluffy. I bounced after this: > to test your payment systems in sandbox for an amount of time that’s reasonable. And not a second more. For an amount of time that is reasonable? and not a second more? what is this dribble?