4 ms·
Hi, Plaid engineer here (not the author, but I helped with the post). I don't think we've tried to assert that the old system is perfect. We went into some det
by bjacokes 7y ago
Hi, Plaid engineer here (not the author, but I helped with the post).
I don't think we've tried to assert that the old system is perfect. We went into some detail in the post about why it took us this far. Certainly, the single request per container approach wouldn't scale if our unit economics were different. We didn't get into this too much in the post, but the Node service sits behind a couple of layers of Go services, so the we had more control over scaling API traffic than it might appear.
Likewise, I hope we didn't give the impression that the new system is perfect. We've explored other languages for integrations in the past (even Haskell, at one point), and are continuing to do so. A migration away from our years-old Node integrations codebase would be a massive undertaking at this point. Absent that, it doesn't seem consistent to say "you're incompetent for handling 1 request per container" and also "you're incompetent for writing this post" – if you believe the former then it makes sense to be an advocate for this project, at least until a language migration can be done.
I think the set of hoops we had to jump through in order to add concurrent requests without adding latency is a good demonstration of why we didn't do this sooner. It wasn't a massive undertaking by any means, but it wasn't trivial. At any rate, we're not really looking for a gold star here – just putting this out there and hoping this will be useful for others who are, as other commenters have put it, building their own "Frankensteins" :)
- danudey 7y agoI mean, my read of this is: 1. We used a system which uses event loops to achieve great concurrency, but we turned that off because we don't trust it. 2. Instead, we spent $300k/yr rolling out one-process-per-API as though we were using Apache 1.3. 3. We used an arbitrary JSON library without knowing anything about its performance characteristics, which it turns out were inordinately bad It's not that this wasn't a great exercise in engineering and problem-solving, or that it's not a great demonstration of how to solve scaling problems at scale, those are definitely true. It's more that "we spent $300k/yr more than we needed to so our engineers didn't need to learn how to use our technology stack properly." I'm not meaning to be harsh, I've kludged enough garbage into production in my lifetime, but more that the fact that you got into that situation in the first place gives a poor impression of either your development team or your development processes.
- bjacokes 7y agoI don't disagree with most of what you're saying. The Nth engineer at a startup rarely looks with admiration at decisions made by the (N/10)th engineer – but it was those decisions which helped the company grow to its current size. Likewise, I think most of us will be happy if the company 10x's again. Then some super-duper-senior engineer can look at the decisions we're making now – they're not perfect, but we're doing the best we can with our current knowledge and resources – and gripe about them. The circle of life goes on. FWIW, I don't believe Node was chosen specifically for its concurrency – it was just the language chosen for the entire stack by the company's founding engineers, and lives on in just this one service.
- GordonS 7y agoGah, I can't believe you're still trying to justify this madness! In a parallel universe, a barely-competent engineer would have designed something more far more obvious, simpler, performant - all while using less hours, and not borderline-fraudulently wasting substantial amounts of your VC's money. If the company 10x's again, it won't be because of poor engineering, it'll be because of marketing and VC's who don't know how you're wasting their money. If the company 0.1x's, it'll likely be because of a security breach because of appalling design.
- acrispino 7y agoSo, in your estimation, plaid's engineering is a bunch of less-than-half-competent madmen, who might as well be committing fraud, correct? Is that overly negative or just the right amount?
- GordonS 7y agoAh, a serious question as an aside to my last (scathing) comment - does Plaid have architects? What about architecture and/or code reviews? I'd be very interested in reading about that.
- bjacokes 7y ago
- GordonS 7y agoReading the article, I'd fully expected a post-mortem at the end, describing how architecture and code review processes were going to be tightened up to ensure a monstrosity like this never happened again - that would have been transparent, interesting, and given me confidence in Plaid's engineering. Instead, you've peppered this thread with comments that kind-of, sort-of justify the approach taken. I'm sorry, but this approach cannot be justified - it's overly complex, and far from the simplest or most obvious approach. I'm truely shocked that Plaid has produced an architecture like this, and doubly so that Plaid would try to justify it. My guess here (and given the attempts at justification, this is me being really charitable) is that a junior dev was given too much leeway, and did some resume-driven-development, just so they could say they'd worked with 4k containers.