16 ms·
We 30x'd our Node parallelism
- 7777fps 7y ago> We were running 4,000 Node containers (or "workers") for our bank integration service. The service was originally designed such that each worker would process only a single request at a time. This design lessened the impact of integrations that accidentally blocked the event loop, and allowed us to ignore the variability in resource usage across different integrations. But since our total capacity was capped at 4,000 concurrent requests, the system did not gracefully scale. I can't be the only person who reads stories like this and wonders how they arrived at that solution in the first place? Failing to scale because their previous approach to scaling was a worker per request, a model which was roundly moved away from, because that's how CGI and Apache modules worked and it didn't scale well. I thought one of the key selling points with Node was an fully async standard library, enabling better scaling in process. But then you read stories like this, and I find it hard to relate to the original problem.
- rynop 7y agoAgreed. This is the second scratch my head moment from the Plaid engineering team blog recently.
- phoe-krk 7y ago> I thought one of the key selling points with Node was an fully async standard library, enabling better scaling in process. We still have an event loop that is trivially blocked by very simple programmer errors, destroying the whole advantage that you describe here. The fact that Node ships a fully asynchronous standard library doesn't in any way fix the fact that Node is a runtime for a language that itself is a mistake.
- nicoburns 7y ago> We still have an event loop that is trivially blocked by very simple programmer errors, destroying the whole advantage that you describe here. So they fixed the issue that some requests blocked... by making all requests blocking.
- phoe-krk 7y agoThis is the worst kind of software engineering. There is a massive deadlocking design mistake in the centre of the language - literally a huge red button with DO NOT PRESS printed on it. Thousands of programmers pass it by every single day, or hour, or minute, and the creators of the runtime insist that it is impossible to fix that button whatsoever; instead, all users need to work around it by ensuring that their code in no way presses that red button on purpose or even by accident. These people insist that it is impossible to program normally and in a language that is actually sane and does not advertise obvious and gaping design mistakes as "features of the language". These people advertise the analogue of Python's Global Interpreter Lock as the core foundation of their language. These people advertise Node and the language it implements as practical for implementing multithreaded applications. Posts such as this show what sort of bullshit it is; it is only practical to use Node for parallelism if each single Node instance is only ever run single-threaded. You don't parallelize by running multiple threads, you parallelize by running multiple Node runtimes. This is no longer an act against productivity or usability. This is simply insane and shows one of the most basic things that are wrong about Node's language and approach. It is impossible to write a multithreaded program if your language of choice makes it trivial, and practically unavoidable, to globally lock your whole runtime with every single line of code you write and import as your dependencies.
- jononor 7y agoThe blog in question indicates some uncommon and questionable engineering practices with Node.js. There are likely hundreds of success stories for every one like that. The first Node.js service I wrote and maintained, processed thousands of requests in parallel and was successfully in production until the company it was developed for ran out of money.
- x86_64Ubuntu 7y agoBut how many of these "questionable" practices get held up as engineering marvels by the creators of the monstrosity? When you look at this blog post, you can see the author really felt a sense of pride in this Frankenstein and wanted to show the whole village.
- duxup 7y ago>by very simple programmer errors I can't help but also feel that is also an issue and in another given language this issue might not happen ... but they'd hit another. It's so easy to say "don't do X because problem Y won't happen" but hard to predict what happens when you move from (language, platform, or whatever) X to (language, platform, or whatever) Z.... and I suspect people often hit issues and realize that maybe Y wasn't the problem. I see it all the time and I feel like "Wait guies I'm not sure we're fixing the right thing!?!?!" This article raises a lot more questions than answers IMO.
- deleted 7y ago[deleted]
- deleted 7y ago[deleted]
- davedx 7y ago> trivially blocked by very simple programmer errors Can you give an example please? I think it's much easier to block a thread with C#'s async programming model than node's...
- dmitriid 7y agoNode only has one thread. Everything else follows.
- coddle-hark 7y agoNo it doesn’t. We’ve had good models for concurrency in single-threaded systems for a while now.
- PixyMisa 7y agoThe model being "don't do that".
- dmitriid 7y agoYou say no it doesn't and then speak about concurrency models for single-threaded systems. Choose one :) Node can't not be single-threaded. Because Javascript is. Node.js is single threaded. It has a single event loop in a single thread, and all the "concurrency" is simply queued on that loop. It offloads some tasks to libuv for some system-related tasks but that's it. And the thread pool that libuv creates is very limited. Anything that doesn't end up in libuv (that is, probably vast majority of user code) will only ever run in one thread... because Javascript is single-threaded hence V8 is single-threaded hence Node.js you get the gist. And of course, node.js even has a separate documentation section titled "Don't Block the Event Loop (or the Worker Pool)" [1] because it's trivial to block the event loop. [1] https://nodejs.org/en/docs/guides/dont-block-the-event-loop/ https://nodejs.org/en/docs/guides/dont-block-the-event-loop/
- willglynn 7y agoNot OP, but… concurrency is not the same as parallelism. This is an important hair to split: https://stackoverflow.com/a/24684037/1026671 https://stackoverflow.com/a/24684037/1026671 "Concurrency" doesn't warrant scare quotes even when describing a single threaded program. It's true that Node.js allows one task to block many others, but that's an implementation detail of Node.js, not a guaranteed result which necessarily applies to any program using only one OS thread. Other programs suffer from analogous starvation/deadlock/livelock/priority inversion problems, but those would be implementation details too, not guaranteed results from using multiple OS threads.
- earthboundkid 7y agoAsync is just modern cooperative multitasking, and just like the 90s, it's easy to accidentally lock the whole system.
- phoe-krk 7y agoWe are no longer in the 90s. The code has increased in volume a hundredfold and it comes from literally everywhere. You can no longer trust everything on your machine or your network to be bug-free or otherwise non-hostile. Creating a system in the 21st century that tries to follow ideals from the 90s gives us the kind of idiotism that we can witness here.
- jerf 7y agoI think you may have misinterpreted earthboundkid; the claim isn't that it worked in the 90s, the claim is that it was already broken in the 90s. You are otherwise on the right track, though Node does technically have one advantage, which is that it is a cooperatively-scheduled island in a preemptively-scheluded overall OS. In the 1990s, when the cooperatively-scheduled program was not cooperative, you locked the machine, not the process [1]. There is a reason why Apple went very aggressive with the OSX rewrite; the previous Systems had basically written themselves into a corner where they had to use cooperative multitasking because so much code made use of the implicit promises it provides, yet they could no longer afford to compete with Microsoft if they didn't get off it it, because the complexity just kept going up, up, up and the problem was going to continue getting exponentially worse. For a Node program, you only have to account for the Node program itself, not everything running on the computer. Still, you're in the same exponentially-growing-complexity trap (with a very initially-safe-seeming low exponent, but it still gets you in the end), you just reset yourself back to a point earlier on the curve. [1] There are various details, caveats, interrupts, etc, the picture is more complicated than one sentence can convey, but the principle still held and it was still possible to wedge the machine fairly badly for varying periods of time with simple bad code.
- winrid 7y agoI mean, if you write slow code in a high throughput environment you'll just kill the CPU from context switching between threads instead. It doesn't matter if it's in an event loop or thread per request. Architect things correctly.
- Crystalin 7y agoI got the same feeling. We use node and usually split to 500 concurrent requests per process. Still interesting...
- eshyong 7y agoI feel like this article is missing a crucial piece of information: why was integration code was blocking the event loop in the first place?...
- techterrier 7y agoI wonder if all this was at root, pickup up jobs from a message queue and they only wanted each process to only have one job in flight at once.
- Vesuvium 7y agoThe way it happens always: many people working to a solution, not agreeing on one and then comprising on something in the middle, even if it makes no sense.
- api 7y agoI wonder what percentage of the massive compute power of huge cloud data centers is spent just chugging away on ugly clunky hacks to run bad code?
- asdfman123 7y agoQuite a large percentage. I worked on a site that got 1 request per second and they were able to handle it by spinning up like 20 VMs. Turns out they were just using Entity Framework wrong. Whoops. But also, though, you have to consider that most places aren't Plaid, and most places developer time is more expensive than throwing an extra machine at the problem.
- liveoneggs 7y agoPretty sure apache + cgi would scale better :)
- bjacokes 7y agoThere are a couple of reasons that the legacy scaling model was viable for us. As mentioned in the post, only 1/10 of our traffic was from the API, which gave us a roundabout way to scale by diverting resources. And it's only viable to use this model of scaling when the business value of a request is high – we were originally quite happy to spin up more containers when we reached our scaling limit. That's the pragmatic reason why we were processing one request per container. In terms of what issues caused us to move away from parallelism in the first place, it was all the CPU-bound stuff that you might expect: ReDoS-style issues, post-processing arrays in very large edge cases, programmer error, etc.
- abalone 7y ago> In terms of what issues caused us to move away from parallelism in the first place, it was all the CPU-bound stuff that you might expect: ReDoS-style issues, post-processing arrays in very large edge cases, programmer error, etc. But these are not parallelism problems. These are single threading problems, which the core problem with Node.js, not parallelism in general. Hence I think the question stands: why did you choose node for this?
- bjacokes 7y agoIt was chosen about 6 years ago when the product was first being developed, so most of us on the engineering team weren't around when the decision was made. The main choice we're making at this point is: what's the impact and ROI of a language migration vs getting Node to work as well as we can?
- ubu7737 7y ago> what's the impact and ROI of a language migration vs getting Node to work as well as we can? Hire an architect costs what? Putting the genie back in the bottle was a problem Plaid baked into its early success, which is common with startups hiring engineers with zero architecture knowledge.
- ilaksh 7y agoThey didn't actually understand Node very well at first and then later they figured it out.
- sergiotapia 7y agoAt some point you have to take a step back and realize you've grown beyond your tech and reach out for something else. Elixir sounds like a great fit for these problems. For example Discord reached out to Rust and built tiny Rust components that are called from Elixir for their server user list. Some servers have 200,000+ people online, and Elixir wasn't cutting it performance wise. Rust, boom now it works.
- asdfman123 7y ago> I can't be the only person who reads stories like this and wonders how they arrived at that solution in the first place? Here's how it probably worked: they liked Node, they liked containers, they put Node into containers and it worked, and they stuck with it as the user base grew.
- Turbots 7y agoUsing an interpreted language and complaining about performance is always a funny thing to read. Use Java, C++, Rust, Go or C if you want pure performance. Use Java with something like Spring + a Reactive stack to scale up to much higher concurrent numbers of requests. You'll have the same easy programming model that Node provides and you'll have both performance from the language and the more efficient execution. If you're worried about the startup times and memory footprint that comes with a JVM based application (although this has improved substantially the last few years), go to a natively compiled language like C, C++, Rust or Go. Or compile your Java apps using GraalVM native images.
- praseodym 7y agoRelated, from the article: > We hypothesized that increasing the Node maximum heap size from the default 1.7GB may help. To solve this problem, we started running Node with the max heap size set to 6GB [..], which was an arbitrary higher value that still fit within our EC2 instances. Sounds like they were utilizing their EC2 instances very poorly. Why not run more workers per instance, or switch to an instance type with less RAM (or more CPUs)?
- tracker1 7y agoThey were using ECS, it also looks like they had to work through a couple bottlenecks to get multiple requests per node working well... I think they could get further by using the newer workers api, since gRPC doesn't work with the cluster module.
- crazygringo 7y agoYeah, I don't get it either, at all. The original poster wrote below: > In terms of what issues caused us to move away from parallelism in the first place, it was all the CPU-bound stuff that you might expect: ReDoS-style issues, post-processing arrays in very large edge cases, programmer error, etc. But it's trivial (a single line) in Node to place breaks in CPU processing to allow the event loop to fire, and as for "programmer error"... many commenters below are also complaining async programming is too hard or finicky. But that's like complaining about C because pointers are hard, or Java because OOP is hard, or databases because planning indexes is hard. Once you "get" async, pointers, OOP, or indexes, it's easy. And it's part of your job as a professional programmer to get it. Async is no trickier than anything else. The setup in the first place makes absolutely no sense to me, using a language exactly opposite of how it's meant to be.
- markandrewj 7y agoThe issue is that many developers that are coming from synchronous programming don't get asynchronous programming. They could both improve the code by not writing blocking code, and also using something like the cluster module (https://nodejs.org/api/cluster.html https://nodejs.org/api/cluster.html).
- crazygringo 7y agoI get that. What I don't get is how nobody treats it as an issue when developers coming from Python or Java to C don't get pointers. The assumption is that you learn. But for some reason, people think it's "OK" to not get async, that it's the language's fault rather than the programmer's. That's what I don't understand. It's like a different cultural standard gets applied.
- markandrewj 7y agoI agree with you, but I think many developers get reluctant to change when they have been doing something one way for a long time, especially if they feel that one way works fine. I can also understand the position, as sometimes it can be fatiguing when technologies are constantly changing. For this project though, if they are actively going to avoid asynchronous programming, they may have been better off choosing a synchronous language.
- tluyben2 7y ago> I can't be the only person who reads stories like this and wonders how they arrived at that solution in the first place? No you are not. I wonder which CTO would allow this; like everyone here, the exact case is not really clear (or at least why this solution is a great solution for it), but this sounds like a weird solution (and expensive) to some issue. I really don't understand these 'solutions' and I am almost 100% sure I (with a team! but the point that this is not the best solution for the problem) can whip up something far simpler and more efficient for this problem. But ofcourse there are problems that might fit?
- rynop 7y agoI’d be curious to hear your reevaluation of moving this to Lambda after some of the major announcements during re:invent. My guess is some of the reasons you went ECS have been addressed with these announcements. Obviously some of the new features are still preview, but would be interested to hear your analysis none the less.
- bdcravens 7y agoOftentimes there's a several month delay from when stuff is announced at re:invent and when it's GA. I don't think anyone would ever make technical decisions based on announcements; they would wait until they could touch it and actually create a proof of concept. In other words, the "analysis" is nonexistent, since there's nothing to analyze.
- tyingq 7y agoDoes node have something similar to how apcu is used with PHP? That is, an mmap based kv store so that if you choose to run more than one node process on a single server, it has a fast kv cache? I'm aware you can use redis or similar, but a simple mmap kv store is simpler and faster for a single server use case.
- ddorian43 7y agoYou can use something like LMDB on every language.
- godot 7y agoI totally see what you mean, coming from a PHP world myself a few years ago. The key thing to note is that node.js (like many other languages including Java) starts a server process that basically does not stop until you explicitly restart it (or it crashes); unlike PHP where every request starts a brand new process on a clean slate (hence needing APCu to store a local memory cache per server). Meaning, what you can accomplish with APCu in PHP can be trivially accomplished by a simple Object in node.js (i.e. a map/hash), by virtue of having a require cache (hence every time you require'd the lib it returns the same instance of the object). If you want a simple open source lib to do exactly that for you and provide an easy to use API, you can use something like https://www.npmjs.com/package/tmp-cache https://www.npmjs.com/package/tmp-cache .
- tyingq 7y ago
- rauchp 7y agoThat was an interesting read, thanks for linking to it. It's hard finding articles online discussing Node and performance, most people just dismiss it as an unviable option due to scale and speed concerns. 30x really is quite the jump though. > Each Node worker runs a gRPC server Not going to lie, this kind of surprised me. When I think of a Node backend I think of ExpressJS. Not because I think Express is better, but because it's been pushed around in the past few years as the fastest, simplest way of running a backend. Yet, if you're going to be running a gRPC server, why not use a more performant language with better multithreading support? I thought this article was about them optimizing a grandfathered-in solution (such as Express), but I can't tell why they built out a gRPC server in Node in the first place.
- bjacokes 7y agoOur integrations are primarily written in Node, which was the original language used for everything at Plaid. Almost all of those original services (except for integrations) have been migrated to Go or Python at this point. We've standardized on gRPC as our wire format, so we stayed consistent and used gRPC in Node. With perfect hindsight, it's a fair point that all the pros and cons could net out to another language being best for our integrations. Integrations are the largest and most quickly-changing codebase at Plaid, so such a migration would be a massive undertaking. We definitely didn't want to block scalability improvements on doing a language migration.
- mnutt 7y agoI've been hoping that the Cloudflare folks will open source parts of their Workers; they seem to have figured out a secure, performant way to run untrusted javascript at scale.
- jrockway 7y agoThe Node gRPC implementation is fine. It uses the C++ implementation which is the gold standard. It has Prometheus and OpenTracing interceptors. You basically give nothing up by using it, if your team wants to write a language that runs on node.
- tyingq 7y ago"Only 10% of Plaid's data pulls involve a user who is present" Since they provide an API, it seems like some of the calls where they think a user isn't present might actually have one present.
- CamouflagedKiwi 7y agoThe other 90% are not triggered by the API, they are "periodic transaction updates" - presumably they refresh once a day or something.
- tyingq 7y agoYeah, I read that, but it's not clear exactly what those calls are. It sorta sounds like making assumptions on how their users are using the API. In fact, it sounds like they think "linking an account" is the only "user present" API call: "Only 10% of Plaid's data pulls involve a user who is present and linking their account to an app"
- namibj 7y agoNo, I'd read this as "linking their account to an app" meaning that the plain account's API credentials are configured in the app, so the app can call the plaid API (presumably interactively on user interaction).
- bjacokes 7y agoWe thread knowledge of whether a data pull was initiated by the API or by our cron-style service into our load-balancing layer, so this ends up being pretty straight-forward.
- tyingq 7y agoAhh, got it. The "present and linking their account" part threw me off. Sounded like only the "linking" call was getting the fast lane.
- spamizbad 7y agoThe only way this makes sense to me is if they have to contend with lots of expensive parsing, event sequencing, and throttling requirements. Payment APIs, bank websites, etc can be quite byzantine. I could understand how one might code yourself into a corner with a monolithic node app and basically just say "F-it, we're doing this synchronously!" I don't even think it's a terribly bad thing to do assuming it favors feature velocity.... but at that point, I'd recommend moving away from Node towards something like Python. And if you wanted to dip your toes back into async plumbing land, explore Go or Elixir.
- duxup 7y agoThat was my thought to. They've got a problem where they've got no idea what a given transaction costs and some unpredictable amount of transactions result in some serious work that holds up the event queue. God knows they could be waiting for some reel to reel tape to spin up somewhere...
- coddle-hark 7y agoThe whole point of async I/O is to be able to do something useful while waiting for tape to spin up. I don’t buy it.
- duxup 7y agoThe article certainly raises more questions than answers that's for sure.
- spamizbad 7y agoBut you need to know if you can do that something first, or if you've done that something too many times in the last N minutes (and could get blocked, forcing thousands of other somethings to get endlessly queued). Or if that something could take too long, and actually you could be doing 200 other somethings in the same time etc. It's not that simple.
- Slartie 7y agoBut the whole point of synchronous I/O is to isolate the programmer from having to think about that spinning tape up takes a non-zero time. I have a feeling that this gets lost sometimes in all that "async I/O is the GREATEST!" craze. Async is nice - if you can handle it. But this is not easy to do in complex systems and processes. It is certainly easier to work with an old-fashioned process that blocks when waiting for whatever you need to wait for, and just scale by letting the OS run lots of those in parallel. Sure, it's less efficient. But it's easier for the devs to handle. I just read the hidden undertone of this article as "our devs aren't that smart after all".
- bfrog 7y agoIt's not clear from the article why they were only able to run one request per node process, but that alone would make it questionable why use Node at all then. The entire point of the environment has been nixed. The article is quite confounding to understand how they arrived at that point in the first place.
- Scarbutt 7y agoI don't want to be that guy, but why did they start with nodejs for something like this instead of using the JVM or Go?
- caiobegotti 7y agoThey probably already had some decent experience with Node and it solved their initial problem well enough, refactoring or rewriting costs usually make engineering managers frown upon (wrongfully) and so it becomes much harder to fix this in the long term.
- hatch_q 7y agoI have experience with node, am frontend developer... but would never in wildest dreams use Node for any kind of production backend. And it's not even the question of programming language - the problem is NPM and whole package management which is inherently insecure.
- meritt 7y agoMy guess is because their system is primarily issuing HTTP requests and extracting data out of responses: html, xml, json, plaintext, etc. Web scraping is a messy business and using a language that allows you to be flexible with string manipulation and types goes a long way toward sanity.
- calibas 7y agoHow is Javascript better at string manipulation? I've never encountered anything special there that I can't do in just about every other language. Javascript just has more helper functions out of the box.
- meritt 7y agoI wouldn't characterize it as "better" but specifically easier and more flexible for the people writing and maintaining these scrapers. I'm also speaking more broadly about scripting languages (not javascript specifically) vs the aforementioned JVM or Go, and the ease with which you can deal with inconsistent, frequently changing, and often completely invalid inputs from a wide variety of data sources. Plaid's use case here is automating logins, responding to captchas, manipulating those on-screen virtual keypads to respond to security questions, chaining together multiple HTTP requests, and then parsing out frequently invalid, rapidly changing, and just plain broken content from a wide multitude of banking websites.
- deedubaya 7y agoA good example of avoiding premature optimization. I'd imagine delaying tackling this problem freed them up to tackle problems that impact users.
- coddle-hark 7y agoThis only holds if they didn’t pour hours into the original solution. Setting up and managing 4000 node services doesn’t sound like a quick hack.
- deedubaya 7y agoWriting and maintaining concurrent code for greenfield projects is relatively hard compared to sync code. Provisioning and deploying with ECS is usually just mouse clicks.
- bjacokes 7y agoWhile we were worried about event loop blockages causing outages, another more subtle problem would have been if event loop blockages doubled our user-facing latency. (If you read the section on latency ratios, you'll see that comparing parallel vs non-parallel workers was the most useful stat in figuring out how effectively we were using the event loop.) It definitely gave us peace-of-mind to know that event loop blockages wouldn't have an effect beyond the requests they're processing. Honestly, the accounting for which would've been higher impact – investing in parallelism earlier, or adding infrastructure and having more resources to devote to other pressing needs – is difficult to do, even in retrospect. There was surprisingly little effort required to get to 4,000 node containers in an ECS cluster, other than deploy speed issues which we talked about in a previous post [1]. But it's possible this migration process would have been easier if we had done it sooner. [1] https://blog.plaid.com/how-we-reduced-deployment-times-by-95/ https://blog.plaid.com/how-we-reduced-deployment-times-by-95...
- ubu7737 7y ago> But it's possible this migration process would have been easier if we had done it sooner. What the f*? Of course it would have been easier if you had done it sooner. What you lacked was the willpower from decision-makers who had growth of dollar-signs in their eyes. You've littered this thread with comments explaining how every move you made was based on ROI. That's the kiss of death for architecture concerns, and bizarrely it puts Node.js on the list of runtimes for data/stream processing backends. No matter how many times you explain how you made these decisions, I can't help getting the feeling you were wearing horse blinders. Edit: I find it impossible to imagine that nobody on the engineering team ever shouted, Hey look out! We are basically a Web farm for banking-related requests, this is insane! Surely you've heard from those people and they were let go.
- mceachen 7y agoIn case anyone else gets excited by JSONStream, know that the package hasn't been updated in over a year, and the GitHub repo was archived by the author with no link to a successor.
- fenwick67 7y agoOboe has a similar API, can't speak for performance though. http://www.oboejs.com/examples http://www.oboejs.com/examples
- contrahax 7y agoI'm maintaining a fork here that incorporates all of the valid open PRs from the original repo + some more updates: https://github.com/contra/JSONStream https://github.com/contra/JSONStream It isn't published on NPM (you can use it as a git dependency) but if people are interested I can.
- mceachen 7y agoThanks for sharing! Why don't you publish releases?
- nosianu 7y agoThey write (somewhere in the middle) > Since V8 implements a stop-the-world GC, new tasks will inevitably receive less CPU time, reducing the worker’s throughput But there is this Google blog post vom January 2019: https://v8.dev/blog/trash-talk https://v8.dev/blog/trash-talk > Over the past years the V8 garbage collector (GC) has changed a lot. The Orinoco project has taken a sequential, stop-the-world garbage collector and transformed it into a mostly parallel and concurrent collector with incremental fallback. So I guess they used an older node.js version. The current LTS version is 12.x and it is from around the middle of this year. --- PS: If the blog author reads this, there is an accessibility problem with the Google-hosted inline images. If I try - without ad blocker - in an anonymous window I see none of the inline images. Logged into Google with my own account I can see some but not all the images. Apparently which images I can see depends on being logged in to my Google account? I also tried IE Edge just to see if the browser makes a difference - no inline images visible there either.
- greentrust 7y agoDitto, images weren't showing up for me
- jimbo1qaz 7y agoWhen I try to view the image in a new tab, I get: Your client does not have permission to get URL /Iw-RdHoPjbwuSAqJHK3C0Sy8m29NqzeHPtmJ7CVFuYqwr4CbwpGjwn9O4bcDNtCf_hLD4FGc75nkQYnJBgyA-CT2ikBDWQD-nAtqxXa4Lw2yDuh_-ywcsDaer6m4LyVtljwfrajO from this server. (Client IP address: [redacted]) Rate-limit exceeded That’s all we know.
- bjacokes 7y agoFixed the images about half an hour ago, sorry about this!
- mnutt 7y agoI’d be curious to hear more about the circumstances that ended up with a blocked runloop. Are there hundreds of junior engineers, or perhaps third parties writing code that you don’t control? I have seen people accidentally write blocking code, but not at such an egregious rate that we couldn’t catch it in code review, or at worst the runloop detector would alert on it in prod and we would roll back the deploy. For instances where you actually know you need lots of CPU, there are now strategies for offloading that specific work, although they have taken a while to get nice and easy to use.
- bjacokes 7y agoSure, one example I remember off the top of my head is a bank that sometimes returned duplicate transaction data, so an engineer had called ramda.uniq on the transaction array. Transactions are nested objects and slow to compare, so when you find an account with 100,000 transactions... kaboom. Some scenarios are more subtle, but a common theme is that the amount of data in an account can vary by many orders of magnitude.
- kevstev 7y agoI was building scalable node applications a few years ago for a very large e-commerce player- millions of customers. I think node.js is a great platform, but its apparent simplicity means there are hordes, and I mean like 90+% of the community, that can "just get things done" without understanding what is going on under the hood at all. And to be fair, for most startupy types of companies that need to iterate fast, that is what you want to optimize for. My interview screening question was pretty simple- "Is node.js single threaded or multithreaded?" And to most, they spit back the blogspam headline- "Single threaded!" I think the most correct answer is "its complicated" but would accept that because most people would say that is the "right" answer. So I would follow up with- "what exactly happens in a default installation if we have say... 5 requests come in at exactly the same time to just return some static content from disk?" (Node's default threadpool is 4). And here is where you could see their understanding just fell apart. Some would say they would be handled entirely synchronously, others completely in parallel- but then had no idea what the cause of the parallelism was. Very few actually understood that node is an event loop executing javascript backed by a threadpool for async operations. Before reading this post, I was like eh this is a waste of time- its typical medium bullshit- they almost certainly found they were doing some blocking call in the event loop and then removed it and voila, 30x speedup. It was interesting because it was a lot worse! They spent all this time and hard work figuring out everything but what was taking so long in the event loop, and it seems that was the last place they actually looked. Anyway, node can be a highly scalable platform (https://changelog.com/podcast/116 https://changelog.com/podcast/116) but you need to understand it or else it will bite you in the foot. When I was last doing this stuff, upwards of 80% of our time was being spent essentially just JSON.parse()'ing, and we were looking to move to protobufs to avoid that.
- pier25 7y agoHere is a SO answer that expands a bit more on kevstev excellent comment: https://stackoverflow.com/a/22644735/816478 https://stackoverflow.com/a/22644735/816478
- x86_64Ubuntu 7y agoWhat does happen with the 5th request.
- deleted 7y ago[deleted]
- awinter-py 7y agoCompared to a compiled language, node / JIT langs make it difficult to know what will be fast in prod. V8 JIT means that things like order of keys in an object or number of different calls to a function might affect whether your function gets optimized. And there's no easy way to find out if a JS function is falling back to slow mode or to tell the buildsystem 'this is a hot path, don't let me write code that deopts this call'.
- PixyMisa 7y agoNobody involved in this project should be allowed to ever be in the same room as a computer again.
- jrockway 7y agoWhy? They had a 12 factor -ish app that scaled the normal way; run more copies. Eventually that got expensive. They had the observability to figure out what was making it expensive and whether or not their fixes had an effect. They then saved $300,000. Seems like everything went right to me. I would be worried if the blog post was "we randomly tweaked some stuff and we can't measure it but it's a little better" or "we rewrote it in go and in the rewrite introduced 87 new bugs while fixing 42 old bugs". They engineered a solution, built from good investment in infrastructure, rather than ninja-ing a hack. That, to me, is a very good thing. A lot of people seem deeply upset that Node was involved, but I think that's a red herring. The problem they had -- allocate a large chunk of memory, keep a reference to it while it is slowly sent to another server, free memory -- is going to happen in any language. (I don't super agree with their solution of "make the server faster" because one day it's going to be slow for some other reason and this problem will crop up again. Instead they probably just need a fixed amount of memory to dedicate to this process and to drop the debug payload when the buffer is full. Or just put it in the request path if it's crucial that it be produced every time no matter what. At least that will apply backpressure to calling services, pop the circuit breaker, and redirect requests to a region where S3 isn't broken. But I don't think the debug information is THAT important ;)
- PixyMisa 7y agoTo save $300,000 they first needed to waste $300,000 by reinventing a problem that was solved in 1967.
- dwild 7y agoI don't know how many software engineer they got on that team, but considering how much they raised and how much their product is used, $300k seems actually quite cheap for something that people consider here as being an awfully big mistake.
- vmarchaud 7y agoI've encounted different issues with NodeJS services in the past (and still do) both with CPU bottleneck and Heap allocations. So i wrote openprofiling-node [0] during this summer to help me profile my apps directly in production and export the result in a S3 bucket. I believe it may help someone else here so i'm posting it [0]: https://github.com/vmarchaud/openprofiling-node https://github.com/vmarchaud/openprofiling-node
- jdc0589 7y agoOn a positive note: this was a good write up. On a negative note: FOR THE LOVE OF ALL THAT IS HOLY, HOW DID THIS HAPPEN.
- supermatt 7y agoIronic. Linked images failing to display due to "Rate limit exceeded"...
- FanaHOVA 7y ago$300k is $300k, but they just raised $250M last year, is this a really good use of time for their engineering team? That's a little above ~0.1% of capital.
- deleted 7y ago[deleted]
- bdcravens 7y agoThat's just one of the benefits. > our system is more robust to increases in external request latencies or spikes in API traffic from our customers
- dwild 7y agoWhy wouldn't it be? You save 300k, that's an engineer salary... that's pretty much the meaning of a job, building value that's higher than your salary. This clearly took less than a year of engineer time. Seems like they got their value out of that employee.
- GordonS 7y agoI don't like to be overly negative, especially when a company/team is being transparent about what they're doing and giving insight into their engineering practices - but has anyone else's estimation of Plaid's engineering team just gone down the toilet? This blog post gives me the impression that either Plaid is filled with either junior or incompetent engineers - to scale to 4k containers serving 1 request each for an API workload is absolute insanity. These engineers are building stuff for banking. Banking!! There is literally no way I'm going near Plaid with a very long bargepole after reading this. It I was someone senior at Plaid, I'd be pulling this blog post before it harms reputation any further.
- sicromoft 7y agoThis comment says more about you than it does about Plaid. Their "insane" design met business requirements successfully enough to grow them into a multi-billion dollar company. Did you consider the likely (and more charitable) explanation that they were aware their design was "bad", but had higher priorities until now? If I were you, I'd be pulling your comment before it harms your reputation any further. :)
- GordonS 7y ago> Did you consider the likely (and more charitable) explanation that they were aware their design was "bad" I think marketing and VC valuations grew them into a multi-billion dollar company; whether they remain so, to a large part relies on how fast they burn through VC cash - so, not looking too good on that front... No even half-way competent engineer would come up with such a complex, unperformant solution to a simple problem - I think a higher priority should be hiring engineers who actually have a clue what they're doing. As for meeting business requirements... while this might have worked for a while, it was plainly not a good way to meet them, and given Plaid are in the banking sector, really doesn't bode well for the future (I'm having flashforwards already to security breaches, plaintext passwords etc...).
- ubu7737 7y agoDownvoted for angering the VC class...
- pdimitar 7y ago...Or you could just use Erlang or Elixir, where concurrency and parallelism come pretty much out of the box, with very little effort required for you to fine-tune the desired policy / strategy. The insistence on using Javascript is just beyond lunacy at this point.
- timmy-turner 7y agoWell, if Elixir had a Typesystem like Javascript has, I'd instantly switch to it. But atm I'm staying with Node because of Typescript.
- pdimitar 7y agoTrue, it doesn't have it. Between pattern matching and function guards however, it has a decent way to protect against common errors. The true treasure is Erlang / Elixir's runtime though. The parallelism, the self-healing, the preemptive scheduling.
- Phil_Latio 7y ago> We were running 4,000 Node containers LOL
- CyanLite2 7y agoTLDR: How to spend millions of dollars of our investors' money because we hired junior devs who chose a framework that was trendy but couldn't scale.
- mirekrusin 7y ago4k containers? That's microservices going macro big time.