3 ms·
> Working in the cloud forces you to address the hard problems first. It also forces you to address all the non-existent problems first, the ones you just wish
by sunrunner 1y ago
> Working in the cloud forces you to address the hard problems first.
It also forces you to address all the non-existent problems first, the ones you just wish you had like all the larger companies that genuinely have to deal with thousands of file upload per second.
And don't forget all the new infrastructure you added to do the job of just receiving the file in your app server and putting it into the place it was going to go anyway but via separate components that all always seem to end up with individual repositories, separate deployment pipelines, and that can't be effectively tested in isolation without going into their target environment.
And all the additional monitoring you need on each of the individual components that were added, particularly on those helpful background workers to make sure they're actually getting triggered (you won't know they're failing if they never got called in the first place due to misconfiguration).
And you're now likely locked into your upload system being directly coupled to your cloud vendor. Oh wait, you used Minio to provide a backend-agnostic intermediate layer? Great, that's another layer that needs managing.
Is a content delivery network better suited to handling concurrent file uploads from millions of concurrent users than your app server? I'd honestly hope so, that's what it's designed for. Was it necessary? I'd like to see the numbers first.
At the end of the day, every system design decision is a trade off and almost always involves some kind of additional complexity for some benefit. It might be worth the cost, but a lot of these system designs don't need this many moving parts to achieve the same results and this only serves to add complexity without solving a direct problem.
If you're actually that company, good for you and genuinely congratulations on the business success. The problem is that companies that don't currently and may never need that are being sold system designs that, while technically more than capable, are over-designed for the problem they're solving.
- themafia 1y ago> the ones you just wish you had You will have these problems. Not as often as the larger companies but to imagine that they simply don't exist is the opposite of sound engineering. > if they never got called in the first place due to misconfiguration Centralized logging is built into all these platforms. Debugging these issues is one of the things that becomes absurdly easy. > likely locked into your upload system The protocol provided by S3 is available through dozens of vendors. > Was it necessary? It only matters if it is of equivalent or lessor cost. > every system design decision is a trade off Yet you explicitly ignore these. > are being sold system designs No, I just read the documentation, and then built it. That's one of those "trade offs" you're willingly ignoring.
- sunrunner 1y ago> You will have these problems. Not as often as the larger companies but to imagine that they simply don't exist is the opposite of sound engineering. A lot of those failure mode examples seem well suited to client-side retries and appropriate rate limiting. If we're talking file uploads then sure, there absolutely are going to be cases where the benefits of having clients go to the third-party is more beneficial than costly (high variance in allowed upload size would be one to consider), but for simple upload cases I'm not so convinced that high-level client retries aren't something that would work. > if they never got called in the first place due to misconfiguration I find it hard to believe that having more components to monitor will ever be simpler than fewer. If we're being specific about vendors, the AWS console is IMHO the absolute worst place to go for a good centralized logging experience, so you almost certainly end up shipping your logs into a better centralized logging system that has more useful monitoring and visualisation features than CloudWatch and has the added benefit of not being the AWS console. The cost here? Financial, time, and complexity/moving parts for moving data from one to the other. Oh and don't forget to keep monitoring on the log shipping component too, that can also fail (and needs updates). > The protocol provided by S3 is available through dozens of vendors. It's become a de facto standard for sure, and is helpful for other vendors to re-implement it but at varying levels of compatibility. > It only matters if it is of equivalent or lessor cost. This is precisely the point, I'm saying that adding boxes in the system diagram is a guaranteed cost as much as a potential benefit. > Yet you explicitly ignore these I repeatedly mentioned things that to me count as complexity that should be considered. Additional moving parts/independent components, the associated monitoring required, repository sprawl, etc. > No, I just read the documentation, and then built it. I also just 'read the documention and built it', but other comments in the thread allude to vendor-specific training pushing for not only vendor-specific solutions (no surprise) but also the use of vendor-specific technology that maybe wasn't necessary for a reliable system. Why use a simple pull-based API with open standards when you can tie everything up in the world of proprietary vendor solutions that have their own common API?
- jamesblonde 1y ago> The protocol provided by S3 is available through dozens of vendors. But not all of the S3 API is supported by other vendors - the asynchronous triggers for lambdas and the CloudTrail logs that you write code to parse.
- j45 1y agoEnjoyed reading this, thanks for writing it. People often don't know how different might be easier for their case. Following others, or the best practices, when they might not apply in their case can lead to to social proof architecture a little too often.