26 ms·
Scaling up the Prime Video audio/video monitoring service and reducing costs
- amne 3y agobut .. but .. there's no buzz words in this solution. monolith? ew!
- dragonwriter 3y ago“container”
- marcopicentini 3y agoThey could save millions by migrating to Digital Ocean or Hetzner (+Cloud66).
- h05sz487b 3y agoSo one team at Prime for one specific application learned, that serverless was not the ideal compute model for their workload. Wow.
- basitmustafa 3y agoThe headline is a bit of a misnomer. This happens in large businesses all the time (which isn't to say it's "good", hardly is, but it suggests the causation is incorrect here, which then indicates the conclusion is entirely off-base): 1) We have sexy new product! Everyone use it so we have some use-case stories to tell and we look credible! Who cares if it's not the right tool for the job! We need a splashy way to use hackneyed business speak like "we're eating our own dog food" at the next user con so all the IT middle managers there will fight over early access and adoption. PROFIT! (Screams of technology teams in the background of "a knife is the most expensive, useless pry tool you can buy, but whatever, you are not listening, mmmkay"). 2) A few quarters/years later (if you're lucky and you made it or someone with enough gravity in their title finally saw the light): Why is expense so high in this business unit? This is insane! Let's go back to a more sane architecture. (Screams of technology teams going back to what was working in the first place, but was not sexy nor necessarily new now that no one is watching and hype cycle is over) Does this mean that serverless is useless? Dumb? Uneconomical? No way. For bursty, very short running workloads, it can be GREAT and INCREDIBLY economical. What is useless and "dumb" is whomever thought that Prime Video's encoding workloads were going to do anything but increase cost and were somehow a fit for a system whose business case specifically necessitates bursty, shorter workloads that are primarily scale-to-zero for significant periods of the day/week/month. It was a marketing stunt gone horribly wrong: intentional or not, but that doesn't repudiate the value of "serverless" for the right workloads, it just proves you better really understand the technology and the business case and the scale economics, and that goes for any technology.
- sen 3y agoEverything old is new again.
- asim 3y agoThis. All trends are cyclical. Microservices have a purpose. Monoliths have a purpose. They are not mutually exclusive. One is the path to the other but there may also be resets along the way. I spent 10 years doing microservices and now I'm back to a monolith. It's a refreshing change but it's also a project in its infancy. Breaking that out over time will only happen as and when needed.
- rbanffy 3y agoNot really - what they realized is that the billing model was not well aligned to what they were doing.
- deleted 3y ago[deleted]
- samwillis 3y agoNext they will transition to on premises hardware from the cloud to save another 90%.... oh wait...
- pbd 3y agolol. Amazon is literally where microservices became mainstream.
- selcuka 3y agoI imagine they will transition to bare metal as the next step.
- yolovoe 3y agoDon't. There's no benefit to using metal as opposed to the largest virt (which will take up the entire server anyways) pretty much. Metal just tends to be somewhat less reliable. Source: I work here.
- selcuka 3y agoSure, mine was a tongue-in-cheek comment, but there are cost benefits of bare metal in some use cases, especially if your workload is more or less predictable.
- clnq 3y agoIt turns out taking it offline has yet another 90% reduction in cost.
- NBJack 3y agoAnd vastly improves security!
- deleted 3y ago[deleted]
- rbanffy 3y agoFrom Amazon's PoV, AWS is on-prem ;-)
- ikiris 3y agoSending video frames between services is expensive, also doing per state transition hosting on things doing state transitions multiple times per second in a single stream is also expensive... Like, did they even think about cost when designing this the first time?
- Vosporos 3y agowhy should they, they're richer than God!
- eru 3y agoYou don't get (and stay rich) by wasting all your resources.
- ocdtrekkie 3y agoConsidering they don't actually pay the bill for this and it is internal accounting, probably not. Belt tightening has probably pushed cloud providers to figure out if they're wasting stuff they could put to better use, and I assume when it launched and nobody was watching Prime Video, inefficiencies were both smaller and less noticeable.
- bagels 3y agoTheir team/org has resource budgets too.
- ocdtrekkie 3y agoOh, absolutely. This makes the Prime Video team look more profitable on paper. But also all streaming services were pretty much launched with an expectation of taking losses for years, so Prime Video being expensive doesn't look unusual for a while. And since it's an internal cost, it's not actually Amazon paying someone, there's really not a significant reason for someone outside the Prime Video team to say "hey, you're too big of an AWS customer". More than likely, Prime Video making their numbers look better makes AWS' numbers look (slightly) worse, because they're doing a little less business. In the overarching grand scheme of things, this will save Amazon some amount of physical computing resources they weren't getting paid by an outside customer for, but good luck figuring out how much that actual real world savings is.
- ad-astra 3y agoStoring individual frames in S3??? Insanity! Their initial distributed architecture is unbelievable.
- selcuka 3y agoThey should've serialised bitmaps to JSON and used SQS instead. /s
- sgtnoodle 3y agoOver 15 years ago now, I was an intern at Toyota. We were working with an in-house python based framework for doing cool/terrible drive-by-wire things with test cars. I had a project to work around a bottleneck of the framework. It could only process about 70 CAN frames per second before running out of CPU. The vehicle's CAN bus had several thousand per second, though. At the time I was able to fix the problem by adding filtering to the CAN adapter's kernel module. A couple years later, I worked on replacing the python based framework with C++. I discovered the underlying root cause of the bottleneck. Someone (cough my manager) had figured out a very "pythonic" way to extract bit-packed fields from the 64-bit CAN frame payloads. They converted every 8-byte payload buffer into a canonical binary representation, i.e. ascii strings of 1's and 0's. They then used string slicing syntax to extract fields. Finally, they casted the resulting substrings back to integers. Awesome! I've since used python many times to process CAN frames in realtime, scaling up to thousands of frames per second without the CPU breaking a sweat. One trick is to use integer bit shifts and masks rather than string printing, slicing and parsing...
- bberrry 3y agoI'd be surprised if this doesn't get taken down as it casts AWS lambda in an unfavorable light (and rightly so). That's the impression I have of Amazon's leadership but maybe I'm wrong.
- dpwm 3y ago> We designed our initial solution as a distributed system using serverless components (for example, AWS Step Functions or AWS Lambda), which was a good choice for building the service quickly. The message seems more that they outgrew AWS lambda but that lambda was a good choice at first.
- supriyo-biswas 3y agoThe post literally says that they could hit only 5% of the expected workload with their server less architecture, so IMO it is still quite negative.
- simplotek 3y ago> The post literally says that they could hit only 5% of the expected workload with their server less architecture, so IMO it is still quite negative. Emphasis on "their server less architecture". Sometimes good tools are used poorly. For example they describe a high throughout workload, and each workload spread through a bunch of lambdas that handled bite size bits of the workflow. Also, they managed the workflow with step functions. Just imagine the number of network calls involved to run a single job, let alone all the work pulling data to/from a data store like S3 into/out of a lambda. I'd guess the bulk of their wall time was IO to setup the computation. Of course you get far better performance if you get rid of all these interfaces.
- deanCommie 3y ago5% of the expected workload for a Prime Video service is probably more than 99% of the workloads of the readers.
- qaq 3y ago
- bawolff 3y agoI'm pretty convinced that microservices are one of those things that make sense 5% of the time and the other 95% is cargo culting.
- wg0 3y agoAs for me,I have been trying to discover that 5% that cannot be done without microservices.
- bawolff 3y agoLike any tool, there is nothing that cannot be done without microservices. However that doesn't mean they never make sense. Microservices have certain costs and certain benefits. I can believe there are certain situations where the benefits outweigh the costs. Its just not most situations. But that doesn't mean it never exists. I could believe it makes sense in extremely large apps with huge number of different groups working on them, where the communication complexity outweighs the other complexities microservices bring.
- qaq 3y agoThat's very hypothetical. Building a distributed system just might always be way more resource intensive than comparable monlith regardless of number of developers involved.
- avereveard 3y agoRemove reliability intercorrelations, so for example that your cart api and payment gateway is up and collecting orders no matter to what happens to the front-end services. But then for perfect decorrelation you'd also need independent databases behind the microservices, and queues between them for horizontal communication,and few are actually going all in with that, and so fall in the 95% where they go trough te motion and the effort of splitting microservices,but reap no actual benefit from it.
- kristopolous 3y agoThere's a bunch of things I'd like to have them do. If they could span across machines like clusters that would be amazing. If I could trivially package them up and deploy them locally with intrinsically less effort and wall-time then the old way, that'd be amazing. If I could somehow get the horizontal scaling promises and redundancy as some kind of built-in, like I can with say, memcache, that's be cool. If I could do these kinds of "hard" things with them more trivially, that'd be really nice. There's a lot of things I want them to do but it's a god-damn bull-riding rodeo every time I try to get there. And before you reply, I know you're an expert and can do all these things trivially. That's amazing. The vast majority of the industry creates a giant fragile spaghetti knot with them and I am not a full time k8s admin nor do I want this to be a career trajectory. It should be like you know, wine, ffmpeg, imagemagick, virtualbox, lua, qemu, redis, gnuplot, lvm2, gdb, ssh, sqlite; tools like that. It's pretty easy to get them to do really nice things. Those things deliver on their promises and potential pretty nicely. It's nice that nobody feels a need to hype curl or squid. They just work. Isn't that nice? I mean look at gdb's website: https://www.sourceware.org/gdb/ https://www.sourceware.org/gdb/ it doesn't even have CSS animations --- in fact, it doesn't even have CSS.
- asdfman123 3y agoWhat are the developers doing, though, if they’re not diagnosing why Reaper isn’t communicating with the Zanzibar service registry?
- raverbashing 3y agoI love how some of developers jumped on the serverless bandwagon with some of the least "serverless" workloads first "Let's make our entire website serverless now" erm, no? It's cargo culting of the worse kind
- selcuka 3y agoIt's the same story as NoSQL. "Let's migrate our transactional data that requires strict referential integrity to CouchDB... Oh, wait..."
- quickthrower2 3y agoWaiting for the next wave of tech misuse due to LLMs and ML!
- Garlef 3y agoThere's a lot of room for "chinese-room" automation: Write a verbal description of your cloud function and let the LLM simulate the execution. Very cheap to develop. Very expensive to execute.
- quickthrower2 3y agoI think that would be done via code creation unless the function needs LLM qualities. But LLM TDD where both the tests and code are autogenerated could be a thing for sure. And it will be microservices so that each service is easy to generate by LLM!
- bobsmooth 3y agoPrompt: Is the following number even or odd...
- selcuka 3y ago> Prompt: Is Avogadro's number even? > ChatGPT: Yes, Avogadro's number is even. The value of Avogadro's number is approximately 6.022 x 10^23, and since it ends with the digit 2, it is an even number. Right answer, wrong reasoning.
- ggm 3y agoamazon product ditches amazon product for another amazon product? feels very strongly they just moved from one AWS platform to another. delay between asynchronous communicating processes differs in these architectures and I suspect they were unable to orchestrate microservices to match the RPC "inside" a monolith model. Nobody can: It only matters if your IPC is causing delay you can avoid. Most of us aren't in a room where the real cost is high: 90% of computers are more than 90% idle 90% of the time. Amazon is not in that cohort.
- 1-6 3y agoPerhaps Amazon reached peak saturation for its video streaming services so it no longer needed unknown unknowns from holding it back from using a more efficient monolithic architecture. Distributing services across multiple machines is certainly more scalable but all those API calls can add up.
- bberrry 3y agoI wouldn't call it a monolith as the number of instances could be scaled up. Mono implies single instance. They just combined multiple microservices into a larger one.
- config_yml 3y agoI am not sure if you’re joking.
- rbanffy 3y agoWe could call it a polylith.
- bberrry 3y agoI don't see whats funny about my statement. Please elaborate on your definition of monoliths vs scalable microservices.
- lastangryman 3y agoTo most people, "mono" refers to a single codebase, not a single deployed instance. I've worked on many monoliths that run multiple instances in production. Microservices are no more or less scalable than a monolith. The main benefit of Microservices is allowing multiple teams to work independently from each other without everyone "stepping on each others toes". You can have scalable monoliths and unscalable microservices.
- bberrry 3y agoThank you for clarifying. It's seems my definition doesn't correspond to others'. Then do we lack a word for a monolithic application that handles everything and that can only have one instance running?
- declnz 3y ago"Singleton" sounds like a good word here (though I've never heard it used in that context)
- christkv 3y agoI've never seen successful micro services if the starting point is not a monolith. The most successful ones I've seen are hybrid ones where some parts needed to be scaled are refactored as a micro service to run in parallel.
- lastangryman 3y agoBang on. A friend I work with used to say "microservices are for scaling teams, not tech" which I liked. Even with monolith -> microservices I've seen it go wrong. One Go application I worked on it would take a senior engineer a week to add a basic CRUD endpoint as the code had been split in to microservices along the wrong boundaries. There was a ridiculous amount of wiring up and service to service calls that needed done. I remember suggesting a monolith might be more appropriate, and was told it used to be a monolith but had been "refactored to microservices"... This type of stuff can literally kill early stage companies.
- vivegi 3y agoOf all the video streaming services I have used, PrimeVideo is the one where the video/audio sync becomes terrible progressively. It is pretty bad. It happens in 8 out of 10 movies. There is some misconfiguration in their AV transcoding pipeline. And here, we have an article talking about Monolith vs. Microservices improving user experience.
- nicoco 3y agoNevers had any issue with bittorrented AWS.
- sgtnoodle 3y agoOf all the streaming services that have irritated me, I can't recall any serious technical problems with prime. I suppose I have a vague memory of poor AV sync that could have been on prime, it was always a problem at the start of streaming that would work itself out after a few seconds. Netflix's shiny new compression scheme a couple years ago didn't work on my Sony TV's buggy silicon. The only way I got that fixed was by knowing someone on the inside. Hulu usually can't make it through an episode without the video freezing at least once. Sometimes it just refuses to work at all until I completely reboot the TV. HBO Max's UI is just really cheesy and slow, but whatever it's fine. Paramount+ is my new favorite to hate on. The UI is maddeningly glitchy and lethargic. I pay for no ads, but it plays ads anyway, on Star Trek episodes from 1996. It doesn't remember progress in a show more than once every week or two, just enough to remind you that it's supposed to be a feature. On my phone, it doesn't hide the typical menu overlays unless I do a complex sequence of finger taps. One time I tried to file a bug report from inside the logged-into app, and I got an email back claiming that they would love to consider my concerns but can't because they don't have an account associated with my email address.
- vivegi 3y agoFor me, the sync is fine at the start of playback on PrimeVideo. It just becomes bad progressively (which leads me to believe they have used a video framerate that is ever so slightly different from the source and have keyframes insertion after a longer than optimal duration; similarly sample rate mismatch for output audio relative to input audio stream could be a potential cause). And I use a FireStick, FWIW. BTW, their own trascoder product MediaConvert seems to have this issue (It is possible that it could be user error too in how they have used the product or setup the parameters). [1] My guess is PrimeVideo dogfoods MediaConvert and they also have this issue. They could have fixed it for newer content, but previously transcoded content still has issues (which will remain until they are re-transcoded). [1]: https://repost.aws/questions/QUGajgu4zKTlewlTg1M96i_Q/questions/QUGajgu4zKTlewlTg1M96i_Q/mediaconvert-introduces-audio-video-sync-issue https://repost.aws/questions/QUGajgu4zKTlewlTg1M96i_Q/questi...?
- fastest963 3y agoClickbait title. The expensive part was passing around individual frames and the associated S3 operations. It's not clear if they could've kept a distributed architecture but made the work units be chunks of frames or even whole videos. Monoliths can inefficiently use S3 and other cloud services to rack up a huge bill.
- LVB 3y ago> Moving the solution to Amazon EC2 and Amazon ECS also allowed us to use the Amazon EC2 compute saving plans that will help drive costs down even further. So various parts of Amazon have to work through the AWS same pricing programs that the rest of us do?
- zoover2020 3y agoThere are internal discount rates per service (IMR), but there's no such thing as free lunch Also, Prime Video isn't part of AWS but the consumer / devices / other part of (retail) Amazon. Source: worked there
- ldargin 3y agoYes, to keep track of costs.
- Garlef 3y agoAlso they might be actually be different legal entities.
- dannyobrien 3y agoAm I right in understanding this is just their defect-detection system?
- AndrewPGameDev 3y agoYes, this is just the defect detector and not the actual video streaming service.
- iamflimflam1 3y agoThis really is a click bait title. They are talking about their video quality monitoring service, not their video streaming service. It’s something they use to check for defects in the video stream - hence the storing of individual frames in S3. Original title: Scaling up the Prime Video audio/video monitoring service and reducing costs by 90%
- debdut 3y agoThe subtitle is "The move from a distributed microservices architecture to a monolith application helped achieve higher scale, resilience, and reduce costs." And the article itself mentions the 90% cost reduction. So the title seems pretty much in-line with the original intent.
- ojkelly 3y agoBut, by omission is reads that Prime Video rebuilt their stack without serverless and got a 90% cost reduction. This post is going to pick up a lot of traction and I suspect these comments are going to bikeshed monolith vs microservices for the next day. On reading it, this is for a video quality monitoring system, that needs to consume and process video. Generally a compute and time intensive task. Something not always suited to severless, particularly when it’s not easy to parallelise. The task at hand doesn’t sound ideally suited to serverless, but the existence of the post shows that’s not readily obvious. So it’s a valuable post to explain a scenario where a few big machines is the best call. But the sensationalism of the headline, would suggest all serverless is expensive and wasteful. When in reality the same is true for a non-ideal workload on a monolith.
- moonchrome 3y agoServerless has such bullshit insidious pricing that makes it seem like you're saving money only to figure out you're in shit once you're knee deep in it. For example you'll have to read fine print to find out that 256MB lambda will have the compute power of a 90s desktop PC because compute scales with memory. And to get access to "one core" of compute you have to use like 2GB of memory. Now you may say "serverless isn't geared towards compute" - but this kind of CPU bottlenecking affects rudimentary stuff - like using any framework that does some upfront optimizations will murder your first request/cold start performance - EF Core ORM expression compiler will take seconds to cold start the model/queries ! For comparison I can run ~100 integration tests (with entire context bootstrap for each) against a real database in that time on my desktop machine. It's unbelievably slow - unless you're doing trivial "reparse this JSON and manually concat shit to a DB query" kind of workloads. You could say those frameworks aren't suited for serverless - or you could say that the pricing is designed to screw over people trying to port these kinds of workloads to serverless.
- bagels 3y agoI don't want to come off too harsh on this, but it sounds like the service didn't meet the initial design requirements? Some of this would have been really easy to predict (eg. hitting account limits) if they simply took the time to calculate how many workflow transitions they'd need to execute for the load.
- jjevanoorschot 3y agoThe title is editorialised to be clickbait. The original title is "Scaling up the Prime Video audio/video monitoring service and reducing costs by 90%". They changed a single service, the Prime Video audio/video monitoring service, from a few Lambda and Step Function components into a 'monolith'. This monolith is still one of presumably many services within Prime Video.
- oaiey 3y agoThe worth here is that Amazon is writing about not going into AWS PaaS native programming (what Lambda is) because it is too expensive for them. That has some newsworthiness and the title kind of reflects that.
- dragonwriter 3y ago> The worth here is that Amazon is writing about not going into AWS PaaS native programming (what Lambda is) because it is too expensive for them. …and going to a newer AWS service (ECS), instead.
- kstenerud 3y agoThe subtitle is "The move from a distributed microservices architecture to a monolith application helped achieve higher scale, resilience, and reduce costs." And the article itself mentions the 90% cost reduction. So the title seems pretty much in-line with the original intent.
- lastangryman 3y agoMy word. I'm sort of gob smacked this article exists. I know there are nuances in the article, but my first impression was it's saying "we went back to basics and stopped using needless expensive AWS stuff that caused us to completely over architect our application and the results were much better". Which is good lesson, and a good story, but there's a kind of irony it's come from an internal Amazon team. As another poster commented, I wouldn't be surprised if it's taken down at some point.
- abrookewood 3y agoYep, expect the Lambda team to raise hell.
- oblio 3y agoEspecially now that it's on the frontpage of HN :-)))
- djtango 3y agoProbably an unpopular take and my experience is almost 10 years old, but I would be surprised to see the Amazon I worked at try to bury something like this. If the product isn't what the customer wants, it isn't what the customer wants - move on and build something the customer wants. Yes agreed there were some funny business like not selling Chromecast, but the guiding principle was generally to make things customers want...
- onion2k 3y agoDo you think the Lambda team want people to use as many of their services as possible even when it's not actually appropriate and there are better architectures and approaches available? I doubt that. They probably understand that Lambda is a good service for some things and not for others, and using it as a part of deploying things to AWS is a great idea but using it where it doesn't fit makes all of AWS look bad (in particular, hard to use and expensive.)
- adql 3y ago
- abrookewood 3y agoI'm waiting for the AWS Lambda team to talk to Marketing to get this taken down ...
- Havoc 3y agoFeels more like the initial version was a prototype not meant to scale Wouldn’t have expected prime to be pushing around images on s3
- bjornsing 3y agoIt’s expensive to store individual video frames in s3 for no good reason? Go figure…
- mparnisari 3y agoAs a former AWS employee I can almost guarantee that the person that made the original design got a promotion over it.
- IceHegel 3y agoThey put individual video frames as images in S3. That’s ubsurdly dumb. It’s like putting a frame buffer on an HDD.
- _joel 3y agoIn any normal company, not one that profits from such dumbness :)
- throwaway2990 3y agoAs a never been AWS employee I can almost guarantee you the original design was most likely simple and the use of lambdas and step functions a good choice and not expensive but the functionality grew and the cost sky rocketed. This is only normal evolution of a service.
- bhouston 3y agoI wish this was a good condemnation of microservices in a general use case but it is very specific to the task at hand. Honestly, the original architecture was insane though. They needed to monitor encoding quality for video streams so they decided to save each encoded video frame as a separate image on S3 and pass it around to various machines for processing. That is a massive data explosion and very inefficient. It makes a lot more sense that they now look for defects directly on the machines that are encoding the video. Another architecture that would work is to stream the encoded video from the encoding machines to other machines to decode and inspect. That would work as well. And again avoid the inefficiencies with saving and passing around individual images.
- amluto 3y ago> Another architecture that would work is to stream the encoded video from the encoding machines to other machines to decode and inspect. That would work as well. And again avoid the inefficiencies with saving and passing around individual images. No, that’s still a bad architecture. Bandwidth within AWS may be “free” within the same AZ, but it’s very limited. Until you get to very very large instance types, you max out at 30 Gbps instance networking, and even the largest types only hit 200 Gbps. A single 1080p uncompressed stream is 3 Gbps or so. There is no way you can effectively use any of the large M7g instances to decode and stream uncompressed video.(Maybe the very smallest, but that has its own issues.) In contrast, if you decode and process the data on the same machine, you can very easily fit enough buffers in memory, getting the full memory bandwidth, which is more like 1Tbps. If you can process partial frames so you never write whole frames to memory, you can live in cache for even more bandwidth and improved multi core scalability.
- bhouston 3y agoAh. I was thinking that the encoding machines were not bandwidth limited but rather cpu limited as they were doing expensive encoding algorithms. So I was thinking the streams were streaming out at less than real time. I figured this was better than the dual/multi encode method I think they are now relying upon when all the detection code doesn’t fit on the same machine as the encoder.
- shri_krishna 3y agoThe only time it makes sense to use edge/serverless anything is lightweight APIs and rendering HTML to end users so they get the page loaded as quickly as possible. That's the only use case good for edge. And any supporting infra that can help deliver rendered pages asap (like kv store on the edge for storing sessions, lightweight database on the edge for user profile data, queues etc). Anything that requires decent amount of processing should not live on the edge/serverless. It defeats the purpose.
- dragonwriter 3y ago> The only time it makes sense to use edge/serverless anything is lightweight APIs and rendering HTML to end users so they get the page loaded as quickly as possible. That’s the only use case good for edge. Serverless and edge aren’t the same thing.
- shri_krishna 3y agoNope. Edge is just serverless that is closer to your user to reduce the number of network hops. Both are essentially the same when it comes to technical functionality. They run on limited resources and should not be used for intensive workloads.
- dragonwriter 3y agoEdge compute is serverless, but most serverless is not edge.
- shri_krishna 3y agoYou are just getting unnecessarily pedantic here. I was talking about computing resource usage. Both are the same when it comes to resource consumption being limited. It is like saying Oracle/Postgres/MySQL/MSQL. If you say they are different in some X functionality, yeah duh they are different in X functionality. However, they are all SQL databases. Same way, Edge/Serverless is both running on limited compute resources (which is the point of the article and point I was making). Both differing in functionality X (of latency/closeness to your user) has nothing to do with either the article or my answer.
- prisonguard 3y agoThis is astonishing coming from Amazon.
- camgunz 3y agoIf you came to me with a design that included passing individual video frames through S3 instead of RAM I would honestly think you were joking. What a wild article.
- IceHegel 3y agoI’m all for big, fast, monoliths - but I’m not sure I want to hear it from the team that saved video frames to s3 in their AWS Step Function video encoder.
- vinay_ys 3y agoAWS is truly a customer first company. I been AWS customer in its early days (2006-2012) and then recently (2022-now). And they have been consistent in being customer-first. In the last year, they have proactively helped us cut our AWS spend by multiples. I'm not surprised at all by this article coming from within Amazon. Kudos for maintaining such a culture.
- deleted 3y ago[deleted]
- alexchamberlain 3y agoSomewhat interesting article, but this isn't a monolith, at least not by a microservice fanboy definition. The product (Prime Video) is still built using many business oriented services. Furthermore, this service appears to be developed and operated by a single team. That being said, there are some lessons here - there are good ideas in most design paradigms, but if you take them to the extreme, you're going to see some weird side effects. Understand the benefits and engineer a balanced solution.
- EVa5I7bHFq9mnYK 3y agoI guess what AWS sells is not servers, but software to manage them automatically, to load balance, to replicate etc. Once, in a short time, GPT can write such (pretty standard) software for you, Amazon will, too, go down.
- _joel 3y agoYou're vastly oversimplifying this, imho. It's not just being able to write something and get AI to write terraform for you (it doesn't do it all that well atm in reality, for anything complex). You can't automate the people who you need to convince to make those decisions internally, on the whole, at least :)
- iLoveOncall 3y agoSure, ChatGPT will automate in a short time what tens of thousands of top engineers have built over a decade.
- EVa5I7bHFq9mnYK 3y agoOf course not. It will help millions of small companies to write scripts so they won't need AWS anymore.
- vdchuyen 3y ago[dead]
- kiesel 3y agoThis is less an example of why serverless was bad but rather an example where using non-suitable services for tasks they were not meant for. In this case they were using AWS Step functions that are known to be expensive ($0.025 per 1,000 state transitions) and they wrote: > Our service performed multiple state transitions for every second of the stream Secondly, they were using large amounts of S3 requests to temporarily store and download each video frame which became a cost factor. They had a hammer - and every problem looked like a nail. In my experience this happens to every developer at a certain stage when he/she gets in touch with a new technology; it doesn't mean that the tech itself is bad - it depends on the scenario, though.
- INTPenis 3y agoThese days when project managers of new products seek my advice as a solutions architect I tend to suggest they create a minimally viable product that is written modularly so it can scale, but deploy it very simply on a few servers just like we used to 15 years ago. Scaling is definitely a good thing, microservices make scaling easier, no doubt about that. But an MVP rarely needs k8s level scaling, it just needs to be written well so it can scale in the future.
- smitty1e 3y agoTwo naive ideas that may be OK as a going-in position: - granularity - bandwidth negligibility Breaking everything down to a gnat's ass might improve testability, but is testability the product? Do I really need a Java stack trace that reads like an Andrew Wiles proof?[1] Maybe I do, at scale. Then there is the non-zero cost of the packet shuffling. Every edge in the aechitctural graph, not just the nodes, costs. But we just throw a waiter into the code and move on to the next line. No biggie. What was most interesting was "It also increased our scaling capabilities." Granularity was supposed to let "serverless" absorb the entire universe, I thought. At a higher level of abstraction, maybe The Famous Article is a map/reduce job: the requirements dissolved into solution, and a proper number of components precipitated out. [1] https://en.m.wikipedia.org/wiki/Wiles%27s_proof_of_Fermat%27s_Last_Theorem https://en.m.wikipedia.org/wiki/Wiles%27s_proof_of_Fermat%27...
- ThouYS 3y agohahahahahahaha
- hubraumhugo 3y agoI'll launch a consulting business focused on migrations from microservices to monoliths and from the cloud to in-house. Pricing would be a % of the saving over the first year.
- vrglvrglvrgl 3y ago[dead]
- kreco 3y agoAt this point isn't the lesson to use serverless stack for fast iterative processes then use a custom solution once you know exactly what you want? I have 0 experience with serverless/cloud. Just a thought.
- yakshaving_jgt 3y agoI think the lesson ought to be that you should start by writing one computer program and running it on one computer.
- uber1geek 3y agoI am big fan of django's apps model ... what I like to call a "Modular Monolith". Being an early engineer at most of my stints, I have build and scaled multiple startups using the approach and it has never failed me, the pitfalls of micro-services is not worth it unless absolutely necessary. I always made it a point to group by business-logic rather than separate at whatever curve ball "new-tech" throws at me.
- ishanjain28 3y ago> The main scaling bottleneck in the architecture was the orchestration management that was implemented using AWS Step Functions. *Our service performed multiple state transitions for every second of the stream*(???), so we quickly reached account limits. Besides that, AWS Step Functions charges users per state transition. This is so obvious in my head. I can't think of a single good reason where a SFN makes sense here.
- jpgvm 3y agoDead horse and all that but please just stick to Boring Tech, it is better for your mental health, not to mention your business, development velocity, defect rate, etc. Most importantly it's good for mental health though.
- noobermin 3y agoNot good for resume padding hype chasers. Especially the managerial types who never need to actual write the code.
- tnsengimana 3y agoOver engineering at its best. I tend to see microservices as a doubled edged sword and in this case, there was no need for them. Also, the pricing of AWS quickly goes up as you go from EC2 -> Fargate -> Lambda. I don't know why on earth someone would build microservices at the lambda-level.
- datadeft 3y agoTwo things: - When people use the solution -> problem path instead of problem -> proposals -> cost analysis -> solution they get what they deserve. - It is possible to optimize most infrastructures and code, it depends how much obviously but I have seen such percentages before The real question is: why didn't they chose the right stack for their problem the begin with?
- babbledabbler 3y agoBreaking things into tiny functions and putting them on many different servers incurs tradeoff costs in both complexity and compute. There is a complexity cost in having to deal with the setup, security, and orchestration of those functions, and a compute cost because if the overall system is running constantly it will be less efficient and therefore more expensive than running on one box.
- makkes 3y agoI agree on the tradeoffs you have to make. The main cost driver here was storage and traffic, though.
- babbledabbler 3y agoGood point. "communication" should also be on the list. I don't think storage is technically the tradeoff in this case even though it's S3. It's the traffic between those components that's costing them.
- samsquire 3y agoI've been having lots of thoughts lately about how you build a) a system that can respond to scale b) for the affordable price possible c) scaling infrastructure spend with income I love the anecdotes about just buying a Hetzner server which can handle a surprising amount. One of my ideas is a company that maintains an incremental infrastructure that can grow to handle extreme levels of traffic - the infrastructure itself mutates over time.
- joanne123 3y ago[dead]
- LASR 3y agoThis is not a discussion of monolith vs serverless. This is some terrible engineering all over that was "fixed". Some excerpts: > This eliminated the need for the S3 bucket as the intermediate storage for video frames because our data transfer now happened in the memory. My candid reaction: Seriously? WTF? I am honestly surprised that someone thought it was a good idea to shuffle video frames over the wire to S3 and then back down to run some buffer computations. Fixing the problem and then calling it a win? But I think I understand what might have lead to this. At AWS, there is an emphasis on using their own services. So when use cases that don't fit well on top of AWS services come up, there is internal pressure to shoehorn it anyway. Hence these sorts of decisions.
- chank 3y ago> Fixing the problem and then calling it a win? It is a win. Just not the win they're aluding to.
- ripper1138 3y agoThis is what L6 and L7 are building at Amazon, meanwhile in sys design interviews I’m being asked to design solutions for a gaming platform with 50M concurrent users.
- adql 3y ago> This is not a discussion of monolith vs serverless. This is some terrible engineering all over that was "fixed". I feel that's like 95% of the "we migrated from X to Y and now it is better"; most of improvements coming from rewriting app/infrastructure after learning the lessons with only small part sometimes being the change in tech
- tylerdurden91 3y agoTo the contrary, from my time at Amazon, I felt that developers want to use more high level AWS services. Unfortunately, the landscape of AWS services is so rapidly evolving that Amazon engineers themselves cant keep up and end up using the wrong service. As mentioned in other comments, there are options such as Fargate, that would still technically be "serverless" and still yield similar cost reductions. Not to mention that AWS also has Step functions express for "on host orchestration" use cases. This seems like a case where the original architecture wasn't very well researched and nor was the new one.
- boredumb 3y agoAWS has a great business model of people over "optimizing" their architecture using new toys from amazon and being charged through the nose for it. It's amazing how clients that are doing a few requests per second will want a fully distributed, serverless, microservice + dynamodb + s3 + athena + etc + etc, in order to serve a semi-static web app and print some reports off throughout the day and pay 10-50k a month when the entire thing could run on a few nodes and even a managed RDS instance for a thousand bucks a month. I would argue at this point that early optimization of architecture is astronomically worse than even* your co-worker that keeps turning all of your non-critical, low-volume iterable functions into lanes to utilize SIMD instructions. Some irony in my anecdotal experiences is that most places that don't have the traffic to justify the cost of these super distributed service architectures also see a performance penalty from introducing network calls and marshaling costs
- abluecloud 3y agoIt's honestly like a cult and a desire to want to "do it right" on AWS. The last few projects I've spent so much time setting up code deploy, load balancers, certificates, SES, route 53... This newest project, I've gone to heroku with everything being basically a few clicks to get setup.
- boredumb 3y agoIt seems like another case of the road to hell being paved with good intentions, most places want/need redundancy and some managed devops and so a few ec2 instances and managed RDS is affordable enough and checks a lot of boxes, but after people start down this path it seems almost irresistible to start drilling down into managed kubernetes, spark jobs, to start ingesting some events we'll just introduce glue, and that plugs right into S3, and look how easy it is to plug in athena, add some quick alerting with cloudwatch, and the next thing you know you're vendor locked and having to hire a full time devops person with AWS experience to configure, manage and keep on top of it all.
- steveBK123 3y agoSo guys we need Lambdas + Step Functions + SES + SQS + SNS + MSK + AWS Batch + S3 + Lakeformation + Cloudformation + Athena + EMR + Redshift + Aurora + SageMaker + Cloudtrail + Codepipeline + maybe some EC2s to run AWS CLI on them. Don't forget to configure Route53, VPC, IAM and an ELB. Great - ready to start writing your app now? Oh wow one of those components as configured with the other components isn't behaving as expected - time to contact AWS support! Cynically I think CTOs see all this stuff and think they'll turn all their expensive on-shore devs into cheaper DevOps because AWS is magic and you don't need to write hard app code anymore. I'd counter that AWS forces expensive on-shore devs into having to wear an entire new hat and be half a DevOps engineer to figure out how to make their code work on this alphabet soup instead of a Linux server.
- revskill 3y agoLambda , Steps functions,... is just pure marketing scam to me , because the price is ridiculous too high for 99% of real world use case. They're though good enough to deliver an MVP quickly, but that's all about it.
- stuaxo 3y agoLocality of reference matters. It's fine to split things up, but we have to be careful how we do it + aware of the overheads.
- dfgdrgf 3y ago[flagged]
- alpos 3y ago"We built a video stream processor by splitting every 1080p+, multi hour long, 30-60fps video into individual images and copying them across networks multiple times." Not surprising that didn't go will. This strikes me as a punching bag example. Anyone who has worked with images, video, 3d models, or even just really large blocks of text or numbers before (any kind of actually "big data") knows how much work goes into NOT copying the frames/files around unnecessarily, even in memory. Copying them across network is just a completely naive first pass at implementing something like this. Video processing is very definitely a job you want to bring the functions to the data for. That is why graphics card APIs are built the way they are. You don't see OpenGL offering a ton of functions to copy the framebuffers into ram so you can work on them there only to copy them back to the video card. And if you did do that, you will quickly find out that you can be 10x to 100x more efficient by just learning compute shaders or OpenCL. You could do this in a distributed fashion though, but it would have to look more like Hadoop jobs. I predict the final answer here, if they want to be reasonably fast as well, is going to be sending the videos to G4 instances and switching the detectors over to a shader language. In general, if the data is much bigger than the code in bytes, move the code, not the data. IO is almost always the most expensive part of any data processing job. If you're going to do highly scalable data processing, you need to be measuring how much time you spend on IO versus actually running your processing job, per record. That will make it dead obvious where you should spend your optimization efforts.
- Guid_NewGuid 3y agoTo be fair it is somewhat a punching bag example but I think what people are reacting to, but maybe not articulating well, is the presumption for microservices by the powers-that-be. Of course the only rational take on monoliths versus microservices is "use the right tool for the job". But systems design interviews, FAANG, 'thought leaders', etc basically ignore this nuance in favour of something like the following. Question: design pastebin (edit, I of course mean a URL shortener not pastebin) Rational first pass but wrong Answer: Have a monolith that chucks the URL in the database. Whereas the only winning answer is going to have a bunch of services, separate persistence and caching, a CDN, load balancing, replicas, probably a DNS and a service mesh chucked in for good measure. I think this article shows that this is training and producing people who can't even think of the obvious first answer they have been so thoroughly indoctrinated.
- huksley 3y agoAWS Step Functions are bad for so many reasons. Scaling, pricing, developer experience, etc. It is clearly made by people who don't really understand (or does not care) how distributed workflows work. And pricing are prohibiting to run it at scale. In my opinion it should be free to use, provided you glue together other AWS services with it.
- bob1029 3y agoI think serverless has its place, but this problem doesn't seem like a fantastic fit. We are looking into serverless as a way to exhibit to our customers that we are strictly following certain pre-packaged compliance models. Cost & performance are a distant 2nd concern to security & compliance for us. And to be clear - we aren't necessarily talking about actual security - this is more about making a B2B client feel more secure by way of our standardized operating model. The thinking goes something like - If we don't have direct access to any servers, hard drives or databases, there aren't any major audit points to discuss. Storage of PII is the hottest topic in our industry and we can sidestep entire aspects of The Auditor's main quest line by avoiding certain technology choices. If we decided to go with an on-prem setup and rack our own servers, we'd have to endure uncomfortable levels of compliance. Put differently, if you want to achieve something like PCI-DSS or ITAR compliance without having to covert your [home] office into a SCIF, serverless can be a fantastic thing to consider. If performance & cost are the primary considerations and you don't have auditors breathing down your neck, maybe stick with simpler tech.
- InvOfSmallC 3y agoOverall, like it's stated in the article, it would be a case-by-case choice what to use. My experience tells me it's always a good idea to start with the monolith but I don't know much about PII to tell you your idea is over-engineered. I feel there are better ways though. Also because you don't need to use Lambda to not be on-prem EC2 is enough.
- jdub 3y agoI work in streaming video, specialise on AWS, and have enjoyed using Step Functions for certain (non-video) projects. I am _astonished_ that Step Functions + S3 was even considered as a starting point for defect detection in streaming video. Astonished.
- chrismsimpson 3y agoThis article is going to keep me employed for some time yet
- time4tea 3y agoMicroservices are just an architectural pattern, and like all patterns there are places where they are highly appropriate, and others where they are inappropriate. Same for cloud, same for <pattern> If everything is a hammer you'll hurt your thumb/hand/arm. At least now (for some time) the pattern is named, so broadly when talking about this sort of thing, the name conjures up the same/similar image in everyones heads. There are all sorts of inputs to the choice of architectural patterns, including budget, scalability (up and down), criticality, security, secrecy, team skills and knowledge, preference, organisational layout, organisation size, vendor landscape, existing contracts, legal jurisdiction ....
- bcoughlan 3y agoI quite like the idea of viewing run cost as architecture fitness function https://www.thoughtworks.com/radar/techniques/run-cost-as-architecture-fitness-function https://www.thoughtworks.com/radar/techniques/run-cost-as-ar.... If your architecture has a high cost to develop, test and run when a cheaper architecture meets your needs, it's a sign that you have overengineered. In my experience there is an order-of-magnitude increase in complexity by adopting microservices that only starts to pay off when your org and user base are huge.
- retrac98 3y agoYou can read more of these sorts of posts at https://www.microservice-stories.com/ https://www.microservice-stories.com/
- deterministic 3y agoMicro-services is BS invented by cloud providers to solve problems you don’t have at 10x the cost. The worst software systems I have ever seen were micro-services. One of them is more than 20 years old. The WTF count per minute is exponential.
- sjinta 3y agoI think using "Monolith" for what they ended up with is badly chosen. Basically they just made a service less granular (or less micro, if you will).
- terom 3y ago> The second cost problem we discovered was about the way we were passing video frames (images) around different components. To reduce computationally expensive video conversion jobs, we built a microservice that splits videos into frames and temporarily uploads images to an Amazon Simple Storage Service (Amazon S3) bucket. Defect detectors (where each of them also runs as a separate microservice) then download images and processed it concurrently using AWS Lambda. However, the high number of Tier-1 calls to the S3 bucket was expensive. Taking "malloc for the Internet" [1] a bit /too/ literally there. [1] https://aws.amazon.com/blogs/aws/eight-years-and-counting-of-cloud-computing/ https://aws.amazon.com/blogs/aws/eight-years-and-counting-of...
- debdut 3y agoWhat! They changed the title. Tells you something
- tylerdurden91 3y agoI think what most people are missing here is that they used AWS Step Functions in the wrong place. Part of the blame here is that in over enthusiasm of trying to get more users, AWS doesn't properly educate customers when to use which service. Worse, for each use case AWS has about dozens of options making the choice incredibly hard. In this case, they probably should have used Step Functions Express, which charges based on duration as opposed to number of transitions and they're looking for "on host orchestration" like orchestrate a bunch of things which usually are done in small time and are done over & over many times. Step functions is better when workflows are running longer, and exactly once semantics are needed. Link for reading differences between Express & standard step functions: https://docs.aws.amazon.com/en_us/step-functions/latest/dg/concepts-standard-vs-express.html https://docs.aws.amazon.com/en_us/step-functions/latest/dg/c.... This also exemplifies the fact that I learned while being at Amazon & AWS that Amazon themselves dont know how best to use AWS. This being one of the great examples. I'll share 1 more: - In my team within AWS, we were building a new service, and someone proposed to build a whole new micro service to monitor the progress of requests to ensure we dont drop requests. As soon I mentioned about visibility timeout in SQS queues, the whole need for the service went away. Saving Amazon money ($$) & time (also $$). But if I or someone else didn't mention, we would have built it. I dont think serverless is a silver bullet, but I don't think this is a great example of when not to use serverless. It helps to know the differences between various services and when to use what. PS: Ex Amazon & AWS here. I have nothing to gain or lose by AWS usage going up or down. I'm currently using a serverless architecture for my new startup which may bias my opinions here.
- tylerdurden91 3y agoWorth mentioning as mentioned in other comments that moving video data around at that scale was a bad choice to begin with. They could have considered fargate and avoided moving the data around so much as well and realized similar reductions in cost. So the wins are not really coming from moving to monolith as much as they're coming from optimizing unnecessary data transfers. If the article said fargate, which is technically still serverless we could have avoided a whole microservice vs monolith debate or serverless vs hosts/instances debate.
- tyingq 3y agoSeems somewhat curious that they didn't at least include Fargate. Feels like they jumped all the way from the typical overengineered setup into using AWS in a way that's very close to just "I need virtual machines".
- tylerdurden91 3y agoAbsolutely. Neither fargate nor step functions express. Seems like they did not evaluate all the options before making the jump.
- recursivedoubts 3y agohttps://grugbrain.dev/#grug-on-microservices https://grugbrain.dev/#grug-on-microservices > grug wonder why big brain take hardest problem, factoring system correctly, and introduce network call too > seem very confusing to grug
- Cthulhu_ 3y agoThey basically underestimated the cost of moving millions of small files to and from S3; it kinda makes sense if they want to save those images for a long time, but in this case it was for semi-real-time error detection, which is much faster to do in-memory.
- Alifatisk 3y agoGuess who happy DHH was reading this
- Alifatisk 3y agohow*
- cmrdporcupine 3y agoI'm happy to see -- in the discussion here-- the continued backlash against microservices and the deleterious effects it has had on software complexity, and data modelling. But I think it's interesting that if we took a time machine back to 2014 or 2015 the tone here would be quite different, and microservices were all the rage on this forum as I recall. I like to hope that the industry learns from its failed trends, but I'm now old enough to see this is rarely the case.
- basilgohar 3y agoI wonder if this is, in some way, a kind of signalling of where AWS wants to go – maybe they want to shift more towards dedicated hosting rather than all of these separate services?
- zenbowman 3y agoShipping around individual video frames between components is really an astonishingly bad idea. Microservices seem to be a decent idea with a terrible name. The idea of running services that are small enough that they can be managed by a single team makes sense - it enables each team to deploy their own stuff. But if you break things down further, where you need multiple "services" to perform a single task, and you have a single team managing multiple services - all you do is increase operational & computational overhead.
- dahwolf 3y agoRarely will I defend Amazon in anything, but I'll make an exception. In my experience, AWS/Amazon people do not force you or even direct you to a particular architectural choice. They are relatively indifferent about it. Instead, trend-driven architectures seem to come from the tech community themselves. It's the customers often making the wrong choice.
- finikytou 3y agothey just coded a step functions monolith...