44 ms·
Mistakes I've Made in AWS
- sebazzz 5y agoIn summary: Either overprovisioning, or not realising every extra CPU cycle or I/O operation costs extra money. This is, of course, the real way "the cloud" makes money. Carefully tuned, it can no doubt be cheaper than do-it-yourself, however, it is also quite easy to make a lot of costs.
- gizdan 5y agoContrary to popular believe, the case for going to the Cloud isn't cost saving, it is flexibility and value for money. It'll likely cost you around the same if you run it in a DC, but you won't have features like auto scaling, increased security, and much more.
- luckylion 5y agoAbout the same? Last I checked for our somewhat static work load on a bunch of webservers, AWS would be x10 in pricing. Not to mention that you need someone who has deep AWS knowledge and experience to manage your system, just like you need someone who manages your dedicated servers in a DC. It's great for workloads that fluctuate extremely, or require massive scaling in very short time. Not sure about the increased security. If you run your images on EC2, it's still up to you to not mess up the config.
- sokoloff 5y agoLightsail is more reasonably priced for a lot of simple web serving use cases.
- fermentation 5y agoIt’s also super easy to use. I have an instance hosting a game server for friends. I might be wasting money since the server sits idle about half the time though.
- maccard 5y agoWe're actively planning an aws workload right now, and with reserved instances for the baseline workload, the pricing is closer to 1.5-2x, but the cost savings of only needing to scale up for a couple of hours per week make up for that. Yes it would be cheaper to run out own infra for the baseload and burst into aws, but that adds operational load onto the development team, which defeats the purpose of going with AWS in the first place
- jimmaswell 5y agoHow many workloads actually fluctuate so extremely and unpredictably?
- genewitch 5y agoThere was a coupon site that I moved to AWS because all the customers were in EU, but our servers were in Los Angeles. I implemented it in two t1.micro servers, db+nginx. Every night at 10PM in LA the first week I'd get a slew of notifications from monitoring that the CPUs were pegged out. I think I put squid or something on the nginx instance so that the most common stuff came out of the proxy instead of having to be php-unblinkered. However another thing that could have been done is to use a slightly larger instance for the DB to get more ram for cache, and then use a single front end until 9PM, spin another up, and spin one down at like 1AM. This is obviously a trivial example, but lots of sites and apps get everyone refreshing at once maybe once or twice a day, and that's when scaling helps, at least from a "costs $X per hour, so don't use it unless we need it". Loads were under 20% the other 23 hours during the day, averaged.
- Viliam1234 5y agoFrom the perspective of a developer, flexibility is a double-edged weapon. Before cloud: we have database quota of a few gigabytes, and once in a few years we need to justify to management why the quota should be doubled. After cloud: whenever we add a new table, or a new column, or import lots of data, the invoice slightly increases, the management notices, and we need to justify the extra megabytes.
- hughrr 5y agoBiggest mistake I’ve made: Shifting any non trivial infrastructure into AWS verbatim is always more expensive than running it yourself. You need to rearchitect it carefully around the PaaS services to make a cost saving or even break even. An extreme example of this is it cousin who works for a small dev company doing LOB stuff. They moved their SQL box into EC2 and it’s costing more to run that single RDS instance than their entire legacy infra cost was per year. I’d still rather use AWS though. The biggest gain is not technology but not having to argue with several vendor sales teams or file a PO and wait for finance to approve it. All I do is click a button and the thing’s there.
- macpete42 5y agoI can confirm that: cloud helps to evade the incompetent sales and infrastructure teams in many companies. Saving money never works once your product scales out.
- AndrewDucker 5y agoIt's always more expensive to have someone else run your infrastructure than to do it yourself unless it's something you only use intermittently. If you need 5 seconds of compute time per day then running that as a Lambda makes perfect sense. If you need a database server that's available 24/7 then I can't see how hosting that on Amazon could be cheaper. (Unless you're employing a full time ops person to look after that one server, in which case you'll have to do your own maths.)
- hughrr 5y agoYep that. Lambda is a massive win for me personally. I have some scraping and processing stuff that runs daily. Costs me $0.60 a month to run it even outside of free tier which is less than a cheap DO or linode box and I don’t have to look after the OS.
- AndrewDucker 5y agoSame. Powershell script that collects links from Pinboard and posts them to my blog. It would be massively overkill to run a whole server for that. (And Microsoft charges me about £0.15 per month)
- danjac 5y agoI've made it a habit to absolutely avoid any and all AWS services for any side projects, unless it's on the employer's dime. I'd rather pay a bit more per month for a flat-fee Digital Ocean droplet. Maybe I'll end up paying a few dollars more than I would with the equivalent AWS setup, but I'll rest easy knowing I won't get a surprise bill thanks to the opaque and byzantine billing. I mean, there are consultancies whose entire premise is expertise on AWS billing, so the chance of AWS newbie-me running up many thousands because I forgot to switch off service A or had the wrong setting for service B is non-zero. And the general advice is "don't worry, call their customer support and they'll refund you". Um, seriously? If I want to spend a morning on hold to deal with a huge unplanned bill I'll call my local tax office, thank you. Which sucks as I learn best by building things in my spare time, but AWS makes that learning process a bit more stressful than I'd prefer.
- tomxor 5y agoPretty much summarises my decision to use Linode, at a small company AWS presents a bigger monetary risk and drain on precious developer time and mental overhead than relatively small savings it might return at smaller scales... I also actually like Linode as a company and enjoy using their services and management interface; Amazon is challenging to be positive about.
- wly_cdgr 5y agoAlso use Linode, they're great. Their docs are a treasure Seems absolutely insane to use AWS for small personal/learning projects (unless the goal is to learn AWS for career purposes, I guess). It'd be like using Unreal Engine to make your 2d indie game Always use the smallest and simplest solution that'll do the job. Simple solutions are not just as good for simple jobs...they're better
- remram 5y ago> It'd be like using Unreal Engine to make your 2d indie game Except that would be free.
- helsinkiandrew 5y agoNothing for me compares to the time I purchased 2 reserved EC2 instances for about $5K on my personal account rather than companies. I can still remember that sinking feeling as I realized what I'd done. Amazon refunded the next day.
- stingraycharles 5y agoIt’s incredibly easy to spend a lot of money on the cloud, indeed. I remember using Google Cloud’s translate API on a bunch of documents — it took several hours for the bill to pop up at $1500. This was a hobby / personal project of mine, Google did not refund it, because of course I should have read the pricing more carefully.
- dotancohen 5y agoThis is the advantage with AWS, they _will_ refund mistakes. I've seen it happen twice, and both times were resolved quickly with a rep on the phone. I've also once had an issue with my own personal account. Five minutes with a rep on the phone saved not my bank account, but my website and hosted services, because my credit card was cancelled and it would be another few months before I could get another.
- nobleach 5y agoAmazon as a company tends to side with the customer. Their whole mantra is that it's not worth chasing after x* amount of dollars. Now repeat offenses? No, you're not getting away with using their services for free. (You mention twice, but I imagine 4 or 5 times, and they're going to fault you without escalating the issue) *within reason... you're not going to serve up an app all month long and skip out on a million dollar bill.
- dehrmann 5y ago> Amazon as a company tends to side with the customer. Completely agree. Google might be learning parts of this with GCP, but historically, customer-obsessed isn't in Google's DNA.
- unglaublich 5y agoBut what mistakes did he make? Did he screw up the bill? Did he fail to keep services available? I only read facts about the ins and outs of AWS' billing and credits system.
- weird-eye-issue 5y agoIf you run out of CPU or IOPS burst balance then your system will suddenly slow to a crawl and it can easily cause downtime or in the case of background jobs it will cause long queues or never ending jobs. Learned that the hard way, a couple times. One time I optimized DB access which fixed the IOPS usage, and then that caused more CPU usage on the app servers which caused them to run out of CPU burst... Fun times. Switched from one burst issue to another.
- sethammons 5y agoAnd that is scaling systems. Open up one bottleneck to discover the next. Rinse and repeat to gain experience.
- weird-eye-issue 5y agoI disagree this is about scaling systems. This is actually more about using the wrong instance type. If this was on a typical VPS it wouldn't have ever happened. The baseline CPU level on these burst instances is so low that for any long running task using even like 40% CPU it gets throttled so hard it brings everything down. I would have been totally fine if this was a $5 DigitalOcean VPS
- lanstin 5y ago“It is easier to push the bottle neck to another piece of the system than to remove it.” http://www.faqs.org/rfcs/rfc1925.html http://www.faqs.org/rfcs/rfc1925.html
- jcims 5y agoI feel like large enterprises primarily see AWS as a way to outsource capital expenses.
- jnieminen 5y agoAWS and Azure are a permission to spend.
- lloydatkinson 5y agoAre they really though? A serverless event driven architecture system I’m working on literally costs less than £10 a month on Azure. Running full blown VMs instead of cheaper more appropriate technologies like containers or functions will always cost more.
- TriNetra 5y agoAs per our calculation for CloudAlarm [0], as we reach a few hundred users, it'd be cheaper to use a dedicated instance than serverless (Azure Functions) design. So it may vary from system to system depending the amount of work you perform for each user. 0: https://cloudalarm.in/ https://cloudalarm.in/ – btw, you may wish to have daily budgeted pace based alerts using it – to inform you when the usage spikes up (much faster than Azure's consumption threshold based alerts).
- StopHammoTime 5y agoThis is literally the main reason a lot of companies use AWS. In Australia, it is very hard for Government Departments to get capital expenses approved for infrastructure as it requires a lot of rigmarole. However, once you’re in AWS its OpEx, who cares as long as you don’t break the budget too soon before EOFY.
- thanatos519 5y agoSo basically this ... «Ah, I see you have the machine that goes ping. This is my favorite. You see we lease it back from the company we sold it to and that way it comes under the monthly current budget and not the capital account.» https://www.youtube.com/watch?v=tKodtNFpzBA https://www.youtube.com/watch?v=tKodtNFpzBA
- nickjj 5y agoMy favorite billing mistake was forgetting to delete an unused elastic IP address and then realizing I was being charged $34 / month for 2 months just to have it exist while doing nothing. Edit: It's exactly $33.62 and I was mistaken on what caused it. It came from having a NAT Gateway just idling which is $0.045 per hour x 747 hours = $33.62 on us-east-1. I know it's not the biggest mistake ever, but these things creep up on you when you use CloudFormation and it continuously fails to delete resources so you're left having to manually trace through a bunch of resources. It's easy to leave things hanging.
- jrochkind1 5y agounused Elastic IP pricing looks to me like $3.60/month on their pricing page. ($0.005 per hour). What am I missing to get to $34/month? (Or did you have 10 of em?) https://aws.amazon.com/ec2/pricing/on-demand/#Elastic_IP_Addresses https://aws.amazon.com/ec2/pricing/on-demand/#Elastic_IP_Add...
- nickjj 5y agoThanks, I edited my post to correct it. It was a single NAT gateway that's $33.62 / month.
- noir_lord 5y agoI nearly made myself a very nice footgun not long since. So MediaConvert (video transcoding), direct s3 upload to s3 bucket, bucket fires event to my application, my application builds the job and submits it to media convert with the output bucket as the destination. Straight forward enough, unless you happen to be copying a config tired and put your input/output buckets as the same bucket... Fortunately previous-me was paranoid enough to have put in an if check and die if they where the same but otherwise that could have cost a lot of money.
- swyx 5y agowhy would MediaConvert not build that if check in? perhaps a good feature request for them.
- noir_lord 5y agoBecause you can write back to the same bucket at different prefixes if you want to. It's simply simpler to split the buckets in my case. I added further checks to not only check the bucket made sense but also that the inbound and outbound had the correct prefixes. So if another person does the same it'll catch both ways.
- fukmbas 5y agoMistake #1: using AWS lol
- arno1 5y agoDiscover Akash Network! Censorship-resistant, permissionless, and self-sovereign, Akash Network is the world’s first open source cloud. It's at the early stages, the amount of deployments is steadily growing! Soon GPU compute and persistent storage! As well as you can already become a provider and earn AKT tokens (which are neat, driven by the Cosmos based blockchain) https://akash.network https://akash.network https://akashlytics.com/price-compare https://akashlytics.com/price-compare
- tedk-42 5y agoFew easy ones as well: 1) Terminating instances that had ephemeral disks with stuff you needed while thinking the EBS volumes would remain 2) Leaving NAT gateways lying around or ELBs that do nothing and have no instances attached. 3) Public S3 buckets - arguably the most common one that can lead to security incidents 4) Debugging security groups/Network ACLs and straight up break networking for something without knowing it. Reverse of that would be you want to fix something quickly and open 0.0.0.0/0 to everyone and never get around to tightening up the firewall later on.
- jnieminen 5y agoI was playing with the Azure "free" tier. Even I tried to be extremely careful with it, after a while noticed that I had left a storage blob for a VM hanging around and some external IPv4 address. I will continue to use Hetzner online for my own stuff instead running this on "public cloud".
- hacker_newz 5y agoI don't understand how anyone could complain about public S3 buckets. You have to go out of your way to do this.
- igammarays 5y agoAWS is complexity-as-a-service. This is why, as a one-man company, I went baremetal[1]. One flat price, screaming fast performance, and massive scalability if you get a beefy enough machine[2]. I don't have time to fiddle with k8s, try to figure out AWS billing/performance tradeoffs, or deal with untraceable performance issues due to noisy neighbours and VM overhead. My disaster recovery plan is a simple DB dump script to S3, and I know I can get another baremetal server up and running in less than 20 minutes. [1] with IBM Cloud 1 year free startup credits [2] Let's Encrypt and StackOverflow run their entire databases on a single beefy baremetal machine. https://letsencrypt.org/2021/01/21/next-gen-database-servers.html https://letsencrypt.org/2021/01/21/next-gen-database-servers...
- pibefision 5y ago+1 also it's easy to use Docker containers and Traeffik as reverse proxy to manage many services.
- tomerbd 5y agoWhich scripting or which infra do you use for automatic installation/configuration of your server?
- chrisandchris 5y agoNot OP, but I did the same and I use - Ansible for the low-level stuff (like network, mounts, iSCSI, configuration files) - Terraform for high-level stuff (like DB users) In my case, as I have several services that use a lot of RAM running, I couldn‘t afford The Cloud but can easily afford a colocation. I don‘t mind the maintenance (it‘s a couple hours each month) and I don‘t care much if services are down a few hours. If you need something running 24/7 with 99.9%, colocation will be more expensive just because of the human you need.
- candiddevmike 5y agoWhy wouldn't you use Ansible for the high level stuff? It can easily manage DBs and you wouldn't need another tool.
- daneel_w 5y ago"Technically they are a smidgen slower than Intel for certain workloads." In my experience, after migrating several servers with quite varying workloads, they're faster than Intel - and more than a smidgen. Just as is the general case with current AMD Ryzen vs Intel.
- defaultname 5y agoOn a price sensitive project I almost exclusively used spot instances at a dramatically reduced price over on-demand. It forced me to built high availability elements into the design at the outside, though ultimately spot instances got shut down no less frequently than my experience with on demand maintenance and individual machine outages. Obviously mileage will vary, but going in I was under the impression that spot instances were on the knife's edge, when with a decent pricing strategy they're as robust as on demand at a fraction of the cost.
- doomslice 5y agoWe use GCPs equivalent of spot instances (preemptibles) to great effect as well. It actually works better at larger scale since a smaller % of your machines get preempted at a given time.
- noogle 5y agoSpot instances for GPU are shutdown within hours. As frequently touted in favor of AWS, engineer time is the most important thing. The time to adapt the code to frequent failure, and the delays in getting the results, costs money as well, negating the financial saving from spot instances.
- defaultname 5y agoDesigning to remain robust in the face of failures is compulsory for any project of any significance. Or at least it should be, though a lot of projects go on a wing and a prayer that nothing will go awry and "save" those engineering hours until a catastrophe at some future point. It basically just prioritized what already should be a priority. I have no doubt that fringe/niche instances have more competitive spot behaviors, though how you set your bid range dramatically impacts how you survive through competition, but I had vanilla instances last for literal years (note that by default the spot requisition has a lifespan of one year so you have to modify that) at per hour pricing somewhere in the range of 1/5th on demand. But mileage will vary. I don't use those spot instances anymore as my projects are much better financed now, and I have significant compute on other platforms including bare metal in colocation facilities. However when I did I stayed silent about it, feeling almost like it was a secret that would be ruined if others knew about it.
- StratusBen 5y ago[Disclosure] I'm Co-Founder and CEO of http://vantage.sh/ http://vantage.sh/, a cloud cost platform for AWS. Previously I was a product manager at AWS and DigitalOcean. Since the author and so many people are commenting about AWS costs (and in particular, choosing cheaper EC2 instances and EBS volumes) I thought I'd mention that Vantage has recommendations that look to tell you for these exact things so you don't get tripped up / spend more than you have to. If you have "antiquated" EC2 instances or EBS volumes, Vantage will give you a recommendation for which instance to switch to and how much money you'll save. The first $2,500/month in AWS costs are also tracked for free so people get a lot of value out of the free tier and can save significant parts of their bills when developing on AWS.
- 7sidedmarble 5y agoRespect that you are all about that grindset for your product in this thread, but it's also a little insane that you need a third party tool to make sense of what's going on in AWS. I'm a bit of a GCP fan, and while it's billing is also arcane, it think it is just a little bit easier to understand and better laid out. For bread and butter stuff like regular VPSs though, AWS is often a little cheaper. But GCPs other cloud offerings are occasionally very respectably priced.
- swyx 5y agoevery $X00 billion dollar business is big enough that third party tools will always be desired because the default experience wont be good enough for some part of the market. question is whether or not that part is big enough to warrant its own venture scale business, as with Vantage :)
- smoldesu 5y agoI'd frankly just prefer to use a VPS. The fact that I need to have a payment stack alongside my technology one is just ridiculous to me.
- deleted 5y ago[deleted]
- lysecret 5y agoOk im going to admit to a mistake revolving around NAT gateways and Lambdas. So, i basically wanted to connect a Lambda to a Postgres / RDS database, for that I had to put into a private VPC, but the lambdas still had to talk to the world (a lot) so i just put a nat gateway around it no biggy. Well, end of the story on one day i produced 2000 Euro in cost for the Nat gateway haha
- hacker_newz 5y agoWhy would you need a nat gateway for a lambda?
- lysecret 5y agobecause it had to talk to a database which was in a private VPC, and at the same time make http requests.(in my original setup i changed that :D)
- wly_cdgr 5y agoHeh, I like how Amazon literally took the boost mechanic from arcade racing games for the CPU credits in T2/T3
- jbverschoor 5y agoMost common made mistake: assuming that your data is safe on an EC2 instance (ephemeral storage)
- steveBK123 5y agoOn billing.. they will never do it, but on smaller accounts they could build trust by offering some sort of "prepaid" mode like cell phone services do at the low end. That is - you deposit $X in your account, and AWS nukes your live services if you breach it. The worst that ever happens is you are out sunk cost of the $X you had already deposited.
- calmlynarczyk 5y agoThis is more just "missed optimization opportunities in EC2" than a statement about mistakes in AWS as a whole. If you want to talk systemic AWS mistakes you can make, we accidentally created an infinite event loop between two Lambdas. Racked up a several-hundred-thousand dollar bill in a couple of hours. You can accidentally create this issue across lots of different AWS services if you don't verify you haven't created any loops between resources and don't configure scaling limitations where available. "Infinite" scaling is great until you do it when you didn't mean to. That being said, I think AWS (can't speak for other big providers) does offer a lot of value compared to bare-metal and self-hosting. Their paradigms for things like VPCs, load balancing, and permissions management are something you end up recreating in most every project anyways, so might as well railroad that configuration process. I've experienced how painful companies that tried to run their own infrastructure made things like DB backups and upgrades that it would be hard to go back to a non-managed DB service like RDS for anything other than a personal project. After so many years using AWS at work, I'd never consider anything besides Fargate or Lambda for compute solutions, except maybe Batch if you can't fit scheduled processes into Lambda's time/resource limitations. If you're just going to run VMs on EC2, you're better off with other providers that focus on simple VM hosting.
- itisit 5y ago> Racked up a several-hundred-thousand dollar bill in a couple of hours. Not doubting you, but curious how you hit such a high figure. Can you walk through the math? Are we talking trillions of <10ms requests?
- mfrye0 5y ago> If you want to talk systemic AWS mistakes you can make, we accidentally created an infinite event loop between two Lambdas. Racked up a several-hundred-thousand dollar bill in a couple of hours. I did more or less the same thing, but with a 3rd party webhook. The bill almost killed my company.
- cutemonster 5y ago> The bill almost killed my company. You had to pay although it was a mistake?
- Kiro 5y agoSlightly OT: I love Forge but recently I've started using it for my non-PHP projects which feels... wrong. Are there any similar services that are more agnostic?
- AaronNewcomer 5y agoI’ve recently been thinking about this as well.
- dncornholio 5y agoMistakes? How about the flaws of that what is AWS and there terrible, terrible pricing system that rewards them for your mistakes.
- projectramo 5y agoMy biggest mistake: years ago I ended pushing personal credentials to GitHub at night and waking up to a several thousand dollar bill in the morning. Changed credentials and cancelled all the running instances only to find that I’d missed some. It was resolved by the afternoon.
- judge2020 5y agoThankfully GitHub now runs secret scanning and AWS is a partner. If you did this today AWS will revoke the key before malicious scanners find it. https://docs.github.com/en/code-security/secret-scanning/about-secret-scanning https://docs.github.com/en/code-security/secret-scanning/abo...
- mfrye0 5y agoOne of the biggest mistakes I made is not exploring spot instances and reserved instances earlier. I cut my bill by 70-80%% after paying full price for years... If you have an active web server or backend workers with fairly short jobs, spot instances will work for you.
- thanatos519 5y agoDoes the 'cpu credits' stuff apply to spot instances too? I have been thinking of shortening my animation render time with spot instances, but it only makes sense if I can run every core at 100% for the entire life of the instance.
- mfrye0 5y agoI believe so, if it's a T series. For 100% CPU usage, it's cheaper to just get a non T series instance like a M.
- zackmorris 5y agoI view AWS as a study in doing everything the "bare hands" way. Here are some examples of the old sysadmin ways of doing things vs the modern "web" way: * regions -> self-balancing algorithms like RAFT * roles/permissions -> tokens * IP address filtering -> tokens * CPU clusters -> multicore/containerization/Actor model * S3 -> IPFS or similar content-addressable filesystems It's not just AWS having to deal with this stuff either: * CORS -> Subresource Integrity (SRI) * server languages (CGI) -> Server-Side Includes (SSI) * Javascript -> functional reactive, declarative and data-driven components within static HTML * async -> sandbox processes, fork/join, auto-parallelization (seen mostly in vector languages but extendable to higher-level functions) * CSS -> a formal inheritance spec (analogous to knowing set theory vs working around SQL errata) I could go on forever but I'll stop there. We are living at a very interesting time in the evolution of the web. I think that web dev has reached the point where desktop dev was in the mid-1990s and is ripe for disruption. No disruption will come from the big companies though, so this is your chance to do it from your parents' basement!
- hacker_newz 5y agoDoes this make sense to anyone?
- physicles 5y agoBurst CPU and IOPS has bitten me a couple times over the years. In fact, it’s basically the sole cause of nearly all our downtime in recent history. That’s frustrating. I get that it’s a technical solution to the problem of resource utilization at scale, but they could’ve spent some time making it easier to observe — for example, rescale the CPU or IOPS graphs so that 100% is your max sustained budget, and anything over 100% eats into your quota.
- awinter-py 5y agothere should be a social media platform just for people to list their mistakes