15 ms·
AWS S3: Sometimes you should press the $100k button
- wodenokoto 5y agoI've never been in this situation, but I do wish you could query files with more advanced filters on these blob storage services. - But why SageMaker? - Why do some orgs choose to put almost everything in 1 buckets?
- korostelevm 5y agoFor many at orgs like this, SageMaker is probably the shortest path to an insane amount of compute with a python terminal. Why single bucket? Once someone refers to a bucket as "the" bucket - it is how it will forever be.
- tyingq 5y ago>Why do some orgs choose to put almost everything in 1 buckets? The article seems to be making the case it's because the delimiter makes it seem like there's a real hierarchy. So the ramifications of /bucket/1 /bucket/2 versus /bucket1/ /bucket2/ aren't well known until it's too late.
- charcircuit 5y ago>So the ramifications of /bucket/1 /bucket/2 versus /bucket1/ /bucket2/ aren't well known until it's too late. What's the difference?
- musingsole 5y agoIn the choice between a single bucket with hierarchical paths versus multiple buckets, there's a long list of nuances between either strategy. For the purposes of this article, you can probably have more intuitive, sensible lifecycle policies across multiple buckets than you can trying to set policies on specific paths within a single bucket. Something like "ShortLifeBucket" and "LongLifeBucket" would allow you to have items with similar prefixes (something like a "{bucket}/anApplication/file1.csv" in each bucket) that then have different lifecycle policies
- 8note 5y agoThere's a lack of searchable blogs and recommendations for how many buckets you need, and how much stuff belongs in one. Got any recommended literature?
- liveoneggs 5y ago1 athena? 2 some jobs make a lot of data
- akdor1154 5y ago> But why SageMaker? You could ask the same thing of most times it gets used for ML stuff as well. > Why do some orgs choose to put almost everything in 1 buckets? Anecdote: ours does because we paid (Multinational Consulting Co)™ a couple of million to design our infra for us, and that's what the result was.
- Tehchops 5y agoWe’ve got data in S3 buckets not nearly at that scale and managing them, god forbid trying a mass delete, is absolute tedium.
- amelius 5y agoMass delete also takes an eternity on my Linux desktop machine. The filesystem is hierarchical, but the delete operation still needs to visit all the leaves.
- sokoloff 5y agoIs S3 actually hierarchical? I always took the mental model that the S3 object namespace within a bucket was flat and the treatment of ‘/‘ as different was only a convenient fiction presented in the tooling, which is consistent with the claim in this article.
- cle 5y agoThis is mostly correct, with the additional feature that S3 can efficiently list objects by "key prefix" which helps preserve the illusion.
- sokoloff 5y agoFollowup question: Is there something special about the PRE notations in the example output below? I can list objects by any textual prefix, but I can't tell if the PRE (what we think of as folders) is more efficient than just the substring prefix. Full bucket list, then two text prefix, then an (empty) folder list sokoloff@ Downloads % aws s3 ls s3://foo-asdf PRE bar-folder/ PRE baz-folder/ 2022-02-17 09:25:38 0 bar-file-1.txt 2022-02-17 09:25:42 0 bar-file-2.txt 2022-02-17 09:25:57 0 baz-file-1.txt 2022-02-17 09:25:49 0 baz-file-2.txt sokoloff@ Downloads % aws s3 ls s3://foo-asdf/ba PRE bar-folder/ PRE baz-folder/ 2022-02-17 09:25:38 0 bar-file-1.txt 2022-02-17 09:25:42 0 bar-file-2.txt 2022-02-17 09:25:57 0 baz-file-1.txt 2022-02-17 09:25:49 0 baz-file-2.txt sokoloff@ Downloads % aws s3 ls s3://foo-asdf/bar PRE bar-folder/ 2022-02-17 09:25:38 0 bar-file-1.txt 2022-02-17 09:25:42 0 bar-file-2.txt sokoloff@ Downloads % aws s3 ls s3://foo-asdf/bar-folder PRE bar-folder/
- cj 5y agoOff topic: for people with a "million billion" objects, does the S3 console just completely freeze up for you? I have some large buckets that I'm unable to even interact with via the GUI. I've always wondered if my account is in some weird state or if performance is that bad for everyone. (This is a bucket with maybe 500 million objects, under a hundred terabytes)
- albert_e 5y agoI suggest you raise a support ticket. AFAIK there is server-side paging implemented in the List* API operations that the Console UI should be using so that the number of objects in a bucket should not significantly impact the webpage performance. But who knows what design flaws lurk beneath the console. Curious to know what you find. Does it happen only on opening heavy buckets? or the entire S3 console? Different Browser / incognito / different machine ...dont make a difference?
- liveoneggs 5y agothe newer s3 console works a little better. It gives pagination with "< 1 2 3 ... >"
- base698 5y agoYes, and sometimes even listing can take days. I worked somewhere that a person decided using Twitter Firehose was a good idea for S3. Keyed by tweet per file. Ended up figuring out a way to get them in batches and condense. Ended up costing about $800 per hour to fix coupled with lifecycle changes they mentioned.
- orf 5y ago> Yes, and sometimes even listing can take days. You have a versioned bucket with a lot of delete markers in it. Make sure you've got a lifecycle policy to clean them up.
- properdine 5y agoDoing an S3 object inventory can be a lifesaver here!
- valar_m 5y agoThough it doesn't address the problem in TFA, I recommend setting up billing alerts in AWS. Doesn't solve their issue, but they would have at least known about it sooner.
- solatic 5y agoTL-DR: Object stores are not databases. Don't treat them like one.
- throwaway984393 5y agoTry telling that to developers; they love using S3 as both a database and a filesystem. It's gotten to the point where we need a training for new devs to tell them what not to do in the cloud.
- Quarrelsome 5y agodo you know if such sources exist publicly? I would be most interested in perusing recommended material on the subject.
- mst 5y agoHonestly a Frequently Delivered Answers training for new developers is probably one of the best things you can include in onboarding. Every environment has its footguns, after all.
- solatic 5y agoYou can either train them with a calm tutorial or you can train them with angry billing alerts and shared-pain ex-post-facto muckraking. I, for one, prefer the calm way.
- hinkley 5y agoCommunicating through the filesystem is one of the Classic Blunders. It doesn't come up as often anymore since we generally have so many options at our fingertips, but when push comes to shove you will still discover this idea rattling around in people's skulls.
- ijlx 5y agoClassic Blunders: 1. Never get involved in a land war in Asia 2. Never go in against a Sicilian when death is on the line 3. Never communicate through the filesystem
- ebingdom 5y agoI'm confused about prefixes and sharding: > The files are stored on a physical drive somewhere and indexed someplace else by the entire string app/events/ - called the prefix. The / character is really just a rendered delimiter. You can actually specify whatever you want to be the delimiter for list/scan apis. > Anyway, under the hood, these prefixes are used to shard and partition data in S3 buckets across whatever wires and metal boxes in physical data centers. This is important because prefix design impacts performance in large scale high volume read and write applications. If the delimiter is not set at bucket creation time, but rather can be specified whenever you do a list query, how can the prefix be used to influence where objects are physically stored? Doesn't the prefix depend on what delimiter you use? How can the sharding logic know what the prefix is if it doesn't know the delimiter in advance? For example, if I have a path like `app/events/login-123123.json`, how does S3 know the prefix is `app/events/` without knowing that I'm going to use `/` as the delimiter?
- korostelevm 5y agoAWS does the optimizations over time based on access patterns for the data. Should have made that clearer in the article. The problem becomes unusual burst load - usually from infrequent analytics jobs. The indexing cant respond fast enough.
- ebingdom 5y agoThanks for the clarification. But now I'm confused about the limits: > 3,500 PUT/COPY/POST/DELETE requests per second per prefix > 5,500 GET/HEAD requests per second per prefix Most of those APIs don't even take a delimiter. So for these limits, does the prefix get inferred based on whatever delimiter you've used for previous list requests? What if you've used multiple delimiters in the past? Basically what I'm trying to determine is whether these limits actually mean something concrete (that I can use for capacity planning etc.), or whether their behavior depends on heuristics that S3 uses under the hood. I'm fine with S3 optimizing things under the hood based on access my patterns, but not if it means I can't reason about these limits as an outsider.
- 5y ago
- liveoneggs 5y agoI have caused billing spikes like this before those little warnings were invented and it was always a dark day. They are really a life saver. Lifecycle rules are also welcome. Writing them yourself was always a pain and tended to be expensive with list operations eating up that api calls bill. ---- Once I supported an app that dumped small objects into s3 and begged the dev team to store the small objects in oracle as BLOBS to be concatenated into normal-sized s3 bjects after a reasonable timeout where no new small objects would reasonably be created. They refused (of course) and the bills for managing a bucket with millions and millions of tiny objects were just what you expect. I then went for a compromise solution asking if we could stitch the small objects together after a period of time so they would be eligible for things like infrequent access or glacier but, alas, "dev time is expensive you know" so N figure s3 bills continue as far as I know.
- sharken 5y agoI suppose it's not just dev time on the line, but also the risk of doing the change that is thought to be too high. If I ever get to be a manager I'd go for an idea such as yours. Though I suspect too many managers are too far removed from the technical aspect of things and don't listen nearly enough.
- darkwater 5y ago> I then went for a compromise solution asking if we could stitch the small objects together after a period of time so they would be eligible for things like infrequent access or glacier but, alas, "dev time is expensive you know" so N figure s3 bills continue as far as I know. This hits home so hard that it hurts. In my case is not S3 but compute bills but the core concept is the same.
- WrtCdEvrydy 5y agoBecause the bill isn't a "dev problem". Once you move those bills to "devops", it becomes an infrastructure problem.
- zrail 5y ago
- vdm 5y agoDeleteObjects takes 1000 keys per call. Lifecycle rules can filter by min/max object size. (since Nov 2021)
- vdm 5y agoAthena supports regexp_like(). By loading in an S3 inventory this can match what a wildcard would. Then a Batch Operations job can tag the result. Not easy, but is possible and effective.
- electroly 5y agoThank you for mentioning that lifecycle rule change. I must have missed the announcement; that is exactly the functionality I needed.
- jwalton 5y agoYour website renders as a big empty blue page in Firefox unless I disable tracking protection (and in my case, since I have noscript, I have to enable javascript for "website-files.com", a domain that sounds totally legit).
- mst 5y agoI have tracking protection and ublock origin both enabled and it rendered fine (FF on Win10). (presented as a data point for any poor soul trying to replicate your problem)
- tazjin 5y agoChrome with uBlock Origin on default here, and it renders a big blue empty page for me, too. That's despite dragging in an ungodly amount of assets first. Here's an archive link that works without any tracking, ads, Javascript etc.: https://archive.is/F5KZd https://archive.is/F5KZd
- Sophira 5y agoThe problem is that the DIV that contains the main text has the attribute 'style="opacity:0"'. Presumably, this is something that the JavaScript turns off. A lot of sites like to do things like this for some reason. I haven't figured out why. I like to use Stylus to mitigate these if I can, rather than enabling JavaScript.
- charcircuit 5y agoCan someone explain what happened in the end? From my understanding nothing happened (they deprioritizod the story for fixing it) and they are still blowing through the cloud budget.
- seekayel 5y agoHow I read the article, nothing happened. I think it is a cautionary tale of why you should probably bite the bullet and press the button instead of doing the "easier" thing which ends up being harder and more expensive in the end.
- snowwrestler 5y agoThey didn’t resolve the issue. There’s an important moment in the story, where they realize the fix will incur a one-time fee of $100,000. No one in engineering can sign off on that amount, and no one wants to try to explain it to non-technical execs. They don’t explain why. But it’s probably because they expect a negative response like “how could you let this happen?!” or “I’m not going to pay that, find another way to fix it.” In a lot of organizations it’s easier to live with a steadily growing recurring cost than a one-time fee… even if the total of the steady growth ends up much larger than the one-time fee! It’s not necessarily pathological. Future costs will be paid from future revenue; whereas a big fee has to be paid from cash on-hand now. But sometimes the calculation is not even attempted because of internal culture. When the decision is “keep your head down” instead of “what’s the best financial strategy,” that could hint at even bigger potential issues down the road.
- hogrider 5y agoSounds more like non technical leadership sleeping at the wheel. I mean if they could just afford to lose money like this why bother with all that work to fix it?
- pattycake23 5y agoHere's an article about Shopify running into the S3 prefix rate limit too many times, and tackling it: https://shopify.engineering/future-proofing-our-cloud-storage-usage https://shopify.engineering/future-proofing-our-cloud-storag...
- sciurus 5y agoTheir solution was to introduce entropy into the beginning of the object names, which used to be AWS's recommendation for how to ensure objects are placed in different partitions. AWS claims this is no longer necessary, although how their new design actually handles partitioning is opaque. "This S3 request rate performance increase removes any previous guidance to randomize object prefixes to achieve faster performance. That means you can now use logical or sequential naming patterns in S3 object naming without any performance implications." https://aws.amazon.com/about-aws/whats-new/2018/07/amazon-s3-announces-increased-request-rate-performance/ https://aws.amazon.com/about-aws/whats-new/2018/07/amazon-s3...
- pattycake23 5y agoSeems like it's a much higher rate limit, but it exists none the less, and Shopify's scale has also grown significantly since 2018 (when that article was written) - so it was probably a valid way for them to go.
- sciurus 5y agoI think two things happened that are covered in that blog post 1) The performance per partition increased 2) The way AWS created partitions changed When I was at Mozilla, one thing I worked on was Firefox's crash reporting system. It's S3 storage backend wrote raw crash data with the key in the format `{prefix}/v2/{name_of_thing}/{entropy}/{date}/{id}`. If I remember correctly, we considered this a limitation since the entropy was so far down in the key. However, when we talked to AWS Support they told us their was no longer a need to have the entropy early on; effectively S3 would "figure it out" and partition as needed. EDIT: https://news.ycombinator.com/item?id=30373375 https://news.ycombinator.com/item?id=30373375 is a good related comment.
- lenkite 5y agosigh. My team is facing all these issues. Drowning in data. Crazy S3 bill spikes. And not just S3 - Azure, GCP, Alibaba, etc since we are a multi-cloud product. Earlier, we couldn't even figure out lifecycle policies to expire objects since naturally every PM had a different opinion on the data lifecycle. So it was old-fashioned cleanup jobs that were scheduled and triggered when a byzantine set of conditions were met. Sometimes they were never met - cue bill spike. Thankfully, all the new data privacy & protection regulations are a life-saver. Now, we can blindly delete all associated data when a customer off-boards or trial expires or when data is no longer used for original purpose. Just tell the intransigent PM's that we are strictly following govt regulations.
- CydeWeys 5y agoThe data protection regulations really are so freeing, huh. It's amazing to be able to delete all this stuff without worrying about having to keep it forever.
- whimsicalism 5y agonow this is a spin i havent heard before.
- hvs 5y agoYou haven't heard it because it's not spin, it's from an engineer's point of view. That's not the view you hear in the news when it comes to these things.
- alisonkisk 5y agoEh, Retention and Deletion are both pain for devs. Not having to care is the happy state.
- whimsicalism 5y agoHN seems like an odd place to assume that people only hear about things from the news and aren't engineers themselves. i am a dev that has to deal with these regulations in my day to day. it is a pain, it is not freeing in any sense, and it makes my models worse. granted, i think there are good reasons for it, but it does not make my life easier for sure.
- StratusBen 5y agoOn this topic, it's always surprising to me how few people even seem to know about different storage classes on S3...or even intelligent tiering (which I know carries a cost to it, but allows AWS to manage some of this on your behalf which can be helpful for certain use-cases and teams). We did an analysis of S3 storage levels by profiling 25,000 random S3 buckets a while back for a comparison of Amazon S3 and R2* and nearly 70% of storage in S3 was StandardStorage which just seems crazy high to me. * https://www.vantage.sh/blog/the-opportunity-for-cloudflare-r2 https://www.vantage.sh/blog/the-opportunity-for-cloudflare-r...
- blurker 5y agoI think that it's not just people not knowing about the lifecycle feature, but also that when they start putting data into a bucket they don't know what the lifecycle should be yet. Honestly I think overdoing lifecycle policies is a potentially bigger foot gun than not setting them. If you misuse glacier storage that will really cost you big $$$ quickly! And who wants to be the dev who deleted a bunch of data they shouldn't have? Lifecycle policies are simple in concept, but it's actually not simple to decide what they should be in many cases.
- asim 5y agoThe AWS horror stories never cease to amaze me. It's like we're banging our heads against the wall expecting a different outcome each time. What's more frustrating, the AWS zealots are quite happy to tell you how you're doing it wrong. It's the users fault for misusing the service. The reality is, AWS was built for a specific purpose and demographic of user. It's now complexity and scale makes it unusable for newer devs. I'd argue, we need a completely new experience for the next generation.
- jollybean 5y agoIn this case it is absolutely the user 'doing it wrong'. AWS allows you to store gigantic amounts of data, thus lowering the bar dramatically for the kinds of things that we will keep. This invariably creates a different kind of problem when those thresholds are met. In this case, you have 'so much data you don't know what to do with it'. Akin to having 'really cheap warehouse storage space' that just gets filled up. "It's now complexity and scale makes it unusable for newer devs. I'" No - the 'complexity' bit is a bit of a problem, but not the scale. The 'complexity bit' can be overcome if you stick to some very basic things like running Ec2 instances and very basic security configs. Beyond that, yes it's hard. But the 'equivalent' of having your own infra would be simply to have a bunch of Ec2 instances on AWS and 'that's it' - and that's essentially achievable without much fuss. That's always an option to small companies, i.e. 'just fun some instances' and don't touch anything else.
- rmbyrro 5y agoWhat do you see missing or not well explained in AWS documentation that newer devs wouldn't understand? I started using S3 early in my career and didn't see this problem. I always thought in data retention during design phase. My opinion is that lazy, careless or under time pressure developers will not, and then will get bitten. But it would happen to any tool. Maybe a different problem, but they'll always get bitten hard ...
- chrisjc 5y agoForget newer devs for a moment... I've had years of experience with S3 and sounds like the author of the article has too. Despite my years of experience in programming/DBs/etc, I'm definitely not an amazing developer. But I learned a whole lot of new things from this article that I didn't understand from reading the AWS documentation, let alone think I had to even concern myself with some of these issues. Spotty warnings about transitional request charges? Anyway, kudos to you for always thinking about (and i hope actualizing) retention during policies the design phase. However, while I certainly think devs bare some of this responsibility, I'm sure they're usual met with all of the usual excuses and kicking the can down the road line of reasoning from PM/PO/etc that lead to these kinds of nightmares in the beginning... Then again, it will probably be another developer or system admins' nightmare when it becomes an issue. Even as an experience engineer, I still struggle setting the retention policy at the beginning of a new design... I'd love to hear any advice you have about how manage this incredibly important aspect?
- gfd 5y agoDoes anyone have recommendations on how to compress the data (gzip or parquet).
- pontifier 5y agoDON'T PRESS THAT BUTTON. The egress and early deletion fees on those "cheaper options" killed a company that I had to step in and save.
- pphysch 5y agoOn a related note, suppose the Fed raises rates to mitigate inflation and indirectly kills thousands of zombie companies, including many SaaS renting the cloud. What happens to their data? Does the cloud unilaterally evict/delete it, or does it get handled like an asset -- auctioned off, etc?
- Uehreka 5y ago> does it get handled like an asset -- auctioned off, etc? Who would buy that? I guess if this happened enough then people would start "data salvager" companies that specialize in going through data they have no schema for looking for a way to sell something of it to someone else. I have to imagine the margins in a business like that would be abysmal, and all the while you'd be in a pretty dark place ethically going through data that users never wanted you to have in the first place. Of course, all these questions are moot because if this happened the GDPR would nuke the cloud provider from orbit.
- cmckn 5y agoI’m not aware of a cloud provider that is contractually allowed to do such a thing (except maybe alibaba by way of the CCP). Dying companies get purchased and have their assets pilfered every day, the same thing would happen with cloud assets.
- bpicolo 5y agoIf the dead company stops paying the bills, Amazon can definitely delete that.
- cmckn 5y agoOf course, I meant the idea of auctioning it off or otherwise accessing the customer’s data. When I worked in Azure, I accidentally created some internal resources in a personal account. I didn’t have the ability to delete them after I left; the only way to do so was to cancel the credit card and let the grace period expire.
- gtirloni 5y agoA "TLDR" that is not.
- zitterbewegung 5y agoI was at a presentation where HERE technologies told us that they went from being on the top ten (or top five) S3 users (by data stored) to getting off of that list. This was seen as a big deal obviously.
- dekhn 5y agoI had to chuckle at this article because it reminded me of some of the things I've had to do to clean up data. One time I had to write a special mapreduce that did a multiple-step-map to converted my (deeply nested) directory tree into roughly equally sized partitions (a serial directory listing would have taken too long, and the tree was really unbalanced to partition in one step), then did a second mapreduce to map-delete all the files and reduce the errors down to a report file for later cleanup. This meant we could delete a few hundred terabytes across millions of files in 24 hours, which was a victory.
- lloesche 5y agoI had a similar issue at my last job. Whenever a user created a PR on our open source project artifacts of 1GB size consisting of hundreds of small files would be created and uploaded to a bucket. There was just no process that would ever delete anything. This went on for 7 years and resulted in a multi-petabyte bucket. I wrote some tooling to help me with the cleanup. It's available on Github: https://github.com/someengineering/resoto/tree/main/plugins/aws/resoto_plugin_aws/cmd/ https://github.com/someengineering/resoto/tree/main/plugins/... consisting of two scripts, s3.py and delete.py. It's not exactly meant for end-users, but if you know your way around Python/S3 it might help. I build it for a one-off purge of old data. s3.py takes a `--aws-s3-collect` arg to create the index. It lists one or more buckets and can store the result in a sqlite file. In my case the directory listing of the bucket took almost a week to complete and resulted in a 80GB sqlite. I also added a very simple CLI interface (calling it virtual filesystem would be a stretch) that allows to load the sqlite file and browse the bucket content, summarise "directory" sizes, order by last modification date, etc. It's what starts when calling s3.py without the collect arg. Then there is delete.py which I used to delete objects from the bucket, including all versions (our horrible bucket was versioned which made it extra painful). On a versioned bucket it has to run twice, once to delete the file and once to delete the then created version, if I remember correctly - it's been a year since I built this. Maybe it's useful for someone.
- k__ 5y agoWhat about the lifecycle stuff? I thought, S3 can move stuff to cheaper storage automatically after some time.
- lloesche 5y agoLike I wrote for us it was a one-off job to find and remove 6+ year old build artifacts that would never be needed again. I just looked for the cheapest solution of getting rid of them. I couldn't do it by prefix alone (prod files mixed in the same structure as the build artifacts) which is why delete.py supports patterns (the `--aws-s3-pattern` arg takes a regex). If AWS' own tools work for you it's surely the better solution than my scripts. Esp. if you need something on an ongoing bases.
- wackget 5y agoAs a web developer who has never used anything except locally-hosted databases, can someone explain what kind of system actually produces billions or trillions of files which each need to be individually stored in a low-latency environment? And couldn't that data be stored in an actual database?
- abhishekjha 5y agoAn image service.
- wackget 5y agoYeah that use-case I get. Binary files which would be difficult/impractical to index in a database. However it feels like something at that scale will only ever realistically be dealt with by enterprise-level software, and I'd hazard a guess that most developers - even those reading HN - are not working on enterprise-level systems. So I'm wondering what "regular devs" are using cloud buckets for at such a scale over regular DBs.
- rgallagher27 5y agoThings like mobile/webisite analytics events. User A clicked this menu item, User B viewed this images etc All streamed into S3 in chunks of smallish files. It's cheaper to store them in S3 over a DB and use tools like Athena or Redshift spectrum to query.
- wackget 5y agoWow. What makes it cheaper than using a DB? Is it just because the DB will create some additional metadata about each stored row or something?
- bpicolo 5y agoS3 is often essentially a database in these scenarios. You store columnar data format files in S3, and various analytical systems can query with S3 as a massive backing storage.
- jopsen 5y agoOne of the biggest pains is that cloud services rarely mention what they don't do. I think it's really sad, because when I don't see docs clearly stating the limits, I assume the worst and avoid the service.
- Mave83 5y agoJust avoid the cloud. You get a Ceph storage with the performance of Amazon S3 at the price point of Amazon S3 Glacier in any Datacenter worldwide deployed if you want. There are companies that help you doing this. Feel free to ask if you need help.
- red0point 5y agoI want to know what the absolute cheapest way of doing this is, without having a lot of CapEx. I thought of renting dedicated storage servers (e.g. Hetzner) and slapping Ceph on them. Do you have another, better, idea?
- charcircuit 5y agoYou have to properly administrate those servers else you'll lose all your files and everything will be inaccessible.
- klysm 5y agoAdministrating CEPH is unfortunately hard.
- rizkeyz 5y agoI did the back-of-the-envelope math once. You get a Petabyte of storage today for $60K/year if you buy the hardware (retail disks, server, energy). It actually fits into the corner of a room. What do you get for $60K in AWS S3? Maybe a PB for 3 months (w/o egress). If you replace all your hardware every year, the cloud is 4x more expensive. If you manage to use your getto-cloud for 5 year, you are 20x cheaper than Amazon. To store one TB per person on this planet in 2022, it would take a mere $500M to do that. That's short change for a slightly bigger company these days. I guess by 2030 we should be able to record everything a human says, sees, hears and speaks in an entire life for every human on this planet. And by 2040 we should be able to have machines learning all about human life, expression and intelligence to slowly making sense of all of this.
- arein3 5y ago>I guess by 2030 we should be able to record everything a human says, sees, hears and speaks in an entire life for every human on this planet. That's a very good point. Are you employed? Would you like to join Meta?
- gmiller123456 5y agoI don't get what's going on with on-line storage. You can walk in Best Buy and get a few Tb hard drive for well under $100. Yet every cloud service wants to charge you several times that per year for just 1Tb. I understand drives fail, there's operating cost, and some need extremely low latency. But there seems to be a huge disparity between what a hard drive costs, and what it costs to make it available on the Internet.
- tekknik 5y agoThere’s a difference between a consumer drive and a server drive. Plop that $100 drive in and you may be back in a week or so replacing it.
- gmiller123456 5y agoWhy would you think a drive automatically looses lifespan just because the computer it's in is referred to as a server? Many of my desktop hard drives see more activity that some of my website HDs.
- zmmmmm 5y agoThe rationale for using cloud is so often that it saves you from complexity. It really undermines the whole proposition when you find out that the complexity it shields you from is only skin deep, and in fact you still need a "PhD in AWS" anyway. But as a bonus, now you face huge risks and liabilities from single button pushes and none of those skills you learned are transferrable outside of AWS so you'll have to learn them again for gcloud, again for azure, again for Oracle ....
- hughrr 5y agoFor every $100k bill there’s a hundred of us with 14TB that costs SFA to roll with.
- harshaw 5y agoAWS budgets is a tool for cost containment (among other external services).
- 0x002A 5y agoEach time a developer does something on a cloud platform, that moment the platform might start to profit for two reasons: vendor lock-in and accrued costs in the long term regardless of the unit cost. Anything limitless/easiest has a higher hidden cost attached.
- deleted 5y ago[deleted]
- kondro 5y agoThe minimum size of objects in cheaper storage types is 128KiB. Given the article quotes $100k to run an inventory (and $100k/month in standard storage) it's likely most of your objects are smaller than 128KiB and so probably wouldn't benefit from cheaper storage options (although it's possible this is right on the cusp of the 128KiB limit and could go either way). Honestly, if you have a $1.2m/year storage bill in S3 this would be the time to contact your account manager and try to work out what could be done to improve this. You probably shouldn't be paying list anyway if just the S3 component of your bill is $1.2m/year.
- deleted 5y ago[deleted]
- cyanic 5y agoWe solved the problem of deleting old files early in our development process, as we wanted to avoid situations such as this one. While developing GitFront, we were using S3 to store individual files from git repositories as single objects. Each of our users was able to have multiple repositories with thousands of files, and they needed to be able to delete them. To solve the issue, we implemented a system for storing multiple files inside a single object and a proxy which allows accessing individual files transparently. Deleting a whole repository is now just a single request to S3.
- gnutrino 5y agoLol this post hits close to home.