7 ms·
Having been on both sides (admin and developer), developers are notoriously bad at estimating how much space they need. You can't give them carte blanche to the
by stuff4ben 3y ago
Having been on both sides (admin and developer), developers are notoriously bad at estimating how much space they need. You can't give them carte blanche to the storage because they'll waste it and consume as much as they're given without a thought to conserving it. And then when you put limits in, they'll whine and complain until they get what they want. Being an Artifactory service provider for a large IT dept gave me a direct view in how hard it is to manage storage for developers. And as a developer using Artifactory, I don't want to worry about storage, I just want my builds and CI pipelines to complete.
- geerlingguy 3y agoThis reminds me of a time we were helping a dev team bring logging in house because they weren't liking the features of their logs-as-a-service provider. They set all applications to "debug" level logs in production and were generating multiple gigabytes of logs per hour. They wanted 90 days retention, and the ability to do advanced searching through the live log data so they could debug in production (they didn't really use their dev or stage environments, or have a process for documenting and reproducing bugs).
- b112 3y agoWith silliness like that, you can bet it was the cost of their logging provider, as the feature they didn't like.
- Veserv 3y ago90 days retention is only 2,160 hours. Even at 999 GB/hr that is only ~2160 TB of storage. So, if we stretch the definition of “multiple gigabytes”, is maybe $100k in storage which is around 3-6 developer-months. If we use a more reasonable definition like 10 GB/hr, then that is 20 TB, so maybe $1k in storage which is around 1 developer-day. Seems pretty reasonable to me.
- floorballchamp 3y agoPlus, logs have enormous compression potential since their entropy is so low. That's the property exploited by every logging-as-a-service out there.
- giovannibonetti 3y agoRelated to that, last year Uber's engineering blog mentioned very interesting results with their internal log service [1]. I wonder if there's anything as good in the open-source world. The closest thing I can think of is Clickhouse's "new" JSON type, which is backed by columnar storage with dynamic columns [2]. [1] https://www.uber.com/en-BR/blog/reducing-logging-cost-by-two-orders-of-magnitude-using-clp/ https://www.uber.com/en-BR/blog/reducing-logging-cost-by-two... [2] https://clickhouse.com/docs/en/integrations/data-formats/json https://clickhouse.com/docs/en/integrations/data-formats/jso...
- Veserv 3y agohttps://messagetemplates.org/ https://messagetemplates.org/ The design described there is what Uber should be logging in the first place. Instead they are logging the fully resolved message and then compressing back into the templated form. However, the compression back into the templated form is a good idea if you have third party logs that you want to store where you can not rewrite the logging to generate the correct form in the first place.
- giovannibonetti 3y agoNeat! The only downside of this approach is having to force the developers to use the library, which can work in some companies. On the other hand, other approaches discussed previously like Uber's don't require any change in the application code, which should make adoption way simpler.
- syndicatedjelly 3y ago…what? Without any other context on what they’re working on or the size of the company, an extra developer of cost is automatically reasonable?
- yjftsjthsd-h 3y ago... only 2PB? You might be using a different scale than some of us.
- Dylan16807 3y agoTheir scale was money. Saying something is "only" a single digit number of developer months makes sense in this context. And that was a number hundreds of times higher than what they were replying to, just to make a point.
- Volundr 3y agoA few years ago I joined a company aggressively trying to reduce their AWS costs. My jaw got the floor when I realized they were spending over a million a month in AWS fees. I couldn't understand how they got there with what they were actually doing. I feel like this comment perfectly demonstrates how that happens.
- namaria 3y agoCloud truly monetizes the tar pit.
- ilyt 3y agoDeveloper will do the simplest thing to solve the problem. If the solutions are: * rewrite that part to add retention, or use better compression, or spend next month deciding which data to keep and which can be removed early * Wiggle a thing in panel/API giving it more space The second will win every single time unless there is pushback or it hits the 5% of the developers that actually care to make good architecture not just deliver tickets.
- SOLAR_FIELDS 3y agoAWS also purposefully makes it easy to shoot yourself in the foot. Case in point that we were burned on recently: - set up some service that talks to a s3 bucket - set up that bucket in the same region/datacenter - send a decent but not insane amount of traffic through there (several hundred Gb per day) - assume that you won’t get billed any data transfer fees since you’re talking to a bucket in the same data center - receive massive bill under “EC2-Other” line item for NAT data transfer fees - realize that AWS routes all traffic through NAT gateway by default even though it’s just turning around and going back into the data center it came from and billing exorbitant fees for that - come to the conclusion that this is obviously a racket designed to extract money from unsuspecting people because there is almost no situation where you would want to do that by default and discover that hundreds to thousands of other people have been screwed in the exact same way for years (and there is a documented trail of it[1]) 1: https://www.lastweekinaws.com/blog/the-aws-managed-nat-gateway-is-unpleasant-and-not-recommended/ https://www.lastweekinaws.com/blog/the-aws-managed-nat-gatew...
- breakwaterlabs 3y agoIn what world is 2160TB $100k? Current single disk solutions are around $25/TB for HDDs and ~$100/TB for NVMe. At a minimum you're looking at $54k just for raw capacity-- assuming no backup, no chassis, no networking, and no redundancy. More reasonable estimations would be in excess of $400/TB.
- Veserv 3y agoSure, whatever, a factor of 10 here or there hardly matters. I literally misinterpreted “multiple gigabytes per hour” as 999 GB/hr, not a much more reasonable 10 GB/hr. I literally overestimated data rates by a factor of 10,000% and the number still comes out “reasonable” i.e. a cost that can be paid if the cost/benefit is there. Unless you want to claim storage costs $5,000/TB for 3 MB/s of I/O “multiple gigabytes per hour” with 90 day retention for a team worth of logging is not stupid on its face. Not to say that is a efficient or smart solution, but certainly not a “look at this insane request by developers” the person I was originally responding to was making it out to be. Personally, I would probably question the competence of the team if they had that sort of logging rate with manual logging statements, but I am merely pointing out that “multiple gigabytes per hour” for 90 days is not crazy on its face and a plausible business case could be made for it even with a relatively modest engineering team.
- breakwaterlabs 3y agoMy recent discussions with multiple SAN vendors as well as quoting out cost to DIY storage has that number being far away from "reasonable". I do not claim storage is $5,000/TB but it is substantially higher than the $50/TB you're estimating. It's difficult to estimate the log throughput in this scenario. Cisco on debug all can overload the device's CPU; systems like sssd can generate MB of logs for a single login. All of this is really missing the core issue though. A 2PB system is nontrivial to procure, nontrivial to run, and if you want it to be of any use at all you're going to end up purchasing or implementing some kind of log aggregation system like Splunk. That incurs lifecycle costs like training and implementation, and then you get asked about retention and GDPR.... and in the process, lose sight of whether this thing you've made actually provides any business value. IT is not an ends in itself, and if these logs are unlikely to be used the question is less about dollars-per-developer-hour and more about preventing IT scope creep and the accumulation of cruft that can mature into technical debt.
- justinclift 3y agoFor their use case it sounds like they wanted to index the heck out of it for near instant lookups and similar too. So probably need to double the data size (rough guess) to include the indexes. And it may need some dedicated server nodes just for processing/ingestion/indexing, etc.
- arghwhat 3y agoThe idea that anyone would find storing 20TB of plain text logs for a normal service reasonable is quite amusing. Don't get me wrong, I understand that a single-digit kUSD/month is peanuts against developer productivity gains, but I still wouldn't be able to take a developer making that suggestion seriously. I would also seriously question internal processes, GDPR (or equivalent) compliance, and whether the system actually brings benefit or if it is just lazy "but what if" thinking.
- tcbawo 3y agoThis doesn’t seem terrible if the business benefits justify the costs. There is a cost/benefit to this, presumably.
- deleted 3y ago[deleted]
- ballenf 3y agoAllocated storage should come directly from the consuming team's budget. Divvy up the total storage cost and allocate in proportion to requested limits.
- atoav 3y agoSure, what are you going to bill me for a 30GB VM on a 100TB cluster? Whether I want 30GB or 100GB for an absolute central service for the whole org shouldn't matter. If we are talking about personal pet projects or user accounts — sure — but that wasn't my complaint here.
- akira2501 3y agoYou have to fill out a load chart to fly a plane, they should have to fill out something like a storage chart to get a production allocation. What size are your objects? How many per unit of time and served entity? What is the lifetime of those objects? How is that lifetime managed?
- ido 3y agoIf you agree to add a few months of development time and reduce future velocity to make sure these limits are enforced, sure. Usually adding storage costs about as much as 1 developer’s salary cost for what, an hour? A day?
- ilyt 3y agoIt's not for saving storage. It's for making sure it will actually not overflow. End number doesn't matter, what matters is developers thinking about how long data should be stored and what data should be stored. Not doing that analysis and overprovisioning 4x will just cause disaster in 2 years instead of 6 months.
- stuff4ben 3y agoYou missed the part where I said they are "notoriously bad at estimating". We really do suck at estimating everything... storage, work estimates, etc. Why can't we just say "it'll be done when it's done and I'll use ALL the storage until I'm done"?
- atoav 3y agoI mean in my case it was literally a database file filled with the number of (dummy) people who are currently in our org. So that database size was the size of the project. He just didn't plan for the size of backups (backups were his job, not ours).
- solatic 3y ago> You can't give them carte blanche to the storage because they'll waste it So what? Just buy more. Storage is cheap. It's hard to have a discussion here without understanding the scales involved. Is the problem that they're wasting 100 GB or 100 TB? And if the issue is truly that they're wasting 100 TB, then clamp down on it as part of cost reduction efforts. The truth in most organizations is you get rewarded for eliminating mountains of waste, but trying to prevent the waste in the first place brands you as someone difficult to work with who is standing in the way. Why not lean into that?