6 ms·
In my day job, I have to write reams of complicated code that can slow down the system or make maintenance more annoying...just because a user can do something
by nathanb 11y ago
In my day job, I have to write reams of complicated code that can slow down the system or make maintenance more annoying...just because a user can do something (even though doing that is unsupported).
That's what it means to write business-class software. Nobody worth having as a customer is going to build their business on your platform if your attitude whenever something goes wrong is "you shouldn't have been doing that in the first place".
(I am surprised to hear this story though. I actually found a bug in the NTFS buffer cache a few years ago which was introduced in (as I recall) Windows Server 2012. Maybe the Server organization are way more on the ball than the consumer OS organization, which is definitely possible. But they took it seriously and fixed it in a patch.)
- angersock 11y agoThere's a balance, though. Oftentimes, you serve the business folks better by preventing them from doing something really silly, and instead letting them optimize their workflow not to include silliness. Unfortunately, it's tricky to tell ahead of time which features are bad ideas and which aren't.
- tjradcliffe 11y agoIt's pretty easy at first order: features that users want are good ideas. I agree at second order users want a lot of features that solve their problems in ways that are less clever than one might like, but in the vast majority of cases the right answer from development is, "Users know what they need better than I do, so if they ask for something, I'll do my best to implement it even if I don't totally understand why." That answer exposes the deeper answer, too: "My job as a developer is to understand user needs so that our software can help them fulfill them, so if I don't understand why someone needs a feature I should dig in further before implementing it so I don't implement the wrong thing or the right thing in the wrong way." Not always possible given shedule constraints, though. Admitedly, those rules don't seem to apply in this case, since if the OS allows you to do something that corrupts data, that's a problem no matter why the user wants to do it. If it can corrupt data, the OS shouldn't allow you to do it, end of story.
- nathanb 11y agoIt's not that simple (speaking as someone who works on OS SCSI driver code for a living...). I will amend your statement to say that the OS itself should never corrupt your data, which I agree with entirely. The post was not super clear about exactly what was happening (or it may be that my limited knowledge of Windows storage internals is keeping me from understanding it), but it sounds like the NTFS client requested a cache flush and then was issuing writes during the flush. I don't know what contract these operations have, but it may very well be the case that the user was violating the contract. If Microsoft responded with "don't do that", this may be the case. But wait! Shouldn't Windows prevent the data from being corrupted? Or shouldn't NTFS fail the writes in this case? Possibly. And most likely inserting the checks to make this happen would increase write latency for every NTFS client, even the ones that don't behave in this way. This reminds me of another scenario I encountered, with Veritas VxFS running on top of AIX. The user initiated a space reclamation, which was sending what you can think of as a delete to the storage array. And the user was also writing data to the device at the same time. Due to a race condition (which I can describe for you if you really care), the legitimate user data would sometimes be deleted. Should VxFS have protected users against this case? Yes. Was VxFS violating the SCSI protocol? No. Was the storage array violating the SCSI protocol? No. (Has VxFS fixed this bug, almost three years after I discovered it? No comment.) It's always a lot more complicated than it seems on the outside.
- chetanahuja 11y agoSorry I don't buy that. An operating system's "contract" with the user is the syscall API. There's no room for argument there. Calling a write while flushing the same data from another process (as the OP reported) or thread is perfectly legitimate set of operations. If these operations are not supposed to be running simultaneously, that has to be enforced by the OS kernel. The "right" way would be for the OS to serialize these operations internally. But even returning an error for the write might be an acceptable (though not nice) way to handle it. What's absolutely not acceptable is randomly corrupting data.
- deleted 11y ago[deleted]
- derefr 11y ago> Nobody worth having as a customer is going to build their business on your platform if your attitude whenever something goes wrong is "you shouldn't have been doing that in the first place". My favorite set of APIs is AWS. You know why? Because they've realized they hold two very weighty sticks that they can use when designing, and they've put them in place all over. 1. They can make any arbitrary message to the API cost the user money every time they send it, to disincentivize using that part of the API thoughtlessly. That's whether or not they expect this to be an actual revenue stream at the rates people are charged for reasonable usage. 2. They can put a "soft cap" on any arbitrary resource, so that you have to phone them and get the cap raised if you want more than [some reasonable number] of something. This likewise disincentivizes bad designs that use a nigh-infinite number of costly somethings to accomplish tasks that could be just as easily accomplished some more idiomatic, less costly way. AWS doesn't prevent you from doing stupid things... but it makes you really not want to. I love it.
- click170 11y agoHave you ever tried to setup AWS IAM permissions for a user pursuant to the principle of least privilege? Because Amazon's APIs are about as far from friendly as you can get in this respect. Their docs make it easy to make the mistake of thinking that fine-grained controls are available for most things, but when it comes to really important things like being able to segregate a production and Dev VPC, their APIs basically force you to grant permissions to everything or nothing. Some examples of things I've hit: Not being able to restrict a user to only change a specific routing table Not being able to restrict a user to only change a specific elastic NIC I'm consistently surprised at what's missing from their API and couldn't disagree more about being happy with it.
- vacri 11y agoThe 'fine-grain' of IAM varies considerably depending on which AWS service you're restricting. You can add extra flexibility with 'Conditions', which I'm sure you're aware of, but I think it's a bit of a misrepresentation to paint IAM as being poor quality. AWS is a very complex environment; I can't see how you could have a user-friendly yet fine-grained user control for something that complex. Anything you choose is going to require training in how to use it. I wouldn't say I'm happy about it, but neither am I unhappy, and neither am I happy about anything in the world of security (also in today's task list is updating https cipher lists... again...). Not even the simplest thing in security is easy. For example, the basic concept of a password is simple, but actually implementing it? Ugh - it involves every layer from backend to frontend to user training (the hardest part - no sticky notes, no friendly phone calls, no passing around in emails...). Anyway, for those not used to IAM 'Conditions', an example of use. The following allows Packer (an AMI builder) to destroy any EC2 instance, but only if they have the tag 'name' as 'Packer Builder'. Conditions don't work for everything, so they're not a workaround to get fine-grain everywhere, but they do add a lot of flexibility. "Sid": "AllowInstanceActions", "Effect": "Allow", "Action": [ "ec2:StopInstances", "ec2:TerminateInstances", "ec2:AttachVolume", "ec2:DetachVolume", "ec2:DeleteVolume" ], "Resource": [ "arn:aws:ec2:us-east-1:xxxxxx:instance/*", "arn:aws:ec2:us-east-1:xxxxxx:volume/*", "arn:aws:ec2:us-east-1:xxxxxx:security-group/*" ], "Condition": { "StringEquals": { "ec2:ResourceTag/Name": "Packer Builder" } }