21 ms·
Can anybody make a strong case to me as to why are buffer overflows considered an issue in C when it takes like 10 minutes to write and test an array implementa
by dreta 10y ago
Can anybody make a strong case to me as to why are buffer overflows considered an issue in C when it takes like 10 minutes to write and test an array implementation that prevents that from ever happening? I do agree that C has issues (though in my opinion neighter Rust nor Go address almost any of them) i just don't understand why are buffer overflows such a huge problem in C when the same thing is going to come up when trying to work with memory in Rust.
- ArkyBeagle 10y agoC is the new "goto". Y'all please, please note that dreta said "... an array implementation that prevents that from ever happening..."
- pjmlp 10y agoC is the PHP and JavaScript of the 70's.
- ArkyBeagle 10y ago:) I don't know anything about either.
- ArkyBeagle 10y agoFinal word on the subject - what we have is people who are trying to make Open Software Reputation Points by finding a problem and fixing it, rather than waiting to find a real problem and fixing that. When the figure of "ten minutes" was used - that's really what it should be as a mean or median figure, with some long-tail outliers for knotty cases. While I am (somewhat) sympathetic, I don't miss what it is - it's make-work. There's not some body of material on language, O/S and software design available today that wasn't available 40 years ago. People new better back then, and didn't do better. But the endless, circular, Utopian arguments just don't cut it. It ain't the crate, it's the pilot. I am sorry that oh so many are put upon to actually - gasp - test their code but that's what this is. The economics of a Rust or Go or whatever make perfect sense - in the long run. Just recall what Keynes said about the long run.
- thesmallestcat 10y agoThis is how I feel. You cut the code to 27%. Great. You C99'd the code. Grand. Now you're going to rewrite the whole danged thing in a new language. Wonderful. This is the sort of stuff I cared about so much as a junior dev. But a user asks, are we there yet?
- mden 10y agoBut adding functionality grows in effort as more code is added. At some point it makes sense to improve a code base so you improve the rate at which you are getting "there". In fact there are a lot of code bases, perhaps most, which simply cannot get there in any sane way and fall short of achieving their goal. You might not care how goals are achieved as a user but as an engineer it's worth exploring and discussing different approaches.
- deleted 10y ago[deleted]
- thesmallestcat 10y agoI had no idea!
- ArkyBeagle 10y agoI care more as a user than as an engineer. :) Just to be clear - I don't suspect that the things that pretend to replace C are bad, or no good. I've just seen multiple "pretenders to the throne"[1] and what would appear to have happened is that we just moved the pathology around. [1] please excuse the horrible metaphor. After a few iterations of that, one begins to think that perhaps this is a distinctly human problem that is not necessarily addressable by improved language systems. My greatest concern is that I keep seeing people learn the same things, over and over, on the job. There's no real repository of literature to actually address any of this - each engineer appears to have to learn it mostly from scratch. There's a famous bon mot from pubic choice economics - "something must be done; this is something; this must be done." I just hope all the new crop of languages are not that. Finally: If you can, and if you can learn the right patterns ( isn't that true of all languages, though?), the thing I have gone to, again and again that is better than test equipment in terms of reliability continues to be Tcl.
- qznc 10y agoLibc does not use your safe array. You cannot pass your safe array to read and other syscalls.
- tedunangst 10y agoAs others have noted, one can also write a four line saferead() function.
- geofft 10y agoAnd once you have written a five line safewrite() function, a hundred-line saferecvfrom() function, a 1500-line safeioctl() function, and all of that, and you have a third-party static analysis tool to prove that you're never calling the unsafe read() or ioctl() functions except from the wrappers, what was the advantage of staying in C in the first place?
- jstimpfle 10y agoI see where you are going, but I don't get how you get from 4 to 5 to 100 to 1500?
- kazinator 10y agosafeioctl may have to contain a redundant switch (or more OOP-style dipatch across multiple functions) on the command code in order to convert the safe style arguments to each specific low level ioctl call. That sort of thing can easily explode in line count. I can see exactly what geofft is talking about.
- jstimpfle 10y agoOk - it's true that ioctl has a generic interface (hadn't remembered that). We can think of it as hundreds of functions, not just one. So wrapping ioctl "completely" is correspondingly a lot of work. However the point was of course to make a safe-land layer which contains the needed functionality. The point was never to have a function "safeioctl" in the first place.
- strictfp 10y agoBecause noone ever does that? Or at the very least people don't think of that as idiomatic c.
- JKCalhoun 10y agoOr perhaps just some helper functions in C that wrap array and pointer allocation/access to provide sanity checks. Seems like moving to a new language is rather extreme....
- jdmichal 10y agoWhy? Rust includes index checks by default, so you have zero chance of programmer error. And, to boot, it removes the checks when the compiler decides that the code is provably safe.
- ue_ 10y agoWhy does Rust still not warn at compile time and furthermore actually crash (panic) on OOB access if it has the checks you describe?
- jdmichal 10y agoI never stated whether the checks were static or runtime. In this case, they are runtime checks. What you are talking about is value range analysis. Rust does a limited amount of it, and if it can prove safe indexing it will remove the runtime checks.
- deleted 10y ago[deleted]
- viraptor 10y agoSure, glib for example contains all the wrapped network / file / string handling you need. But then you need to make sure you use it correctly. And that no code goes around that API. And that you didn't make mistakes that invalidate the safety of your abstraction. Moving to an abstraction would potentially solve some overflow issues and requires constant attention. Moving to a safe language likely solves all of them and you get checked by the compiler instead.
- geofft 10y agoIf you want to do this well, as other commenters pointed out, your entire C library interface is going to have to change to accept bounded arrays instead of pointers and lengths as separate arguments, or worse, just pointers and implicit lengths by looking for the NUL character. But then you've given up the ability to directly call into existing code; everything is at best using a little translation layer to check bounds before/after passing buffers to outside functions. So the advantage of staying in the language is minimal (note that you've given up on "C strings" entirely), and if you're going to have all of this, you might as well use a language with compile-time checks (types, etc.) that you're doing all of this right. Rust and Go are not your only choices here. C++ with aggressive use of modern features and aggressive non-use of everything inherited from C is also a fine choice, but enforcing that discipline is hard, and the compiler won't help you a lot. Honestly, a dynamic language like Python or Ruby is also a fine choice in many cases, although NTP might be too latency-sensitive for that to work. (But it might not be! Premature optimization and all that.)
- naasking 10y ago> i just don't understand why are buffer overflows such a huge problem in C when the same thing is going to come up when trying to work with memory in Rust. False. Buffer overflows in C can overwrite the program's memory, so it can be hijacked and supplanted with the attacker's code. This cannot happen in Rust (unless unsafe code has the vulnerability), or any memory safe language. Sure you can implement a safe array/buffer abstraction and use it in your C programs that abort on invalid indexing. Now how many actually do this? Very few given the prevalence of C programs on vulnerability disclosure lists.
- lossolo 10y agoI think he means to work with raw memory so using unsafe keyword and in this case he is right. And you can't implement certain things in Rust if you are on the quest for maximum efficiency without using unsafe.
- eridius 10y agoWorking with `unsafe` code in Rust has the same potential issue to be sure, but very little (if any) of your Rust code should be `unsafe`. You can usually accomplish what you want without having to drop down into raw C pointers. And everywhere that you do have to do this is explicitly called out with the `unsafe` keyword so it's very easy to audit.
- nothrabannosir 10y agoYou can't implement certain things in C if you are on the quest for maximum efficiency. Thankfully, one rarely is, because efficiency isn't binary. It's about trade-offs.
- lossolo 10y agoI could say exactly the same for Rust, that it's not enough even using unsafe keyword and I need to go into assembly. Someone could even say that assembly is not enough and we need FPGA and then ASIC etc. But it's not the point here. The point is that for ex. you can't get safety from out of bounds access without bound checking, doesn't matter which language you use. And bound checking for certain hot paths is not acceptable.
- staticassertion 10y agoBecause your 'safe' implementation will certainly have a performance cost, and won't be the default. This is why, despite C++ providing std::array, you'll still find buffer overflows in C++ code. C++'s std::array provides the safe 'at' function but you're opting into a performance penalty and it's not the more familiar [] syntax. Rust arrays/ vectors are safe-by-default. To use the unchecked, unsafe version requires using the 'unsafe' keyword. let v = vec![0, 1, 2]; unsafe { let x = v.get_unchecked(5); } This means you can basically grep audit for vulnerabilities, and the above code should be very rare.
- sidlls 10y agoHow do you suppose runtime bounds checks are done in Rust? They certainly also incur a performance penalty in not-trivial cases. Also, "safe by grep audit" means "safe according to a human." The argument of course is that it lowers the surface area of what a human must be trusted to verify. I'm still not convinced by that argument, because human error is a thing. And for actual systems programming, "very rare" may not be true.
- staticassertion 10y ago> How do you suppose runtime bounds checks are done in Rust? They certainly also incur a performance penalty in not-trivial cases. Certainly. I didn't intend to imply otherwise. > Also, "safe by grep audit" means "safe according to a human." Again, totally correct here. > The argument of course is that it lowers the surface area of what a human must be trusted to verify. I'm still not convinced by that argument, because human error is a thing. And for actual systems programming, "very rare" may not be true. Well, given a codebase where both safe and unsafe code exists, the amount of unsafe code is strictly less than the amount of both safe and unsafe code. So it does reduce the amount of code needed to audit, even in a very atypical case where a ton of the code is unsafe. It's true that a project may use egregious amounts of unsafe. That would be unfortunate. Rust is still safer than C in that case, since it just defines more behavior (like arithmetic overflow), but I certainly wouldn't pretend that the rust code should be trusted. When writing rust one should certainly strive to write less unsafe code, and to always document the invariants required for unsafe code to be safe. Rust is not 100% safe 100% of the time, I'm only arguing that safe defaults are critical, and that grep auditing is a powerful tool.
- jstewartmobile 10y agoWhen writing modern C in a disciplined way, it is not as bad as the Rustophiles will make it out to be, but still a problem. Scenarios where I frequently end up fixing other people's memory errors: 1. No Error Handling: not checking an error condition on a function that allocates, then using the uninitialized pointer anyway 2. Sloppy Error Handling: jumping to abort from an error without freeing what has already been allocated 3. Faith in \0: still using the old string functions I'm on the fence about the whole thing, so others may be able to field something more compelling.
- pjmlp 10y agoTry fixing those errors in binary libraries or enterprise code.
- moosingin3space 10y agoA properly-designed Rust API will not allow code without error handling to compile, so 1) should be much less relevant in Rust. 2) I can't remember if a panic in Rust calls destructors, which would clean up that memory. Can someone answer that please? 3) is only relevant in FFI scenarios in Rust, and everywhere else is irrelevant because Rust does not require the use of such footguns for basic string manipulation.
- steveklabnik 10y agoA panic _may_ call destructors, but it cannot be relied upon. For example, aborting is a perfectly reasonable panic implementation, and your destructors won't get called. If you have unwinding panics, they will call destructors though.
- Manishearth 10y agoI mean, in case of an abort the kernel generally cleans things up for you :) Not all things, but most things. More exotic panics (like an infinite loop panic on a microcontroller) would not call destructors though. Most panic impls will either unwind (which will call destructors) or abort/exit, so this is usually not a problem.
- ssalazar 10y agoI think it breaks down when that array has to interact with system libraries or the C stdlib in any way. A lot of C string functions have weird gotchas related to terminators and sizes, and any IO you're doing will involve raw buffers being passed into or out of a system IO function that doesn't understand custom array types.
- jdmichal 10y agoObviously, one can program C to do anything, and write all the provably safe abstractions wished. But, that's not really the point. The point is that doing such is not the default. It requires engagement and knowledge of the programmer, especially on distributed projects with loose communication, such as many open source projects. And it only takes one programmer mistake to bring the whole house of cards down. Why allow programmers to make mistakes? That was fine in the 70's when resources for compiler execution were limited. I don't see any reason for it today. I mean, just look at the underhanded C contestants and especially winners for ways in which your program can completely blow up for extremely subtle reasons.
- panic 10y agoAnd it only takes one programmer mistake to bring the whole house of cards down. Isn't this also true of Rust, with its unsafe keyword? None of these languages are completely safe against programmer mistakes.
- afarrell 10y agoIs the use of the unsafe keyword a default? I don't know Rust, but from a user interface perspective, it sounds like it has an affordance of "Hey! Pay particular attention to this bit because it is risky!"
- steveklabnik 10y agoIt is very much not the default.
- Manishearth 10y agoIt's more of a "Hey, trust me here when I say that the enclosed code is actually safe" hint to the compiler. It's used sparingly. Not as sparingly as I'd like, but sparingly enough.
- afarrell 10y agoSure, but code is not just instructions to a compiler, but is also a user interface.
- dbaupp 10y agoThe lack of generics means your array implementation is either going to either: - be implemented with macros and token pasting, and result in a ton of mental overhead because you'll have a pile of types like array_foo for an array of `foo`s, and array_bar for an array of `bar`s, along with a pile of corresponding `foo * array_foo_get(array_foo, size_t)` and `bar * array_bar_get(array_bar, size_t)` functions. - or, have a runtime cost and lose type safety by storing void* and casting when accessing. The first case is even worse than it sounds: e.g. I don't know how you handle arrays of types with spaces in them (like `unsigned char`, or `struct bar`) with a macro. And, we haven't even thought about const correctness yet, which would probably require having const_array_foo, const_array_bar (etc.) types defined too. (And, of course, these only solve one facet of the problems with C's pointers: there's no way to defend against use-after-free or dangling pointers.)
- nwmcsween 10y agoYou're missing main glaring issue with parametric polymorphism, bloat.
- dbaupp 10y agoWhile "unnecessary"/extra generated code is a trade-off one has to consider when choosing to use the specialize/monomorphise-everything implementation of parametric polymorphism (it isn't the only one), it isn't a problem in this case: the types themselves are a compile-time abstraction and don't exist at runtime, and the functions are all tiny (a branch, a memory access and a function call/abort). Additionally, all the functions should be inlined anyway because the function call overhead will likely be as much or more than the actual code, and, more importantly, inlining enables other optimisations (removing the branch, vectorising the memory access, etc.). Once inlined, the code will be the same as the manual/macro-based approach of writing `if` statements around each array[index] access.
- dreta 10y agoNo, you just allocate enough space to store an extra int at the start for the length, and return a typed pointer to the actual data. Then you need an accessor that checks bounds, if you want safe access. Both of these problems are solved by simple macros.
- MaulingMonkey 10y ago> Can anybody make a strong case to me as to why are buffer overflows considered an issue in C when it takes like 10 minutes to write and test an array implementation that prevents that from ever happening? The CVE database. Just because you 'can' write such an array implementation doesn't mean you will, doesn't mean your third party libs will, doesn't mean any of your legacy code uses it, and certainly doesn't mean you will properly test said array implementation correctly. The number of mitigations added to C compilers and OSes dealing mostly with C and C++ code. ASLR, W^X, /GS, -fstack-protector-all, AddressSanitizer, ... - note the lack of similar tools, or demand for them, for, say, JavaScript - despite it enjoying a similar ubiquity. I ask this in bad faith: I encourage you to share a single nontrivial codebase which actually creates the abstraction you've described and religiously adheres to using it throughout. As to why this is in bad faith: I'm definining "nontrivial" here to mean using 3rd party APIs - which will operate on C style arrays, not your project specific safe wrappers - and thus by definition won't be "religiously" sticking to said abstractions when using said APIs. By these definitions, the codebase I'm asking for doesn't exist - by definition. Even relaxing the "third party" rule, I haven't actually worked on a nontrivial C or C++ codebase without buffer overflow problems. Now, e.g. Rust will have the same problems when interacting with C APIs - and nontrivial programs will end up doing so eventually. However, by virtue of the language itself embracing safe-by-default, you're less likely to run into the same problems when consuming Rust APIs. You can also use third party static analysis tools to ensure you're using a "safe C subset" (such as MIRSA C), but "nobody" does that.
- agumonkey 10y agoIt's crazy that it's not solved above the language level, if people really want zero cost abstraction and architecture friendliness at least tooling should check buffer logic and flag the binary in case Warnings have been ignored.
- dhsjxhx 10y ago>if people really want zero cost abstraction That one "if" is (by definition) not zero-cost.
- pcwalton 10y agoBecause: 1. Buffer overflows aren't considered the most insidious issue in C nowadays. That award would probably go to use after free, which is not so easy to fix. 2. In C, it is easier and faster to do the wrong thing. Compare "char buf[256]; strcpy(buf, foo); ..." to "array_t buf = array_create(strlen(foo) + 1); strcpy(buf.ptr, foo); ... array_destroy(buf);" 3. Buffer overflows do not in fact come up routinely in Rust the way they do in C.
- thesmallestcat 10y agoI'm no C expert, but those do two different things (fixed vs. dynamic length, stack vs. heap alloc). And the first one is safe with strncpy, right? Though I don't know why you'd ever have code like this unless you like undefined behavior and want to return the array. The second example seems fine and not particularly difficult or verbose.
- dibanez 10y agoIts useful to be able to remove safety checks for speed. I have a C++ code where all data is in array objects. Bounds checking is a compile time option, and it makes the overall code 2X slower. I can do testing with bounds checking on, but once it gets to a supercomputer that needs to be removed. Address sanitizing by compilers is an even more effective tool for this, especially for C. Bounds checking is critical for security, but if you're only concerned with correct execution then a segfault is not much different from an exception.
- ChemicalWarfare 10y agotechnically you probably could limit yourself to using a "safe" subset of C - basically no pointer arithmetic, no strcopy() etc - but that would defeat the purpose of using C in a first place.
- Sanddancer 10y agoCompiler vendors have been resistant towards putting in such features. Bounds checking slows things down, and the performance race is very much a thing in C compiler implementations -- a compiler that can deliver a few percentage points better code can be a big win to teams working on compute heavy problems. C11 has Annex K which has a lot of safety features, like memory safe arrays. Unfortunately, none of the vendors have implemented it even as an option. Which is a shame because it would solve a lot of problems, with requiring minimal rewrites for a lot of code.
- pjmlp 10y agoAnnex K was required in C99 and made optional in C11, it tells everything on how C vendors see safety. Also it is actually a joke, since it still separates the pointers and length in two separate variables, instead of using some kind of struct. The only thing it does is have functions with better semantics on the terminating nulls.
- josteink 10y ago> Can anybody make a strong case to me as to why are buffer overflows considered an issue in C when it takes like 10 minutes to write and test an array implementation that prevents that from ever happening? That it's not a by-default and forced language-feature and that most developer aren't going to spend those 10 minutes when they need an array. They'll just use the language-provided array-implementation instead. Which in C is very, very unsafe.
- gsdean 10y agoI believe your post is the strong argument you desire aka hubris