5 ms·
I think this is an indicator how fundamentally over-engineered all of the GNU tools are. strings was supposed to be a simple tool that finds bits of data that l
by _ak 12y ago
I think this is an indicator how fundamentally over-engineered all of the GNU tools are. strings was supposed to be a simple tool that finds bits of data that look like human-readable strings. It wasn't meant to parse ELF binaries and suddenly be a security risk, especially since it's one of the first tools you would use in computer forensics.
- _delirium 12y agoI don't think that's the root of the problem. While you could make a simpler strings(1), which would help people who only use that one, more complex stuff like objdump(1) really does need to parse binaries. And that should be possible to do without worrying about security problems: you're just reading a file and extracting some information, which even in the worst case should be possible to do without accidentally executing arbitrary code. It's just that libbfd seems to have a lot of bugs, and because it's written in an unsafe language, such bugs can not only cause incorrect information or crashes, but sometimes attacker-controlled code execution. But if you don't do it via libbfd, you're going to need some library that can parse binaries, since many utilities end up needing to do it, and it shouldn't be impossible to safely do so. An alternative is to implement a subset of full parsing specifically tailored to each utility. In the case of strings(1) that's very simple; in the case of some other utilities it's of intermediate complexity; all the way up to some that need to parse every corner of ELF. Whether that produces a bigger or smaller attack surface depends on a lot of factors: each parser might be simpler, but there are many more of them. FreeBSD was contemplating centralizing more of that into a common libelf, so I don't think it's only GNU who think that's a good idea in principle: https://wiki.freebsd.org/LibElf https://wiki.freebsd.org/LibElf
- cbd1984 12y agoAlthough perhaps not in this case, it's generally true that one person's bloat is another person's functionality. BSD tools are simpler. They're also generally less useful.
- saurik 12y agoThe goal of the tool "strings" has always been to dump the strings table of a binary object. It happens to also have a mode that lets you try to find random strings-like content in any file. It happens to default to this mode if it can't parse the file as an executable. People thereby have gotten somewhat used to using this tool to having this functionality at hand, and use this tool a lot of this purpose. This is not, however, the actual goal or purpose of this tool. The fact that many people use Perl as nothing more than a slightly better version of sed doesn't mean that Perl's ability to write complex object-oriented software is "over-engineering". You just don't know what this tool is actually for, which is OK, but means you can't judge whether or not it is "over-engineered". BSD, apparently going back to at least BSD 4.3, also had a strings tool, and it did the exact same thing: it parsed binary files to dump their strings table. Apple's strings tool has no code heritage from the GNU version, instead being a vague descendent of the one from BSD 4.3. This is how this tool has always worked: stop being part of the noise trying to turn this into a GNU-bash fest :/.
- e12e 12y ago> The fact that many people use Perl as nothing more than a slightly better version of sed doesn't mean that Perl's ability to write complex object-oriented software is "over-engineering". Not sure I agree with you on Perl in particular, but other than that I agree with what you're saying ;-) (That is, I don't think GNU being over-engineered is the (or perhaps even "a") problem here. >> Am I the only person who thinks there's something fundamentally wrong with computing if running "strings" could let someone take over your computer? No, the underlying library (libbfd) is an example of something that should've been fixed a long time ago. Maybe not quite the horror that was/is openssl -- but clearly an example of "old c code that sort of works" -- perhaps in some ways like bash was/is. It's old, it works, but it could use some clean-up (as evidenced by a number of related buffer under/over-flows and whatnot). Note that parsing arbitrary binary (or otherwise) input safely, is a pretty hard problem. There was recently a (resource DOS) bug in libxml2, which have been under quite a lot of scrutiny lately (by virtue of being a brilliant injection vector for malicious code, if a bug can be found). I read this as a two-part bug: one, a lot of people didn't know that strings did more complex parsing than a hexdump and a filter for printable strings (me included) -- and it turns out that the "smart" library isn't terribly robust (in other words: it's typical C code). While it is possible (at least in theory) to write small C utilities that are safe, if you want them to be (wildly) portable, and have sane handling of various kinds of encoded strings, and encoded data, along with handling different endianess -- apparently most people screw up. I think there are two basic camps wrt what should be done: those that think we need something like rust, so that we can have safety without much of a slowdown, and those that say screw it, we're no longer running on 5Mhz (or 50Mhz) cpus, we can take anything up to a 100x (10x) slow down without it really being an issue -- Security/stability/predictability is more important. Those that can't decide between the two, continue writing C like it was still 1989, and we get lots of stuff like this. I'm not sure if it's usually the mix of a "smart" C programmer writing a program that is patched by a "hobby" C programmer, or the fact that getting C right is just too hard -- or that people don't -Wall and don't run fuzzers and static checkers -- but whatever the reason, we keep seeing serious bugs in C programs. I'd like to think some of it could be avoided if people wrote more "bloated" C with copious use of functions, more call-by-value, smaller loops, perhaps more computer generated code -- and other "slow" things (while still being C). But I'm probably hopelessly naive. While I'm definitely not sold on C++ (at least not as viable "better" C for systems programming), I think the old 1998 article[1] by Strostrup on "simple" C and C++ programs illustrate quite well how hard C can be to get reasonably right, for even simple problems. Perhaps rather than waiting for Rust, a reimplementation on large parts of the backbone of our OSs/GNU in Guile, Lua, or some other higher-than-C level language could be worthwhile. As a side note -- does anyone know of any follow up on Stroustrup's article? [1] http://www.stroustrup.com/new_learning.pdf http://www.stroustrup.com/new_learning.pdf
- sneak 12y agoThe problem is not additional features, the problem is unsafe parsers. Adding features is natural and healthy. The damage is that we've been so afraid of parsers for so long that we associate "let's understand this bytestream better so we can be more useful" with implicit danger. Some friends of mine are tackling this problem. You should help them. http://langsec.org http://langsec.org