6 ms·
Why the OpenSSL punycode vulnerability was not detected by fuzz testing (2022)
- stcredzero 2y agoWould it be possible to attach an LLM to a debugger sessions executing all of the fuzzer seeds, and ask it to figure out how to expand coverage?
- swatcoder 2y agoNot to dismiss the novelty and power of LLM's but why would you turn to a black box language interface for that? Wouldn't you expect a designed system to be more efficient, complete, reliable, and auditable? Most or those characteristics are critical to security applications and the nature of LLM's largely run counter to them.
- agilob 2y agoTurns out LLMs are really good at writing fuzzers https://verse.systems/blog/post/2024-03-09-using-llms-to-generate-fuzz-generators/ https://verse.systems/blog/post/2024-03-09-using-llms-to-gen... https://google.github.io/oss-fuzz/research/llms/target_generation/ https://google.github.io/oss-fuzz/research/llms/target_gener...
- cjbprime 2y agoThese links are a little different to the GP comment, though. Both of these cases (which I agree show LLMs being an excellent choice for improving fuzzing coverage) are static analysis, going from the project source code to a new harness. Some issues with that are that the model probably doesn't have enough context to be given all the project source code, so you have to work out which subset to share, including definitions of all relevant symbols (but not too many). It helps a bit that the foundation models were already pre-trained on these large open source projects in oss-fuzz, so they already know something about the project's symbols and definitions from their original training sets -- and even from public discussions about the code! -- but that wouldn't work for a private codebase or for recent changes to a large public one. Then the harness source that the LLM writes might have syntax errors/fail to compile, and you have to deal with that somehow, and the harness source that the LLM writes might be valid but not generate any coverage improvement and you have to deal with that, and so on. GP seems to be talking about instead some form of LLM-aided dynamic analysis, where you are probably using some kind of symbolic execution to generate new seeds, not new harnesses. That's important work too, because I think in this case (disagreeing with the blog post author) the vulnerable function was actually reachable by existing harnesses, just not through the seed corpora (at least the public ones). One approach could be for the LLM to become a kind of a symbolic execution constraint solver, using the debugger as a form of path instrumentation and producing new seeds by predicting what a new input would look like when you invert each interesting constraint that the fuzzing coverage is blocked by, as the debugger hits the test for that constraint (which is difficult because it can require actual computation, not pattern matching, and because of path explosion). Or perhaps more plausibly, the LLM could generate Z3 or other SAT-solver code to define and solve for those constraints to generate new seeds, replacing what is currently extremely tedious and time-consuming work when done by humans.
- swatcoder 2y agoThose demonstrate that they're capable of generating capable ones, which is really cool but also not surpising. What matters for engineering is how that technique compares to others in the context of specific requirements. A big part of "falling for hype" is in mistaking a new and capable tool for the being the right or oprimal tool.
- cjbprime 2y agoIt's fine to have LLM skepticism as a default, but here it's not justified. Google is showing here that the LLM-written harnesses improve massively on the harnesses in oss-fuzz that were written over many years by the combined sum of open source security researchers. Most dramatically, they improved tinyxml2 fuzzing coverage by 31% compared to the existing oss-fuzz harnesses, through an entirely automated flow for harness generation by LLMs. Whatever engineering technique you are imagining would be better is not one that humanity actually applied to the problem before the automated LLM harnesses were written. In general, writing and improving fuzzing harnesses is extremely tedious work that is not being done (or paid for) by nearly enough people to adequately protect critical open source software. The LLM approach is a legitimate breakthrough in the field of open source fuzzing.
- swatcoder 2y agoFair enough, interesting, and plausible! I looked at the first link and saw it as more of a capabilities demo, but didn't dig into the Google one. I'm mostly just encouraging thoughtful reflection on tool choice by raising questions, not making a case against. Very cool.
- _flux 2y agoI suppose it would be nice to require at least 100% line coverage for tested encryption-related functionality, when tested by a fuzzer. Deviations from this could be easily detected in CI. Testing cannot be used to prove that a flaw doesn't exist, only that it does.
- Retr0id 2y ago> Testing cannot be used to prove that a flaw doesn't exist This is not universally true, formal methods can take you a long way, depending on the problem domain. And sometimes you can test for every possible input. This usually isn't feasible, but it's nice when it is.
- dllthomas 2y agoFormal methods aren't "testing". As you say, though, exhaustive testing is sometimes possible.
- bluGill 2y agoIt is normally safe to assume the exhaustive testing isn't possible because the total states to exhaustively tests exceeds the number of atoms in the universe. There are a few exceptions, but in general we should assume it is always impossible to exhaustively test programs. Which means we need to use something else to find what the limits are and test only those, or use formal methods (both would be my recommendation - though I'll admit I have yet to figure out how to use formal methods)
- dllthomas 2y ago> in general we should assume it is always impossible to exhaustively test programs The whole program? Yes. For testing individual components, there's no need to assume. The answer is very likely either clearly yes (the input is a handful of booleans) or clearly no (the input is several integers or even unbounded) and for those cases where it's not immediately clear (one i16? Three u8s?) it's probably still not hard to think through for your particular case in your particular context.
- TacticalCoder 2y agoWhy is OpenSSL using punycode? To do internationalized domain name parsing?
- amiga386 2y agoI would assume it's needed for hostname validation. Is one hostname equivalent to another? Does this wildcard match this hostname? EDIT: I looked it up. It is used to implement RFC 8398 (Internationalized Email Addresses in X.509 Certificates) https://www.rfc-editor.org/rfc/rfc8398#section-5 https://www.rfc-editor.org/rfc/rfc8398#section-5
- cryptonector 2y agoYes.
- arp242 2y ago(2022) Previous: Why CVE-2022-3602 was not detected by fuzz testing - https://news.ycombinator.com/item?id=33693873 https://news.ycombinator.com/item?id=33693873 - Nov 2022 (171 comments)
- dang 2y agoYear added above. Thanks!
- cjbprime 2y agoI'm not sure the post (from 2022) is/was correct. I've looked into it too, and I expect this was reachable by the existing x509 fuzzer. There's a fallacy in assuming that a fuzzer will solve for all reachable code paths in a reasonable time, and that if it doesn't then there must be a problem with the harness. The harness is a reasonable top-level x509 parsing harness, but starting all the way from the network input makes solving those deep constraints unlikely to happen by (feedback-driven) chance, which is what I think happened here. Of course, a harness that started from the punycode parsing -- instead of the top-level X509 parsing -- finds this vulnerability immediately.
- hedora 2y agoI've found that writing randomized unit tests for each small part of a system finds this sort of stuff immediately. In this case, a test that just generated 1,000,000 random strings and passed them to punycode would probably have found the problem (maybe not in the first run, but after a week or two in continuous integration). I try to structure the tests so they can run with dumb random input or with coverage-guided input from a fuzzer. The former usually finds >90% of the bugs the fuzzer would, and does so 100x faster, so it's better than fuzzing during development, but worse during nightly testing. One other advantage for dumb random input is that it works with distributed systems and things written in multiple languages (where coverage information isn't readily available).
- mattgreenrocks 2y agoI really like this idea because it avoids the issue of fuzzers needing to burn tons of CPU just to get down to the actual domain logic, which can have really thorny bugs. Per usual, the idea of "unit" gets contentious quickly, but with appropriate tooling I could foresee adding annotations to code that leverage random input, property-based testing, and a user-defined dictionary of known weird inputs.
- paulddraper 2y ago> maybe not in the first run, but after a week or two in continuous integration You'd use a different seed for each CI run?? That sounds like a nightmare of non-determinism, and lost of trust in CI system in general.