11 ms·
Open-Sourcing ClusterFuzz
- metzmanj 8y agoI work on this. Happy to answer questions if people have any.
- devy 8y agoHi, good stuff! I am curious to know why it was primarily written in Python (according to GitHub 83.3%)?
- saagarjha 8y agoPython is great for writing glue code, so I'm not particularly surprised by the language choice.
- tracker1 8y agoJust generally speaking, code that does orchestration and testing in general is often easier under a dynamic scripted language over something that is built and compiled, even if it winds up as a custom DSL. I think Python is one of the better options here for the broader community support, and tooling. Aside: I tend to reach for node/js often for similar reasons (despite detractors) mostly because I'm more comfortable with it over Python or Ruby, but also because it's already integrated to most of the build/test environments I'm working on anyway.
- aboutruby 8y agoThat's one of the main languages at Google. And there are moving towards more Go since a few years.
- kortilla 8y agoI’m curious why you seem to be surprised? It’s one of the most popular languages (even within Google).
- AndrewKemendo 8y agoHow does the decision get made at Google to open source something? I'm always confused about what gets shut down vs open sourced vs paid product.
- snazz 8y agoI would guess that it has to do with the usefulness of the project outside of Google. This project could be applied to so many other things (as OSS-Fuzz demonstrates), so open-sourcing it makes perfect sense. It isn’t some kind of classified algorithm, either.
- asdkfjasl 8y agoOpen sourcing is usually pushed from the bottom. People decide they care about open sourcing their project at they push for it.
- joshsimmons 8y agoHi there! I work in the Google Open Source Programs Office. Echoing what others have said, it's usually just a matter of an engineer or team deciding it's something they want to do. Other times, it's a strategic choice. We open sourced our open source policies/docs a while back, so if you're so inclined you can dig deeper there. These two links will be of particular interest: https://opensource.google.com/docs/creating/ https://opensource.google.com/docs/creating/ and https://opensource.google.com/docs/why/ https://opensource.google.com/docs/why/
- metzmanj 8y ago+1 I can speak a little bit about what motivated us. We saw from OSS-Fuzz (https://github.com/google/oss-fuzz https://github.com/google/oss-fuzz) that this sort of thing could be widely useful and wanted non-open source code to benefit from making fuzzing easier.
- AndrewKemendo 8y agoThanks Josh! Love the meta-open sourcing.
- legohead 8y agoI am quite ignorant on this subject. I looked briefly through the docs, and still feel a little lost. So before I go too much further, would it be possible to use this for web apps or unity games?
- metzmanj 8y ago>So before I go too much further, would it be possible to use this for web apps or unity games? Web apps, almost certainly no. ClusterFuzz (and fuzzing generally) is most useful for finding bugs in C/C++ code so maybe it could work for unity games? I don't know much about them though.
- Insanity 8y agoI thought unity was mostly written in either C# or an EcmaScript flavour?
- ritz_labringue 8y agoUnity itself (the engine) is written in C++. Game scripts are written in either C# or UnityScript iirc.
- pjmlp 8y agoAlthough the engine is currently written in C++, they are in the process of rewriting parts of it in C#, with help of their HPC# subset and Burst compiler, having some ex-Insomniac Games developers like Mike Acton on the team.
- xiphias2 8y agoIs this the fuzzing tool that was used to find bugs in Bitcoin? I don't see it in the repository.
- metzmanj 8y agoI don't think so. But solidity was recently added to OSS-Fuzz: https://github.com/google/oss-fuzz/tree/master/projects/solidity https://github.com/google/oss-fuzz/tree/master/projects/soli...
- wyldfire 8y agoKeep up the good work, Johnathan! This is a great feature.
- q3k 8y agoAny plans to port this from GCE APIs to something more provider agnostic, like k8s? A lot of us would like to fuzz on on-prem equipment.
- metzmanj 8y agoWe would like to support this use case. For now, you can actually do the fuzzing on prem while communicating with app engine. We do this for our OS X bots since GCE doesn't offer OS X.
- deleted 8y ago[deleted]
- metzmanj 8y agoThe other possibility for completely on-prem use right now is running it using the dev server: https://google.github.io/clusterfuzz/getting-started/local-instance/ https://google.github.io/clusterfuzz/getting-started/local-i...
- j88439h84 8y agoHow does it compare to afl?
- metzmanj 8y agoIt uses AFL. ClusterFuzz is infrastructure for running fuzzers, so we use it to run AFL, libFuzzer, and other domain specific fuzzers we've written. Using it to run AFL gives us a lot of nice things over using AFL on someone's desktop (such as crash deduplication, automatic issue filing, fixed testing, regression ranges etc.)
- tanin 8y agoCongrats, Jonathan!
- metzmanj 8y agoThanks Tanin!
- bobwaycott 8y agoFor those interested in the repo: https://github.com/google/clusterfuzz https://github.com/google/clusterfuzz
- syastrov 8y agoMakes you think about choosing to write software in C / C++ / other non-memory-safe languages when you need 25000 cores churning away to ensure you don’t make mistakes that could cause serious security issues. It makes me wonder why Google wouldn’t put their efforts into using Rust, for example. Of course, server power is cheap, but not for our planet.
- tracker1 8y agoWell, don't most of their internal systems already run tooling that is already written using Linux and Gnu libraries? Wouldn't it make sense to do testing on what they already have available instead of some theoretical ground up replacement effort? Not to say there isn't work on replacement efforts, only that testing what is in use is a practical thing, even if the scale is surprising.
- twblalock 8y agoI know some embedded/kernel devs and they don't take Rust very seriously. I don't think it has a lot of mindshare in industry among the kind of people who currently write C. Even if all new code was written in safe languages we would still need to do fuzzing until all legacy code was rewritten -- operating systems, SSL libraries, browsers, load balancers, etc.
- pcwalton 8y agoI know lots of embedded/kernel devs (I am one) and lots of them are at least interested in Rust. It does have mindshare.
- master-litty 8y agoI truly believe that's the only reason people aren't using Rust as a replacement for C / C++ immediately. The adoption isn't widespread, but I'm rather optimistic it will be. It's getting there.
- twblalock 8y agoI don't think it's the only reason. My C programmer friends like the dangerous features of the language and want to use them. Their programs are designed around the assumption that they can mutate anything in memory whenever they want to. They use mutable global variables. They would strongly dislike Rust's memory protections and ownership concept, and its other safety features.
- polskibus 8y agoIs there a fuzzing tool oriented towards web applications? Something that could generate loads of Selenium cases automatically and verify whether the application crashes, logs an exception or continues to work smoothly??
- zdragnar 8y agoThere are a boatload of pentesting (i.e. penetration testing) tools that use fuzzing. Just be sure that your sysadmin and / or your cloud provider are aware that you intend on running such a test, as that's a really quick way to accidentally bring down servers you weren't anticipating would be connected, or get your IP address banned for DDOS attacks (looking at you, junior QA guy who had good intentions but caused all sorts of havoc). Edit: Just realized I didn't quite address your question fully. https://pentest-tools.com/home https://pentest-tools.com/home is an online service that will run tests, including URL fuzzing and what not. All of the features they offer can also be found in open source and proprietary software. Not sure about saving failed tests as selenium tests for re-running in the future, though I imagine that you'd just re-run the same tool in the first place.
- polskibus 8y agoWhat open source did you mean? I'd like to run such tests on the intranet, without access to any SaaS.
- babayega2 8y agoLast time I tested https://github.com/s0md3v/XSStrike https://github.com/s0md3v/XSStrike whith some quite interesting results. It's important to note that it's not a fizzbuzz tool, but just a pentest one.
- Insanity 8y agoPerhaps a noobie question, but it mentions c/c++ specifically. How does this hold up for Go? Where you have pointers but no pointer arithmetic?
- metzmanj 8y agoThere are tools for fuzzing go: https://github.com/dvyukov/go-fuzz https://github.com/dvyukov/go-fuzz But I think the kinds of bugs found by fuzzing aren't generally security issues in go (I don't know much about go) as they are in C/C++. EDIT: See guidovranken's excellent sibling comment for how fuzzing can still be useful for go.
- Insanity 8y agoAwesome, I will check that out. Thank you!
- guidovranken 8y agoGo software can exhibit a variety of denial-of-service bugs such as slice out-of-bounds access (since there is no try/catch mechanism, this leads to a panic), excessive allocations, excessive computation/timeout (consider "for i := 0; i < N; i++" where N is untrusted), stack overflow due to unbounded recursion (rare because Go has a custom, large stack). My bignum-fuzzer project [1] runs on oss-fuzz and tries to find mismatches between bignum computations across different libraries (OpenSSL, Go, Rust, etc). This is one example of how fuzzing can be useful even if the underlying language is "safe". With some small hacks you can also have Go code coverage instrumentation as a libFuzzer counter. [1] https://github.com/guidovranken/bignum-fuzzer/blob/master/modules/go/lib.go https://github.com/guidovranken/bignum-fuzzer/blob/master/mo...
- staticassertion 8y agoAnd Go isn't memory safe given race conditions. If you're using goroutines you may want to consider fuzzing with the race detector.
- Insanity 8y ago
- painful 8y agoHow about people stop using unsafe languages such as C and C++?
- saagarjha 8y agoEven if everyone switched to writing Rust/Swift/$SAFE_LANGUAGE today, there are many, many projects that are written in C/C++ that aren't going to disappear overnight.
- sametmax 8y agoFuzzing is effective for more than finding memory related bugs. Parsing related crashed, code injection, unexpected state machine path can all also be revealed. Beside, as much I agree with the benefits of a memory safe language, and I do believe in the urgency of promoting techs like rust, C and C++ are going to be part of our lifes for a lot of time. If you gotta use knife, make sure you have the tools to make it sharp.
- paulgdp 8y agoRust without the use of unsafe code is considered by most to be a safe language but nonetheless, it's still useful to fuzz it to find logic bugs (like assert failures) or anything that triggers a panic (like out of bound array access or integer overflow). More info: https://github.com/rust-fuzz https://github.com/rust-fuzz Disclaimer: I am the author of the rust fuzzer honggfuzz-rs.
- stormbeard 8y agoWhat planet do you live on? What do you think embedded/realtime systems, signal processing, graphics, and kernel developers are supposed to use? Also, what do you think these memory-safe, garbage collected, runtime environments are written in?
- saagarjha 8y ago> What do you think embedded/realtime systems, signal processing, graphics, and kernel developers are supposed to use? My guess is that they'd use Rust for new code.
- rarecoil 8y agoThank you for open sourcing this. For those interested in trying multiple cluster-based fuzzing solutions, I'd also like to point at yahoo/yfuzz[1], which is k8s-backed. [1] https://github.com/yahoo/yfuzz https://github.com/yahoo/yfuzz
- guidovranken 8y agoI don't want to hijack the thread subject but here are my thoughts on the usefulness of fuzzing of safe languages. Even in the absence of memory corruption bugs there is a subclass of bugs that can emerge in any general-purpose language, like slowness/hangs, assert failures, panics and excessive resource consumption. Barring those, you can detect invariant violations, (de)serialization inconsistencies (eg. deserialize(serialize(input)) != input, eg. see [1]), different behavior across multiple libraries whose semantics must be identical (crypto currency implementations are notable in this regard as deviation from the spec or canonical implementation in the execution of scripts or smart contracts can lead to chain splits). With some effort you can do differential 64 bit/32 bit fuzzing on the same machine, and I've found interesting discrepancies between the interpretation of numeric values in JSON parsers, which makes sense if you think about it (size_t and float have a different size on each architecture, causing the 32 bit parser to truncate values). This might be applicable to every language that does not guarantee type sizes across architectures like Go (not sure?), but I haven't tested that yet. You can detect path escape/traversal (which is entirely language-agnostic but potentially severe) by asserting that any absolute path that is ever accessed within an app has a legal path, or by fuzzing a path sanitizer specifically. And so on. Code coverage is the primary metric used in fuzzing, but other metrics can be useful as well. I've experimented extensively with metrics such as allocation, code intensity (number of basic blocks executed) (which helped me prove that V8's WASM JIT compiler can be subjected to inputs of average size that take >20 seconds to compile), and stack depth, see also [2]. Any quantifier can be used as a fuzzing metric, for example the largest difference between two variables in your program. Let's say you have a decompression algorithm that takes C as an input and outputs D. Calculate R = len(D) / len(C), so that R is the ratio between compressed input and decompressed output. Use R as a fuzzing metric and the fuzzer will tend to generate inputs that have a high compressed/decompressed size ratio, possibly leading to the discovery of decompression bombs [3]. Wrt. this, libFuzzer now also natively supports custom counters I believe [4]. Based on Rody Kersten's work I implemented libFuzzer-based fuzzing of Java applications supporting code coverage, intensity and allocation metrics [5], and it should not be difficult to plug this into ClusterFuzz/oss-fuzz. Feel free to get in touch if you have any questions or need help. [1] https://github.com/nlohmann/json/blob/develop/test/src/fuzzer-parse_json.cpp https://github.com/nlohmann/json/blob/develop/test/src/fuzze... [2] https://github.com/guidovranken/libfuzzer-gv https://github.com/guidovranken/libfuzzer-gv [3] https://en.wikipedia.org/wiki/Zip_bomb https://en.wikipedia.org/wiki/Zip_bomb [4] https://llvm.org/docs/doxygen/FuzzerExtraCounters_8cpp_source.html https://llvm.org/docs/doxygen/FuzzerExtraCounters_8cpp_sourc... [5] https://github.com/guidovranken/libfuzzer-java https://github.com/guidovranken/libfuzzer-java
- boulos 8y agoDisclosure: I work on Google Cloud. I'm super pleased to see this! Abhishek and the cluterfuzz team were one of our initial customers for Preemptible VMs, still are, and make for a great example. Congrats to the team!
- deleted 8y ago[deleted]