12 ms·
Don’t underestimate grep-based code scanning
- K0nserv 7y agoJust a small note that I would highgly recommend ripgrep[0] over standard grep. It's another modern tool that has been created by leveraging Rust and it's from BurntSushi[1] who is excellent. 0: https://github.com/BurntSushi/ripgrep https://github.com/BurntSushi/ripgrep 1. https://github.com/BurntSushi https://github.com/BurntSushi
- simias 7y agoIn general I think that's very good advice. In this particular instance however a dumb old grep might be superior because it could catch potential security vulnerabilities that are not explicitly hardcoded in the source code by greping through the compilation artifacts for instance. Sure you'll get a bunch of false positives that way but at least you know that nothing is slipping through the cracks.
- burntsushi 7y agoTo be clear, you can disable all smart filtering in ripgrep. e.g., `rg -uuu foo` should be equivalent to `grep -r foo`. And if you want to exhaustively search binary files, then you need to add the `-a` flag to both commands.
- pepper_sauce 7y agoWhy is it better than grep?
- Canada 7y agoIt's way faster, which is great when you're working with big repos. It's designed for recursively searching through a lot of files.
- masklinn 7y agoGrep is really fast at that at the actual search (gnu grep at least), the gain there is mostly that "smarter" tools will ignore e.g. VCS data or binary files by default whereas grep will trawl through your PNGs and git packfiles.
- icebraining 7y agoExplanation from the original author on why GNU grep is fast: https://lists.freebsd.org/pipermail/freebsd-current/2010-August/019310.html https://lists.freebsd.org/pipermail/freebsd-current/2010-Aug... Excerpt: "The result of this is that, in the limit, GNU grep averages fewer than 3 x86 instructions executed for each input byte it actually looks at (and it skips many bytes entirely)."
- masklinn 7y agoThere's also this bit: https://ridiculousfish.com/blog/posts/old-age-and-treachery.html https://ridiculousfish.com/blog/posts/old-age-and-treachery.... However note https://news.ycombinator.com/item?id=19522987 https://news.ycombinator.com/item?id=19522987 > It does not. ripgrep does not use Boyer-Moore in most searches. > In particular, the advice in [the freebsd mailing list post] is generally out of date. although the out of date bits are really the Boyer-Moore ones: https://lobste.rs/s/ycydmd https://lobste.rs/s/ycydmd > much of Mike Haertel’s advice in this post is still good. The bits about literal scanning, avoiding searching line-by-line, and paying attention to your input handling are on the money. > But the stuff about Boyer-Moore is outdated.
- oconnor663 7y agoIf I remember that giant post of benchmarks correctly, there are some big exceptions, particularly around non-ASCII searches.
- Annatar 7y agoHowever fast it is, it's going to have a tough time beating /usr/bin/fgrep.
- K0nserv 7y agoLike a few other tools — ack, ag, and pt — it's specialized for source code, in addition to that it's really fast. The repo contains detailed comparisons with grep and an FAQ.
- seren 7y agoBlog post from the author with pros & cons of rg https://blog.burntsushi.net/ripgrep/ https://blog.burntsushi.net/ripgrep/ (It seems down at the moment you can try the cached version)
- deleted 7y ago[deleted]
- Annatar 7y agoThis program re-invented find + egrep. It is written in Rust which isn't available on anything other than Linux, OS X and Windows. I'll just stick to /usr/bin/find and /usr/bin/egrep programs. That works just fine without having to port Rust before trying to build a re-invention of find and egrep.
- adrianN 7y agoIt "reinvented" it and made it dramatically faster. Rust is also available on FreeBSD. I don't think you actually need Rust to run ripgrep. It's not like it's an interpreted language.
- Annatar 7y agoSince when do FreeBSD executables run on the illumos family of operating systems, as far as I know, there is no freebsd-branded zone yet in illumos? Dramatically faster? Is there any scientific evidence to that? If it were me, I wouldn't rush to make assumptions.
- adrianN 7y agoI tried it and it's much faster for me. The author did some benchmarks. I'm not about to publish a paper. I also never said anything about Illumos.
- dmit 7y agoIllumos is also a supported platform (SPARC and x86-64).
- Annatar 7y agoYou obviously haven't tried it on either of those. They are "second tier", which means one is completely on one's own. What in your opinion would have to be the size of the source code to warrant jumping through the hoops to get this software running, as opposed to a combination of find + xargs + egrep,fgrep,awk?
- _kbh_ 7y agorg + fzf really makes for a great toolset for some quick first pass code review.
- secure 7y agoOne thing which was not immediately obvious to me for a while: the stricter your language’s formatting is, the easier it will be to grep source code. I work a lot with Go, where all code in our repository is gofmt'ed. You can get quite far with regular expressions for finding/analyzing Go code. (And when regexps don’t cut it anymore, Go has excellent infrastructure for working with it programmatically. http://golang.org/s/types-tutorial http://golang.org/s/types-tutorial is a great introduction!)
- nkozyra 7y ago> the stricter your language’s formatting is, the easier it will be to grep source code Well that makes sense, the less reliable your input text, the more complex the regexp.
- felipeerias 7y agoRelated to this, it is generally a very good idea to be strict when naming functions, parameters, variables, etc. so that each concept has exactly one name throughout the codebase.
- polytomous_ 7y agoAlso, check your spelling. It's a pain when items don't show up in search because of spelling issues. I've seen function names misspelled, and then every invocation just doubling down on that misspelling.
- yakubin 7y agoI encountered a similar problem in a c++ codebase and the debugging logs it produced. There was an error which was being reported in the logs as "weak ptr expired" or somethings like that. I grepped for it the whole source code (a gigantic project). No results. Going back and forth several times. Feeling stupid beyond imagination. Then I copy-pasted what was actually printed in the logs into my grep query (previously I was typing it in manually). It quickly found a match. Turns out someone wrote "weak ptr" as "week ptr". Everyone in the team had a good laugh.
- jimmaswell 7y agoStill never going to beat AST-integrated searching like VS has for C#. Which has a regex search too.
- arethuza 7y agoAre there any stand-alone AST based search tools?
- bruce_one 7y agoI've found a few, eg https://www.graspjs.com/ https://www.graspjs.com/ is a Javascript one. I do wonder if there is a multi-language aware one though.
- hchasestevens 7y agoI'm going to plug my own here: astpath, which is AST-based search for Python. https://github.com/hchasestevens/astpath/ https://github.com/hchasestevens/astpath/
- secure 7y agoYes, https://github.com/mvdan/gogrep https://github.com/mvdan/gogrep for Go
- llimllib 7y agogoogling returned https://github.com/azz/ast-grep https://github.com/azz/ast-grep for js
- sethammons 7y agoI have stopped using the AST integrated searching vs code for Go. Instead of clicking to declaration, it is now faster in my larger code base (especially one that uses interfaces a lot) to just search for substrings. The AST search still works most of the time, but sometimes it fails and usually it is just plain slow.
- jpalomaki 7y agoRandom idea: maybe you could supercharge this by introducing to grep some constructs from programming languages. Now you have things like "word character", "whitespace", "start of line". In supercharged version you would have "function", "identifier"
- pjc50 7y agoThen it's not grep, it's something much slower which has to parse the syntax.
- LandR 7y agoVisual Studio with Resharper does this. It's very very fast / almost instant even with hundreds of source code files and millions of lines of code. I hit ctrl+T and then can search everything, this give me a drop down that filters out the more I type, select the item in the dropdown and it goes to that source file. I can also type: /t and search just types /m members /mm methods /u unit tests /f file /fp project /e event /mp property /mf field /ff project folder e.g. /t Foo will find all the Foos /mm SavePhoto will find any methods called SavePhoto Same works in JetBrains Rider for C# stuff. I couldn't dev without this now, and it's all built into my IDE.
- secure 7y agohttps://github.com/mvdan/gogrep https://github.com/mvdan/gogrep does something like this for Go :)
- predakanga 7y agoThere's an old project from Facebook that does something like this, called Pfff[0] It provides a tool called sgrep (syntactical grep) that lets you do some cool tricks, for instance: `sgrep "some_func(X,X)"` returns all calls to some_func with the same argument repeated. `sgrep "some_func(X,Y,...)"` returns all calls to some_func with 3 or more arguments. It's come in very handy for refactoring some troublesome codebases. [0]: https://github.com/facebookarchive/pfff https://github.com/facebookarchive/pfff
- anon1253 7y agoI tend to work a lot in Lisp and XML, both are more or less trees if you squint (with the Lisp syntax famously being the AST due to homoiconicity) and it always makes me wonder if there are better command line tree search or tree diff algorithms out there (extra awesome if it works with git merge strategies). I mean whitespace preference is fine and all, but sometimes you just don’t care :p
- parentheses 7y agoYou can search your codebase using livegrep [0] and get near instant results. [0] https://github.com/livegrep/livegrep https://github.com/livegrep/livegrep
- tannhaeuser 7y agoThe post's core message seems to be lost on HN. It's about screening sources for supposedly insecure and/or injection-prone funcs using simple text scanning (such as strcat, which however is considered in iOS apps when it is a C std API func); supposedly grepability is also about quickly finding code locations of messages and variables. But comments are all about Rust or Go superiority, irrelevant grep implementation details, and AST-based code analysis tools when these are specifically dismissed in TFA as producing too many false positives. Talk about bubbles and echo chambers.
- fredley 7y agoDon't use grep. Use ag[0], which is specifically designed for searching code. It's much faster, honors .gitignore, and the output can be piped back through grep if you like. ag FooBar | grep -v Baz It's in brew/apt/yum etc as `the_silver_searcher` (although brew install ag works fine too). 0: https://github.com/ggreer/the_silver_searcher https://github.com/ggreer/the_silver_searcher
- hawski 7y agoIt's not much faster as in "over 50% faster". It's faster to invoke as you have much less to type to scan recursively with ignore list. However I find that in many projects .gitignore is too extensive, because it includes generated code, that many times is quite informative. Then it's still nice to use those alternative grep-likes, but not by much. Besides, when you can't install easily it's hard to beat something that it's already there and everywhere else.
- tigroferoce 7y agoI second ag too. Not because it's faster, but because it is developer-oriented so it has sane defaults for searching into code.
- war1025 7y agoI find `git grep` to be quite sufficient. A real pain when trying to look through code that isn't in git though.
- Whoaa512 7y agoyou can just add `--no-index` to search in non-git repositories
- Joeboy 7y agoIs there a reason to use ag rather than rg? afair the latter was a lot faster when I tried it (on ubuntu / intel).
- timwaagh 7y ago> If not, the reviewer can quickly dismiss it as a false positive This is were you could be wrong. We would need to give a reason for dismissing it and then the risk officer would need to approve it (or reject it). False positives can be a real pain in the ass.
- jolmg 7y agoThe ability to easily grep for functions in C-like code is why I've come to appreciate projects defining their functions like: int foo_func(void) { You can grep for `^foo_func\b` to get to a declaration or definition, or `^foo_func\b.* {$` to get to a definition or `^foo_func\b.* ;` to get to a declaration. This is instead of using something like `^\w.* \bfoo_func\(`, which is what you'd need for: int foo_func(void) { By the way, anyone know of a way to insert a literal asterisk here without having to follow it up with a space?
- kahirsch 7y agoYes, I've been doing this since I saw it in the BSD source back in the '80s.
- paulddraper 7y agoJust one more reason to love languages with trailing instead of leading types (Scala, Typescript). fooFunc(): Int { "fooFunc returns an integer." Not "An integer is returned by fooFunc."
- switch007 7y agoThe imperative headline strikes again!
- alxmdev 7y agoHold on, strncat and strncpy are considered dangerous too, now? Not just the older versions without the size_t num argument?
- unilynx 7y agoThey don't \0-terminate the target on overflow, so you still need to test for that condition. So most people will have a wrapper around those to ensure the \0 is there.
- guitarbill 7y agoI think BSD has strlcpy and strlcat for exactly this reason
- jandrese 7y agostrlcpy has the braindamage that it returns the length of the source buffer, which means it has to traverse the entire buffer to figure out the length. If you want to copy out the first line from a buffer that happens to be a 10TB mapped file, that strlcpy call will take a long time to finish. If you are using strncpy/strlcpy because you don't trust the src buffer is properly null terminated but you still want to stop the copy at the first null or when the buffer is full, well, you're out of luck because strlcpy is going to blast past the end of the source buffer regardless. I would have been much happier if it had just returned a flag indicating either successful copy (0), buffer was truncated (1), or an error occurred and errno was set (-1). Possible errors could be that the src or dest was NULL or the size was 0 (ERR_BAD_ARGUMENT).
- deathanatos 7y agostrncpy() acts as you describe, but strncat() will terminate; from its man page[1], > the resulting string in dest is always null-terminated. [1]: https://linux.die.net/man/3/strncat https://linux.die.net/man/3/strncat
- deathanatos 7y agoIn addition to what unilynx mentions about strncpy(), the size arguments are also, effectively, the remaining space in the destination buffer, not the entire space in the destination buffer. So, you have to figure that out. It isn't hard (hell, it's trivial) but I think if you're either going to be aware of the pitfalls — and then these functions are mostly not going to help you — or you're not, in which case you're just as likely to pass the wrong value for the size (dest's size/src's size) and overflow the buffer anyways. Honestly, if I had to do more than a trivial amount of string manipulation in C, I'd be wrapping that in a mini library to manage some sort of stronger string type or finding such a library (glib? ICU?) very quickly, depending on needs. std::string was one of the things in C++ that made me question why anyone was still using C, given how much less error-prone it is, comparatively. (std::string is not without problems / only as compared to char * in C.)
- johnny-lee 7y agoI've gone down this road years ago. While there's no install and initial results are quick to appear, the false positives that grep or any string search tool generates will make the cynics shoot down this simple attempt to find problems in the source code. Problems that arose: - what about use of those questionable APIs/constants in strings (perhaps for logging) or in comments? - some of the APIs listed in the article were only questionable when certain values were used - sometimes you can get grep/search tool of choice to play along, but if the API call spans multiple lines or the constant has been assigned to a variable that is used instead, then a plain string search won't help. - it's hard to ignore previously flagged but accepted uses of the API/constants. - so there's a possible bug reported, but devs usually want to see the context of the problem (the code that contains the problem) quickly/easily. Some text editors can grok the grep output and place the cursor at the particular line/character with the problem, some can't. If you go down that road to try and reduce false positives, you'll end up with a parser for your development language of choice.
- time0ut 7y agoI haven't tried this approach, but having spent years using one of the best commercial SAST tools, I'm reluctant to dismiss it too quickly. My SAST generates tons of false positives and is unforgivably slow. If this is orders of magnitude faster, it might be worth the extra false positives. As a side note, my dream is a SAST that comments directly in the PR like a human reviewer would. Maybe that exists?
- johnny-lee 7y agoThe SAST program is probably doing a lot more than a string search tool does. If the SAST has to process C/C++ source code, then the SAST will parse all the #include'd header files. The SAST may track values to determine if illegal/uninitialized values are used. A string search tool will skip doing all of that. If the class of problems you're looking for contains only bad functions/constants, then a string search tool may be fine. But as I mentioned before, the string search tool may get confused if these bad strings occur in strings/comments/irrelevant #if/#else/#elif sections. There are another class of bugs dealing with data values which a string search tool can't deal with easily. As an example, PC-Lint lists the type of problems the program may flag - https://www.gimpel.com/html/lintchks.htm https://www.gimpel.com/html/lintchks.htm. A string search tool won't know about classes and virtual destructors or other concepts relevant to the programming language in question. For the string search tool, you'd either invoke the search string tool several times with different search strings for the same source code or slightly more efficient, have one long search string containing all your search strings as alternate search targets for the string search tool. Either case, when the string search tool spits out a positive result, it won't explain why there is a problem. The dev will have to know or lookup the problem associated with that search result. When I worked on this area, C/C++ compilers stopped at syntax errors. Most have gotten better at flagging popular problems like variable assignments within if statements, operator precedence bugs, and printf-format string bugs. Some divisions at Microsoft required devs to run a lightweight SAST before committing changes to locate possible problems ASAP. It's relatively easy to integrate an SAST into your build system to scan the modified source code just before you're ready to commit the changes.
- KuhlMensch 7y agoI do a few VERY SIMPLE greps. The most useful, is a pre-commit hook to check no blacklisted env vars exist in the commit diff. So, useful. Grepping leans-in to shell. Though if you have other environments available (python, javascript etc), it makes sense to lean-into them e.g I use JavaScript examine my package.json to ensure my dependency SemVers' are "exact". That said, I rarely write static-analysis scripts: In JavaScript-world there is already a plethora of easily configurable linting & type-checking tools. If I wanted to focus in on static-analysis etc I'd probably reach for https://danger.systems/js/ https://danger.systems/js/ SideNote: My CI generates a metrics.csv file, which serves as a "metric catch-all" for any script I might write e.g. grep to count "// TODO" and "test.skip" strings, plus my JavasScript tests generate performance metrics (via monkey-patching React). I don't actually DO ANYTHING with these metrics, but I'm quite happy knowing the CI is chugging away at its little metric diary. One day I'll plug it into something.