22 ms·
Semgrep: Semantic grep for code
- jhgb 5y agoIsn't "grep for code" called just "grep"?
- vinceguidry 5y agoTagline appears to be submitter's, not the project's.
- mynegation 5y agoievans works for r2c
- ievans 5y agoSemgrep started off as a “syntactic grep” but has increasingly become more semantic. So if you want to find all calls to foo that have 1 as the first argument, you just search for foo(1) and even things like x = 1; foo(x); will match. Here’s an elaborate example: https://semgrep.dev/s/ievans:c-dataflow https://semgrep.dev/s/ievans:c-dataflow
- dang 5y agoOk, I've put the word 'semantic' up there so we can not get hung up on title stuff. (Submitted title said "Like Grep but for Code".)
- jhgb 5y agoThat's called "a code walker", isn't it?
- tyingq 5y agoIt does seem potentially good for enforcing standards where the participants are willing. But you can work around it fairly easily. Like the example "python no-prints" rule: https://semgrep.dev/s/sabihb:no-prints https://semgrep.dev/s/sabihb:no-prints Lots of workarounds it wouldn't find, like: import builtins builtins.print("whee")
- sdesol 5y agoI'm guessing the value with the rules system is, you can add new rules easily. So during a code review, if you see somebody using your example, you could create a new rule to catch that. I don't think you can go in with the mindset that it will catch everything, but rather, it's about being able to iterate quickly with your rules.
- tyingq 5y agoSure. I raised it because it keeps using the word "static analysis", which at least wouldn't be fooled by builtins.print(). It's more than grep, for sure, but something less than static analysis tools I've used. And I don't mean that as a knock. It's working across a lot of languages, so I see the tradeoff.
- enriquto 5y ago> You need to enable JavaScript to run this app. Wait, is this a web app? I was expecting a command line tool to navigate my code locally.
- exdsq 5y agoThere’s a demo on the site, I assume that’s it?
- enriquto 5y agoNo, it's just the landing page. Apparently it does not allow to see it without running some javascript.
- sahkopoyta 5y agoWell the demo is there on the landing page
- thomasahle 5y agoThe cli is here: https://semgrep.dev/docs/getting-started/ https://semgrep.dev/docs/getting-started/ You can write stuff like # Check for Python == where the left and right hand sides are the same (often a bug) $ semgrep -e '$X == $X' --lang=py path/to/src
- enriquto 5y agoCool! But this example is a bit simplistic since it can be done just as easily by regular grep: grep -E '(.+) = \1' *.py I have trouble looking at the examples in the project website (many things inside iframes are adblocked). Do you have any example of a search that would be difficult or impossible with grep?
- SavantIdiot 5y agoIt can infer (x==y) if x=1 and y=1, which is grep cannot do.
- deleted 5y ago[deleted]
- SavantIdiot 5y agoSince the capability has never existed, I don't think in terms of being able to semgrep. If that makes any sense. My brain is not wired this way, yet. Like, if you've never tasted lychee, it would never occur to you how to cook with it. I'm going to need to see some useful, real-world examples to jumpstart my brain to think this way.
- underyx 5y agoHey, I work on Semgrep. As a real world example, I just noticed today that Hashicorp uses a whole bunch of Semgrep rules on terraform-provider-aws[0]. I'd recommend reading the `message` keys to know what they intend to match, and then the `patterns` lists below to see how that's accomplished. Alternatively, we curate 1000+ community rules that you can look through as well.[1] [0]: https://github.com/hashicorp/terraform-provider-aws/blob/main/.semgrep.yml https://github.com/hashicorp/terraform-provider-aws/blob/mai... [1]: https://semgrep.dev/r https://semgrep.dev/r
- SavantIdiot 5y agoNice! Thanks. This will certainly help me start to thinking in semantic grep. I can see this being an additional coverage tool and am eager to study it.
- saagarjha 5y agoYou can cook with lychee?
- kesterallen 5y agoTypo in the "Trying Semgrep" screenshot ("ruleste"): https://semgrep.dev/static/media/Step1.df848497.png https://semgrep.dev/static/media/Step1.df848497.png
- hyper_reality 5y agoThis is an excellent tool to have as a security consultant, and it just keeps getting better and better. When approaching a large codebase, it enables you to write custom rules that match on certain antipatterns you've spotted that may be unique to the codebase. That's the real value of the tool, but the repository of per-language rules is also convenient for quickly finding low-hanging fruit (like every use of a potentially injectable function such as exec,system,etc. in PHP). For example, a webapp may have been designed such that authorisation needs to be explicitly added with a line or two to each controller. A semgrep rule can be written to match all the controllers which are missing this line. Then these controllers can be manually reviewed to assess whether unauthorised access should be allowed. Depending on what you are trying to match, this is something that may be very complex or even impossible to implement accurately in plain grep. Some languages like Ruby have powerful static analysis tools (Brakeman) that can also do this, but the benefit of Semgrep is the flexibility across multiple languages and how readable the rulesets are. [1] [1] https://blog.includesecurity.com/2021/01/custom-static-analysis-rules-showdown-brakeman-vs-semgrep/ https://blog.includesecurity.com/2021/01/custom-static-analy...
- tyingq 5y agoI'd be careful with how much of a warm fuzzy the tool gives you. See this example from my other comment in the thread: https://news.ycombinator.com/item?id=26905880 https://news.ycombinator.com/item?id=26905880 If it were really looking at AST level data, that wouldn't have fooled it. I suspect there would be similar issues with your example of ensuring no use of eval() in PHP. So it seems okay to keep your own developers informed, but I wouldn't use it, alone, to vet outside code. PHP has eval-like functionality buried in preg_replace(), assert(), and probably other places. This tool also doesn't seem to dig into namespaced "aliases".
- aseipp 5y ago> If it were really looking at AST level data, that wouldn't have fooled it. Semgrep does look at an AST; but that counterexample is not something you can "fix" solely by looking at an AST. You need actual Python-specific semantic analysis that knows that all "open" functions like 'print' come from the builtins module, and thus are bound to the same identifier. They're literally built into the implementation, it's not something you can "discover" from analyzing existing Python source. Even if you had a perfectly accurate python AST it couldn't "tell" you this fact, it's a priori knowledge, and all analysis engines need a base set of facts like this that they work from. > but I wouldn't use it, alone, to vet outside code. I mean, nobody seems to be suggesting this though, and the OP quite literally stated the major value of the tool is enforcing domain/codebase-specific rules among a team. Which is a really good use for it! There are tons of little useful patterns you can codify this way.
- thesuperbigfrog 5y agoThe name "Semantic Grep" does not give a good idea for what this tool is and what it does. The web page states: "Static analysis at ludicrous speed. Find bugs and enforce code standards" "grep" is short for "global regular expression print". It finds matches for the given regular expression and prints them. "Semantic Grep" is a static analyzer with configurable rules, style checks, etc. It does much more than search and print. Perhaps a better name is needed? Edit: How about "omnilint" or "omnicritic" since semgrep is more of a "lint" (https://en.wikipedia.org/wiki/Lint_(software) https://en.wikipedia.org/wiki/Lint_(software)) or "critic" (https://en.wikipedia.org/wiki/Perl::Critic https://en.wikipedia.org/wiki/Perl::Critic) type of tool that handles multiple languages? Edit2: "Static analysis at ludicrous speed" ==> "turbolint"? ("ludicrous speed" reminds of the hilarious Space Balls scene :) "turbolint, GO!"
- petters 5y agoThe literal meaning of "grep" is not the only meaning. It also means “find snippet in files.”
- thesuperbigfrog 5y agoYes. But at least to me, semgrep looks a lot more like "lint" than "grep".
- prepend 5y agoYou have to find stuff before you lint it.
- smithza 5y agoThere is the common, if informal, definition that grep means "command line text search tool". I read this as "semantic/syntax search tool".
- thesuperbigfrog 5y agoBut semgrep is much closer to a linter / critic program than a grep program.
- rmetzler 5y agoLooks like a useful tool for me and I would like to try it. Go down, see "brew install semgrep" and try to copy paste it. And it's an image :(
- rmetzler 5y agoThere is also a bug in the example rules single pages app. Go to https://semgrep.dev/p/jwt https://semgrep.dev/p/jwt Go to the page 2/5 Click "Run Locally", so you can copy the code close the modal -> you're on page 1/5. Expectation would be to stay on page 2/5. It would also be very useful to be able to filter by language and topic.
- westurner 5y agoIs there a more complete example of how to call semgrep from pre-commit (which gets called before every git commit) in order to prevent e.g. Python print calls (print(), print \\n(), etc.) from being checked in? https://semgrep.dev/docs/extensions/ https://semgrep.dev/docs/extensions/ describes how to do pre-commit. Nvm, here's semgrep's own .pre-commit-config.yml for semgrep itself: https://github.com/returntocorp/semgrep/blob/develop/.pre-commit-config.yaml https://github.com/returntocorp/semgrep/blob/develop/.pre-co...
- theptip 5y agoI've never used the `pre-commit` framework, but it's really simple to wire up arbitrary shell scripts; check out the `.git/hooks` directory in your repo for samples, e.g. `.git/hooks/pre-commit.sample`. You can run any old shell script there, without having to install a python tool.
- westurner 5y agoYeah but that githook will only be installed on that one repo on that one machine. And they may have no or a different version of bash installed (on e.g. MacOS or Windows). IMHO, POSIX-compatible portable shell scripts are more trouble than portable Python scripts. Pre-commit requires Python and pre-commit to be installed (and then it downloads every hook function). This fetches the latest version of every hook defined in the .pre-commit-config.yml: pre-commit autoupdate https://pre-commit.com/#pre-commit-autoupdate https://pre-commit.com/#pre-commit-autoupdate A person could easily `ln -s repo/.hooks/hook*.sh repo/.git/hooks/` after every git clone.
- joj123 5y agoOut of curiosity, Is there value in doing this over (say) running a GitHub Action post commit and failing the build if it finds something nasty?
- Spivak 5y agoIf you can catch it before the commit is even made then why do/wait for a build?
- layer8 5y agoNo Windows support yet: https://github.com/returntocorp/semgrep/issues/1330 https://github.com/returntocorp/semgrep/issues/1330
- twh270 5y agoFrom the thread you link it looks like they're getting close, there's been activity in the past few days. (I'm guessing from your comment that this is important to you (i.e. WSL/Docker is not a solution).)
- IshKebab 5y agoKind of crazy that you can make a tool like this in the modern age that isn't cross platform. Maybe they just can't face the Python packaging nightmare on Windows. Also kind of surprising it's written in Python given that they advertise its speed.
- prepend 5y agoI’m always surprised at stuff I take for granted that doesn’t work on Windows. So yeah, it seems like cross-platform should be easy but since my dev environment is zsh, it’s easy for my stuff to work sort of “everywhere but Windows.” Add to that that the reason things fail on Windows is usually something Windows specific and “their fault.” So it’s unusual for me to fire up a Windows VM just to sanity check my code. And since Windows CI runners cost 2x or more, I don’t usually run cross platform CI.
- IshKebab 5y agoThis is an open source project on Github. They can use Github Actions which has free runners for Windows, Mac and Linux.
- prepend 5y agoWindows runners consume 2x minutes, Mac runners consume 10x minutes [0]. Free projects only get 2,000 minutes per month so running on all three platforms means you only get 1/13th of the minutes of only using Linux. I rarely use anything but Linux runners even on my paid projects. I like saving money, so unless I really need integration testing on Windows or Mac, I don’t do it. [0] https://docs.github.com/en/github/setting-up-and-managing-billing-and-payments-on-github/about-billing-for-github-actions https://docs.github.com/en/github/setting-up-and-managing-bi...
- silasb 5y agoJust the tool that I was looking for. We are looking to do Service linting in our organization as a method of making sure our services don't drift too far apart. Anyone else know of a Service linting tool? OPA/conftest come close but lack syntax parsers for Ruby/Javascript.
- eric_fib 5y agogrep grep
- hn_throwaway_99 5y agoI currently use a highly opinionated ESLint config (based on the airbnb one) together with strict checking in my TypeScript config, and it is configured to run on every commit with husky git hooks. The example given on the Semgrep homepage is an exact match to one that exists in my ESLint config (eslint's no-console rule). How does Semgrep compare to ESLint+a strict tsconfig?
- avodonosov 5y agoHow do you deal with false positives? If the commit hook rejects anything where rules are triggered, a way to force the comrit is needed for the cases when the rule finding is not reanny an issue. Upd: I found that the --no-verify option can be used in many cases
- hn_throwaway_99 5y agoYou can put comments in the code to ignore eslint rules on a specific line, or for the whole file.
- RichieAHB 5y agoIt seems from reading the docs that the “semantic” side is important here. It can track things like redefinitions. E.g. const output = console.log; output(“hello”); This wouldn’t be caught by ESLint (to my knowledge) but would be caught by Semgrep. I think you could do it with ESLint but given the interface for an ESLint plugin exposes an AST, you’d have to track this yourself. I’m assuming Semgrep could stretch to things like enforcing APIs are called with certain optional arguments present (even if the TS types don’t require it). Again, I think with ESLint you’d have to do more juggling with the raw AST. That said, this is my understanding from a quick skim!
- unwind 5y agoWhen tools like this use terms like "legacy languages", and don't show that C is supported unless you click "More Languages", it makes me feel old. :) Still, it seems rather cool, I like the idea of being able to search code at a higher level than just raw source text.
- afro88 5y agoNo swift support yet. What would be involved in adding it?
- mdaniel 5y agohttps://github.com/returntocorp/ocaml-tree-sitter/blob/master/doc/adding-a-language.md#how-to-add-support-for-a-new-language https://github.com/returntocorp/ocaml-tree-sitter/blob/maste... appears to be the general answer to your question, but navigating to the tree-sitter docs shows that tree-sitter has one in progress: https://github.com/tree-sitter/tree-sitter-swift https://github.com/tree-sitter/tree-sitter-swift so hopefully the machinery to incorporate it into semgrep will not be horrific
- vlovich123 5y agoI want the ease of use of their AST specification with the power of clang’s refactor tool. Has anyone attempted to do that?
- more_corn 5y agoI used to use SAST-SCAN but that seems abandonware. I like that this exists. Everyone should go from nothing to something in the SAST space. A free/freemium tool/service for that is pretty great. The first couple runs have found useful results.
- leafmeal 5y agoWhat does this give you over writing a flake8 plugin (for Python at least)? I've found the flake8 API and documentation lacking, so perhaps just a cleaner interface?
- pantuza 5y agoReally outstanding those guardrails rules from semgrep. Useful to enforce code. Thanks for sharing the tool.
- shuringai 5y agoThis is much better alternative to codeQL used by google and does not use a shameless registration-only model! Thanks for sharing
- shuringai 5y ago*github, not google my bad
- pabs3 5y agoDoes it come with a standard set of rules that finds bad code without any false positives out of the box? Or is it more of a tool for people doing code security audits & pentesting who know what they are looking for and want to read the surrounding code?
- realquadrant 5y agoHi, this is very cool. I have been building up a suite of tools to roll out across major open source projects to improve security. I like what I have seen so far, this is a great use case. Whom can I connect with to learn more? And similarity/diff with sourcegraph, also like a lot.
- minusf 5y agoprobably doing something wrong but running the ci ruleset on a tiny django hobby project made all cores spin at 100% after 33% of the progress bar and made the OS almost unresponsive. ctrl-c after 5 minutes and i still had to pkill every semgrep process... never seen the M1 airbook overheat this much before.
- drewdennison 5y agoSemgrep maintainer here. We just ordered our first M1 laptop and will debug. Thanks for the bug report
- wdb 5y agoApparently this is invalid TypeScript (cannot parse it says): try { const parsedURL = new URL(url) requestPath = parsedURL.pathname } catch (error: unknown) { // NOOP } It's complaining about : unknown bit which one of the newer typescript eslint rules enforces.
- Spivak 5y agoApparently import random if random.randint(0,1) == 2: print(“hello”) Is also unparseable.
- mdaniel 5y agois that due to the smartquotes, or that's just an artifact of your HN comment? Perhaps a more pointed set of questions: what is the error it emits, and have you considered submitting that case as an actual bug?
- CGamesPlay 5y agoHow much does the CI service cost? I can't seem to find any information about it on the website without creating an account.
- joj123 5y agoThe CI service is free, with some limitations on how long the findings stay on the dashboard, SSO integration and maybe a few others. The paid version was $40/usr/mo the last time I checked Once we figured it out, it takes us a few minutes to onboard a new repo to Semgrep
- joshuamorton 5y agoThere's lots of confusion about what semgrep does here, which is kind of unfortunate. I haven't touched it much, but I have built a very similar tool (I'm one of the contributors to refex[1], which is a very similar project). The starting point of semantic grep is very useful. When you have a big codebase, you often want to detect antipatterns, or not even antipatterns, but just uses of a thing, say you're renaming a method and want to track down the callers. Being able to act on the AST, instead of hoping you searched up all of the variants of whitespace and line breaks and, depending on the specific example, different uses of argument passing, is really useful. But often when you're semantically grepping, your goal is to replace something with something else (this is what refex was initially built for: to aide in large scale changes in python, as a sort of equivalent to the C++ tools that Google uses). But then you want to shift left even further: once you have a pattern that you want to replace once, you can just enforce that a linter yell at you when anyone does it again. So it's very natural to develop a linter-style thing on top of one of these[2]. This is, as I understand it sort of the same thing that happens in C++: clang-tidy and clang-format are written on top of AST libraries that can be used for ad-hoc analysis and transformations, but you can also just plug them into a linter. The thing is, for most organizations, enforcing code style and best practices is more valuable than apply a refactoring to 10M lines of code, because most organizations don't have 10M lines of code to refactor. That doesn't mean that these tools aren't also useful for ad-hoc transforms and exploratory analysis. They absolutely are! [1]: https://github.com/ssbr/refex https://github.com/ssbr/refex [2]: https://github.com/ssbr/refex/tree/main/refex/fix https://github.com/ssbr/refex/tree/main/refex/fix
- solipsism 5y agoWhat's the status of C++ support?
- carlmr 5y agoNone, so I assume it's not planned in the near future. Hasn't been there the last few times this was posted on HN. A really cool tool even if the language that would profit the most of it isn't supported.
- Annatar 5y agoI click on the link above and I get a seemingly blank page, all because the website uses some JavaScript garbage and violates W3C standards. That's the ridiculous, disgusting state of the information technology industry in the 21st century. I rue the day I decided to do this professionally, and I am deeply ashamed and despondent.
- nojvek 5y agoThe underlying package tree-sitter that semgrep uses is pretty amazing too. It’s an incremental parser for many different languages written in C. It blows my mind how fast it is compared to many tools in js ecosystem. Tree-sitter was parsing millions of files in half a minute. JS, TS, Ruby, yaml, html, Css. It’s quite magical. Such great engineering.
- globular-toast 5y agoWhenever I see "at ludicrous speed" or something to that effect, I now assume it's slow.
- vindarel 5y agoInteresting. Looks similar to Comby: https://comby.dev/ https://comby.dev/ "a tool for searching and changing code structure". Comby is more on rewriting, it has less integration for a CI (though you can do it), it is less geared towards reporting.
- hardon4semgrep 5y agoHow does this compare to the tools available at large companies like Google and Facebook?
- sriram_malhar 5y agoNice looking tool. Is there a way to search for functions in C (other than printf!) whose return value is ignored at the call site?