5 ms·
One thing which was not immediately obvious to me for a while: the stricter your language’s formatting is, the easier it will be to grep source code. I work a
by secure 7y ago
One thing which was not immediately obvious to me for a while: the stricter your language’s formatting is, the easier it will be to grep source code.
I work a lot with Go, where all code in our repository is gofmt'ed. You can get quite far with regular expressions for finding/analyzing Go code.
(And when regexps don’t cut it anymore, Go has excellent infrastructure for working with it programmatically. http://golang.org/s/types-tutorial http://golang.org/s/types-tutorial is a great introduction!)
- nkozyra 7y ago> the stricter your language’s formatting is, the easier it will be to grep source code Well that makes sense, the less reliable your input text, the more complex the regexp.
- felipeerias 7y agoRelated to this, it is generally a very good idea to be strict when naming functions, parameters, variables, etc. so that each concept has exactly one name throughout the codebase.
- polytomous_ 7y agoAlso, check your spelling. It's a pain when items don't show up in search because of spelling issues. I've seen function names misspelled, and then every invocation just doubling down on that misspelling.
- yakubin 7y agoI encountered a similar problem in a c++ codebase and the debugging logs it produced. There was an error which was being reported in the logs as "weak ptr expired" or somethings like that. I grepped for it the whole source code (a gigantic project). No results. Going back and forth several times. Feeling stupid beyond imagination. Then I copy-pasted what was actually printed in the logs into my grep query (previously I was typing it in manually). It quickly found a match. Turns out someone wrote "weak ptr" as "week ptr". Everyone in the team had a good laugh.
- jandrese 7y agoAnd it turns out the pointer expired on sunday morning, except in some regions where it expired on some other day.
- xamuel 7y agoAlso, please think twice before using OO classes as a license to give all your methods useless names like "get", "add", "open", etc.!
- splittingTimes 7y agoBut how do you effectively orhanize/enforce this for a code base of several million LOC where geographically distributed teams are working on different ends of the system all the time? The amount of cross team coordination is staggering.
- paulddraper 7y agoDocumentation, I think.
- peterstensmyr 7y agoFail verification in CI if the change doesn’t pass your checks. You can check anything, such as whether it has a duplicate name already in the codebase.
- splittingTimes 7y agoI would be interested to know, which CI tool can check "that each concept has exactly one name throughout the codebase." I thought code reviews are the only way and then you need to have every Dev aligned and on the same page on this topic... Which never happens. :/
- peterstensmyr 7y agoIt wouldn’t be something provided by the CI tool, you’d have to write the test yourself. At the end of the day it’s just another test, albeit a more complex one than a standard unit test.
- fragsworth 7y agoAnother thing that's not immediately obvious is the longer your search string is in grep, the faster it will find your results.
- herpderperator 7y agoCan you explain why that's the case?
- saghm 7y agoI think it's because there are fewer potential substrings to check for matches, since most of the characters you add to a regex to make it longer also add to the minimum length of the expressions that it can find
- jlebar 7y agoString search algorithms can cleverly skip forward when they don't find a match. They can skip forward more for longer "needle" strings. That said grep is pretty fast period, so probably doesn't make a huge difference in practice, especially if you're IO-bound, which is common.
- jandrese 7y agoThe caveat being if you're Unicode aware then many of the old skipahead strategies don't work as well and have to be rolled back or disabled.
- burntsushi 7y agoThat's not true. They work just fine. Typically, substring search algorithms are implemented at the encoding level, e.g., on UTF-8 directly. If you just treat that as an alphabet of size 256, then algorithms like Boyer-Moore work out of the box. But the skip-ahead stuff isn't the most important thing nowadays. The key is staying in the fast vectorized skip loop as long as possible.
- paulddraper 7y agoI find that line wrapping frequently prevents me from getting complete coverage though. So it works for some cases, not for others. But I do love auto-formatters. (I'm doing web dev at the moment, so Prettier.) It is freeing to not worry about spacing, line breaks, parens, etc. All I have to do is give the computer a valid AST and it does The Right Thing.
- jandrese 7y agoIsn't this what the /s flag is for in your regex? Assuming you are also using /x of course.
- paulddraper 7y agoSort of. But there's indentation, so I need repetition. And if I'm doing whitespace repetition, there's no great advantage in an autoformatter.
- jandrese 7y agoYou do have to be more liberal with '\s*' or '\s+' instead of just ' ' in your patterns, that is true.