6 ms·
Formatting a 25M-line codebase overnight
- exsol 5mo ago[dead]
- andrewstuart 5mo ago[flagged]
- skinfaxi 5mo agoWhy is that terrifying?
- Jtsummers 5mo agoIt's not particularly terrifying. Some people really just don't like Ruby.
- mikedelago 5mo agoSome folks don't like shipping
- fantasizr 5mo agoive yet to see a compelling elitist programming language opinion. especially when used at big successful companies. these companies don't function in spite of their technology choices.
- lstodd 5mo ago> these companies don't function in spite of their technology choices. shows you never worked at "big succesful companies".
- NetOpWibby 5mo agoThe only one that worked on me wasn't even elitist in its framing. Try TypeScript! It makes your JavaScript better! That was enough for me.
- sikozu 5mo agoThe systems have to be written in some kind of programming language, and I think Ruby is a perfectly fine choice.
- Imustaskforhelp 5mo agoNot denying that Ruby is a perfectly fine choice but within the article itself it says that Stripe runs the world's largest Ruby codebase so certainly it might be testing the constraints of the language. The thing I am interested is that I don't suppose that Stripe always had these many LOC's and so I would be curious to know if at any point as the codebase was increasing, were they looking at other new languages which were coming like golang or rust which was more suited for their work or not and what were there decisions/thinking process to continue using ruby.
- clintonb 5mo agoLOC doesn’t have much to do with the “constraints of the language”. Stripe has dabbled in Golang. There is also a growing Java monorepo.
- throwaway041207 5mo agoStripe uses Sorbet which, in my experience, increases LOC.
- mbStavola 5mo agoConsidering that it's been doing so successfully at volume for just over 15 years, I think their language choice was fine.
- benbristow 5mo ago[dead]
- semiquaver 5mo agoI’d hardly call Sorbet Ruby :)
- sixo 5mo agoThis ought to change your mind about Ruby!
- sunrunner 5mo agoThings can always be worse. It could be PHP, for example.
- msla 5mo agoIf you think that's terrifying, imagine all of the essential code written in COBOL and FORTRAN. Skippy the Intern, now retired these thirty years...
- varun_ch 5mo agoI’m shocked at the 25M line part! That is a completely unfathomable amount of code for one codebase. I really want to know more about that.
- jsnell 5mo agoRight, where is the rest of the code?
- mr_mitm 5mo agoThey're up to 42 million now, as per the article
- lukan 5mo agoThat sounds even more insane to me, but I guess most of that code does not really touch financial transactions, otherwise it would be a nightmare being responsible to verify that.
- clintonb 5mo agoRuby code touches financial transactions. Card payments were migrated to Java when I left in 2022. Non-card payments (e.g., ACH, checks, various wallets) were still processed by Ruby. PCI-related/vaulting code lived in its own locked-down repo. I think that was a mix of Go and Ruby. Once you have the foundations in place for account balances and the ledger, processing a payment isn’t that daunting. Those foundations, however, took a lot to build and evolve.
- varun_ch 5mo ago> migrated to Java I want to know more about this
- jamesfinlayson 5mo ago> Once you have the foundations in place for account balances and the ledger, processing a payment isn’t that daunting. Those foundations, however, took a lot to build and evolve. Pretty much. I've worked at places with PHP payment processing that worked just fine, and at a place with C++ payment processing (and no testers) and it worked just fine. I wasn't around when the systems were first built though so not sure if there were tears along the way.
- hokkos 5mo agoNow it makes me wonder, are those 45M LoC are untyped ?
- m12k 5mo agohttps://brandur.org/nanoglyphs/015-ruby-typing#ruby-typing https://brandur.org/nanoglyphs/015-ruby-typing#ruby-typing
- c3ab8ff137 5mo agoNo, Stripe has its own Ruby typechecker - https://sorbet.org/ https://sorbet.org/
- burnte 5mo agoThe floating spiral thing is so distracting I spent more time deleting it in Inspector than reading the article. I feel like they hate their readers. Awful.
- annaspies 5mo agoIf you set `prefers-reduced-motion: reduce`, it goes away
- CrzyLngPwd 5mo agoOne of my first jobs was a small software company writing software for a small number of clients, in MS basic PDS. The lead developer didn't like to bother with formatting code, so I wrote a tool called makenice to format his nasty spaghetti gibberish into something with good indents and layout to make it easier for us normal people to parse. He was furious, literally spun in circles about it right in the office in front of everyone, so I wrote makenasty to format code into the way he appeared to like. I only shared makenasty/nice with a couple of the team, who loved it, as it allowed easy conversion between something readable and something the team lead like. He never knew about makenasty.
- munk-a 5mo agoOutside of the naming - this is a perfectly sane thing to do for developer comfort and can usually be accomplished with simple transformations. There are often limitations (like manually added indentation/spacing for alignment) but as long as you're very intentional about what changes you'll allow and have a good understanding of the language it can be an extremely safe operation.
- e28eta 5mo agoI think git’s naming is actually pretty reasonable: smudge (on checkout) & clean (on stage).
- munk-a 5mo agoOh smudge and clean are excellent names. My singly held objection to the OP was that they called one of the scripts "makenasty" instead of like "makemunkastyle" or something more neutral. I think it's an excellent idea I'd just avoid being judgemental in naming. You can consider my deep love of BSD braces super nasty but I'd prefer you didn't label it that way.
- nitwit005 5mo agoIf he didn't bother formatting code, it would seem impossible to create a tool that formatted code the way he preferred.
- CrzyLngPwd 5mo agoSurely, it no longer needs to be human-readable, and the era of write-only code is finally upon us with the dawn of AI writing our mealtickets. Why bother formatting 25m lines of slop, and why is AI wasting tokens on making code look human-readable anyway?
- sgc 5mo agoEvery LLM I have ever asked about this says they perform better when they receive pretty-printed code because it is easier to see structure and priorities. It has been an almost universal recommendation for me, and it makes sense since LLMs are just mimicking human expression.
- throawayonthe 5mo agoyou asked the llm? i'm confused you do understand it can't "know" how it performs right?
- sgc 5mo agoYou actually think that LLMs are not fed docs on how they work in order to help users interact with them better? Asking an LLM how to use it is based on the reasonable presumption that the company making it will prioritize making it useful for users and work on programming it with its own best practices. Again, it makes perfect sense as well based on how they are trained in the first place. Look at how they tokenize whitespace and you will see why it's useful. Each number of repeating white spaces gets a unique token (so 2 whitespaces = token1, 3 whitespaces = token2) - so it actually does make a very clear reinforcing hierarchy readily available. And we all know if there is anything an LLM needs, it is reinforcement of important points.
- CrzyLngPwd 5mo agoThat doesn't make sense. A compiler doesn't need pretty code to compile; in my tests, when I ask an LLM to deobfuscate code, it doesn't skip a beat.
- hobofan 5mo agoI'm surprised they went with a all-at-once reformat. Even when doing it over a weekend this is bound to mess with a lot of open PRs at their scale. I had to introduce a formatter in a few sizeable codebases in the past (few 100k to few million LOC), and I always did it incrementally via a script that reformatted all files that are not touched in any open PR. The initial run reformatted 95% of all files. Then I ran the script every day for ~two weeks and got up to 99.5% of all files and then manually each time one of the remaining ~dozen PRs that were WIP for longer were merged.
- skydhash 5mo agoYou can always let the team know so that they can apply the formatter on their PR branch.
- jrajav 5mo agoThis is exactly the remedy to the PR issue. I've "lucked" into owning a Prettier formatting pass at two different places now, and did the same process at each - full pass on master, simple step-by-step process to follow to update any PR by running the format script.
- hobofan 5mo agoIn the smaller migrations I did I tried that, but some way or another a decent chunk of the people still managed to get stuck in merge/rebase conflicts. I would almost explicitly not recommend giving that advise to the teams. My rough blueprint for introducing formatter or linter nowadys would be: - Recorded knowledge share session around how to set up the tools for local use 1-2 weeks before the initial rollout, and outline how the process will take place - On the day of the initial rollout send out a reminder + the recording again - Do the initial PR - Incrementally do the rest of the migration, and subscribe to the PRs that drag out the process
- rileymichael 5mo agoboth options have their pros and cons. if you utilize some form of ratcheting[1], you can sneak it in without your team knowing.. but all of your PRs for the foreseeable future will have a ton of reformatting screwing with your git blame. if you do it all at once, someone will have to sort out conflicts, but you can utilize `blame.ignoreRevsFile`[2] so that your history remains useful [1] https://github.com/diffplug/spotless/tree/main/plugin-gradle#ratchet https://github.com/diffplug/spotless/tree/main/plugin-gradle... [2] https://git-scm.com/docs/git-blame#Documentation/git-blame.txt---ignore-revs-filefile https://git-scm.com/docs/git-blame#Documentation/git-blame.t...
- munificent 5mo ago> We chose a Saturday to format the entire codebase to avoid merge conflicts. And while our test suite gave us high confidence we'd gotten everything right, it's always a bit daunting to have a diff so large that GitHub can't render it. The dart formatter has an internal sanity check. It walks through the unformatted and formatted strings in parallel skipping any whitespace. If any non-whitespace characters don't match, it immediately aborts. This ensures that the only thing the formatter changes is whitespace, and makes it much less spooky to run it blind on a huge codebase. That sanity check has saved my ass a couple of times when weird bugs crept in, usually around unusual combinations of language features around new syntax. (Unfortunately, the formatter in the past year has gotten a little more flexible about the kinds of changes it makes, including sometimes moving comments relatively to commas and brackets, so this sanity check skips some punctuation characters too, making it a little less reliable.)
- Terr_ 5mo agoI imagine a fancier version would be to compare the Abstract Syntax Trees.
- caminanteblanco 5mo agoThe only issue is then you're at the mercy of whatever parser your formatter uses to construct the AST
- Terr_ 5mo agoWell, if any (common, non-hobby) parser is thrown off by the reformatting, then it's probably not a safe reformatting either way.
- saghm 5mo agoI've always thought it would make sense for formatters to be baked into the toolchain so that they can reuse the language's parser (presumably exposed as a library) and then be implemented via parsing to AST and then formatted back out so that they're guaranteed to be correct and normalized. This doesn't seem to be how most formatters work in practice though, although I'm not sure if it's because of performance reasons or a lack of support for the parser being exposed in language toolchains.
- cadamsdotcom 5mo agoAn insight about code is that compared to the scale we operate on data, code as text is tiny. Instantaneous git operations and “run this tool over all the code” are the norm even while we wait for LLMs to stream their tokens to stream back so tool calls can operate on it. That insight might seem obvious - but if you stay cognizant of it as you work, you can invent some pretty amazing tooling for yourself & your team.
- nitwit005 5mo ago> Given that complexity, the hypothesis was simple: tackle the hardest syntax first and the rest will follow. Always nice to see. I've seen people fall into the trap of designing for the common case, not realizing most of the code will be to deal with the less common cases.
- sgc 5mo agoIn another field I have heard it called going for the jugular; the vivid description helps get the point across nicely. If you want to master something, you will have to know the hardest part. So just deal with that first and then everything else is easy, because you are dealing with it as somebody who has already mastered the domain.
- comrade1234 5mo agoMan must me nice to have the time to put so much work into tabs.
- Pxtl 5mo agoClean indenting is about saving time so you don't spend way too long getting lost trying to understand what seems like an insane piece of code until you realize it was a mundane bug hidden by incoherent indentation.
- machiaweliczny 5mo agoThe practical case is less time spend on rebasing/formatting code - IMO formatting standard is very helpful as then it's the same as storing AST. There's also better preserving of git blame and that's likely why they have done it as single operation as otherwise you would have everybody messing with that part and now you know there's single commit that touched everything and if blame is on it then you check previous edit.
- failure_arch 5mo ago[dead]
- tmaly 5mo agoHow did I know this was going to be a rewrite in Rust?
- throwatdem12311 5mo agoWhat is even the point of formatting code anymore.
- throawayonthe 5mo agoclean diffs for one
- voidUpdate 5mo agoTo make it look nice and readable
- throwatdem12311 5mo agoFor an agent?
- voidUpdate 5mo agoFor you, the person reading and writing the code
- throwatdem12311 5mo agoPeople are still reading the code?
- voidUpdate 5mo agoI do... I guess people not reading their own code is why products are so buggy and crap these days
- Twirrim 5mo agoWhy are you not reading the code?
- stefantalpalaru 5mo ago[dead]
- dgrin91 5mo agoI don't understand why the felt the need to do a big-bang merge like this. Its a formatter, so the files should be functionally equivalent before and after. Why not just enable it for new files/edit files for a while, then once comfortable apply it to old files in batches? What advantage does the big bang merge give? Seems higher risk for the same reward
- fsckboy 5mo agocould introduce subtle bugs, so doing it all at once while it's on the front of everybody's mind with as much comprehensive review and testing of parts or the whole to everybody's satisfaction. if you don't do it all at once, you'd need to repeat the same amount of testing multiple times. >files should be functionally equivalent before and after when you say something like this, the road you are on is paved with good intentions.
- riffraff 5mo agoYou can also introduce subtle bugs in your own feature development, and if you change formatting there it's also at the front of your mind. I think the main argument for doing a big bang rewrite is that you have a defined before/after, otherwise you're stuck into an endless in-between.
- eigenblake 5mo agoReally reminds me that there's nothing in principle stopping us from storing parse trees and exposing them via something git like so we can avoid even needing to format, let alone also needing to resolve a whole category of merge conflicts based on that formatting. I mean a format is just a theme over your data -- I mean code.
- ryanisnan 5mo agoCool story. The treat at the end was fun as well, thank you!
- hiroto_lemon 5mo ago[dead]