14 ms·
Debugging: Indispensable rules for finding even the most elusive problems (2004)
- deleted 2y ago[deleted]
- mootoday 2y agoDid anyone say debugging? I've followed https://debugbetter.com/ https://debugbetter.com/ for a few weeks and the content has been great!
- deleted 2y ago[deleted]
- spawarotti 2y agoVery good online course on debugging: Software Debugging on Udacity by Andreas Zeller https://www.udacity.com/course/debugging--cs259 https://www.udacity.com/course/debugging--cs259
- belter 2y agoUdacity is owned by Accenture? That is...Surprising.
- apples_oranges 2y agoA good bug is the most fun thing about software development
- jamesblonde 2y agoJust let a LLM create even better bugs for you - Erik Meijer https://www.youtube.com/live/SsJqmV3Wtkg?si=MUoiNbWpsunsZ39y&t=1006 https://www.youtube.com/live/SsJqmV3Wtkg?si=MUoiNbWpsunsZ39y...
- waynecochran 2y agoSometimes I am actually happy when there is a obvious bug. It is like solving a murder mystery.
- BobbyTables2 2y agoAnd often you’re the culprit too!
- deleted 2y ago[deleted]
- D-Coder 2y agoThe victim, the murderer, and the detective.
- k3vinw 2y agoThe unspoken rule is talking to the rubber duck :)
- zarq 2y agoThat is literally #8 in the list
- dkdbejwi383 2y agoRule #10 - read everything twice
- k3vinw 2y agoHmm. Perhaps, but mannequin is not nearly as whimsical sounding as a rubber duck which inspires you to bounce your ideas off of the inanimate object.
- urbandw311er 2y agoRule #10 - it’s probably DNS
- tmountain 2y agoYears ago, my boss thought he was being clever and set our server’s DNS to the root nameservers. We kept getting sporadic timeouts on requests. That took a while to track down… I think I got a pizza out of the deal.
- urbandw311er 2y agoThat sounds hellish
- jsrcout 2y agoI worked at place in the late 90s where that was true, at least for anything Internet related. We were doing (oh so primitive by today's standards...) Web development and it happened so many times. I'd call downstairs and they'd swear DNS was fine, and then 20 minutes to half an hour later, it would all be mysteriously working again. But only if we called down heh heh. On an unrelated note, one of the folks down there explained the DNS setup once and it was like something out of a Stephen King novel. They'd even been told by a recognized industry expert (whose name I sadly can't remember any more) that what they needed to do was impossible, but they still did it. Somehow. They really were great folks, they just had that one quirk but after a while I could just chuckle about it.
- knlb 2y agoI wrote a fairly similar take on this a few years ago (without having read the original book mentioned here) -- https://explog.in/notes/debugging.html https://explog.in/notes/debugging.html Julia Evans also has a very nice zine on debugging: https://wizardzines.com/zines/debugging-guide/ https://wizardzines.com/zines/debugging-guide/
- Hackbraten 2y agoI love Julia Evans’ zine! Bought several copies when it came out, gave some to coworkers and donated one to our office library.
- david_draco 2y agoStep 10, add the bug as a test to the CI to prevent regressions? Make sure the CI fails before the fix and works after the fix.
- soco 2y agoWhat do you do with the years old bug fixes? How fast can one run the CI after a long while of accumulating tests? Do they still make sense to be kept in the long run?
- hobs 2y agoWhy would you want to stop knowing that your old bug fixes still worked in the context of your system? Saying "oh its been good for awhile now" has nothing to do with breaking it in the future.
- simmonmt 2y agoThis is a great problem to have, if (IME) rare. Step 1 Understand the System helps you figure out when tests can be eliminated as no longer relevant and/or which tests can be merged.
- hsbauauvhabzb 2y agoI think for some types of bugs a CI test would be valuable if the developer believes regressions may occur, for other bugs they would be useless.
- jerf 2y agoI'm not particularly passionate about arguing the exact details of "unit" versus "integration" testing, let alone breaking down the granularity beyond that as some do, but I am passionate that they need to be fast, and this is why. By that, I mean, it is a perfectly viable use of engineering time to make changes that deliberately make running the tests faster. A lot of slow tests are slow because nobody has even tried to speed them up. They just wrote something that worked, probably literally years ago, that does something horrible like fully build a docker container and fully initialize a complicated database and fully do an install of the system and starts processes for everything and liberally uses "sleep"-based concurrency control and so on and so forth, which was fine when you were doing that 5 times but becomes a problem when you're trying to run it hundreds of times, and that's a problem, because we really ought to be running it hundreds of thousands or millions of times. I would love to work on a project where we had so many well-optimized automated tests that despite their speed they were still a problem for building. I'm sure there's a few out there, but I doubt it's many.
- qwertox 2y agoMake sure you're editing the correct file on the correct machine.
- chupasaurus 2y agoPoor Yorick!
- spacebanana7 2y agoAlso that it's in the correct folder
- eddd-ddde 2y agoHow much time I've wasted unknowingly editing generated files, out of version files, forgetting to save, ... only god knows.
- ZedZark 2y agoYep, this is a variation of "check the plug" I find myself doing this all the time now I will temporarily add a line to cause a fatal error, to check that it's the right file (and, depending on the situation, also the right line)
- shmoogy 2y agoI'm glad I'm not the only one doing this after I wasted too much time trying to figure out why my docker build was not reflecting the changes ... never again..
- overhead4075 2y agoThis is also covered by "make it fail"
- ajuc 2y agoThat's why you make it break differently first. To see your changes have any effect.
- snowfarthing 2y ago
- waynecochran 2y agoI also think it is worthwhile stepping thru working code with a debugger. The actual control flow reveals what is actually happening and will tell you how to improve the code. It is also a great way to demystify how other's code runs.
- ajross 2y agoI think that fits nicely under rule 1 ("Understand the system"). The rules aren't about tools and methods, they're about core tasks and the reason behind them.
- nthingtohide 2y agoMake sure through pure logic that you have correctly identified the Root Cause. Don't fix other probable causes. This is very important.
- sumtechguy 2y agoThat is rule #3. quit thinking and look. Use whatever tool you need and look at what is going on. The next few rules (4-6) are what you need to do while you are doing step #3.
- nottorp 2y agoThis is necessary sometimes when you’re simply working on an existing code base.
- hugograffiti 2y agoI agree and have found using a time travel debugger very useful because you can go backwards and forwards to figure out exactly what the code is doing. I made a video of me using our debugger to compare two recordings - one where the program worked and one where a very intermittent bug occurred. This was in code I was completely unfamiliar with so would have been hard for me to figure out without this. The video is pretty rubbish to be honest - I could never work in sales - but if you skip the first few minutes it might give you a flavour of what you can do. (I basically started at the end - where it failed - and worked backwards comparing the good and bad recordings) https://www.youtube.com/watch?v=GyKrDvQ2DdI https://www.youtube.com/watch?v=GyKrDvQ2DdI
- berikv 2y agoPersonally, I’d start with divide and conquer. If you’re working on a relevant code base chances are that you can’t learn all the API spec and documentation because it’s just too much.
- berikv 2y agoAlso: Fix every bug twice: Both the implementation and the “call site” — if at all possible
- BobbyTables2 2y agoYe ol’ “belt and suspenders” approach?
- causal 2y agoCheck the plug should be first
- begueradj 2y agoThis is related to the classic debugging book with the same title. I first discovered it here in HN.
- fn-mote 2y agoThe article is a 2024 "review" (really more of a very brief summary) of a 2002 book about debugging. The list is fun for us to look at because it is so familiar. The enticement to read the book is the stories it contains. Plus the hope that it will make our juniors more capable of handling complex situations that require meticulous care... The discussion on the article looks nice but the submitted title breaks the HN rule about numbering (IMO). It's a catchy take on the post anyway. I doubt I would have looked at a more mundane title.
- bananapub 2y ago> The article is a 2024 "review" 2004.
- TheLockranore 2y agoRule 11: If you haven't solved it and reach this rule, one of your assertions is incorrect. Start over.
- duxup 2y agoI’m so bad at #1. I know it is the best route, I do know the system (maybe I wrote it) and yet time and again I don’t take the time to read what I should… and I make assumptions in hopes of speeding up the process/ fix, and I cost myself time…
- nickjj 2y agoFor #4 (divide and conquer), I've found `git bisect` helps a lot. If you have a known good commit and one of dozens or hundreds of commits after that is bad, this can help you identify the bad commit / code in a few steps. Here's a walk through on using it: https://nickjanetakis.com/blog/using-git-bisect-to-help-find-which-commit-broke-something https://nickjanetakis.com/blog/using-git-bisect-to-help-find... I jumped into a pretty big unknown code base in a live consulting call and we found the problem pretty quickly using this method. Without that, the scope of where things could be broken was too big given the context (unfamiliar code base, multiple people working on it, only able to chat with 1 developer on the project, etc.).
- Icathian 2y agoTacking on my own article about git bisect run. It really is an amazing little tool. https://andrewrepp.com/git_bisect_run https://andrewrepp.com/git_bisect_run
- jerf 2y ago"git bisect" is why I maintain the discipline that all commits to the "real" branch, however you define that term, should all individually build and pass all (known-at-the-time) tests and generally be deployable in the sense that they would "work" to the best of your knowledge, even if you do not actually want to deploy that literal release. I use this as my #1 principle, above "I should be able to see every keystroke ever written" or "I want every last 'Fixes.' commit" that is sometimes advocated for here, because those principles make bisect useless. The thing is, I don't even bisect that often... the discipline necessary to maintain that in your source code heavily overlaps with the disciplines to prevent code regression and bugs in the first place, but when I do finally use it, it can pay for itself in literally one shot once a year, because we get bisect out for the biggest, most mysterious bugs, the ones that I know from experience can involve devs staring at code for potentially weeks, and while I'm yet to have a bisect that points at a one-line commit, I've definitely had it hand us multiple-day's-worth of clue in one shot. If I was maintaining that discipline just for bisect we might quibble with the cost/benefits, but since there's a lot of other reasons to maintain that discipline anyhow, it's a big win for those sorts of disciplines.
- deleted 2y ago[deleted]
- heikkilevanto 2y agoSome additional rules: - "It is your own fault". Always suspect your code changes before anything else. It can be a compiler bug or even a hardware error, but those are very rare. - "When you find a bug, go back hunt down its family and friends". Think where else the same kind of thing could have happened, and check those. - "Optimize for the user first, the maintenance programmer second, and last if at all for the computer".
- ajuc 2y agoIt's healthier to assume your code is wrong than otherwise. But it's best to simply bisect the cause-effect chain a few more times and be sure.
- wormlord 2y agoI always have the mindset of "its my fault". My Linux workstation constantly crashing because of the i9-13900k in it was honestly humiliating. Was very relieved when I realized it was the CPU and not some impossible to find code error.
- dehrmann 2y agoLinux is like an abusive relationship in that way--it's always your fault.
- physicles 2y agoThe first one is known in the Pragmatic Programmer as “select isn’t broken.” Summarized at https://blog.codinghorror.com/the-first-rule-of-programming-its-always-your-fault/amp/ https://blog.codinghorror.com/the-first-rule-of-programming-...
- bsammon 2y agoAlternatively, I've found the "Maybe it's a bug. I'll try an make a test case I can report on the mailing list" approach useful at times. Usually, in the process of reducing my error-generating code down to a simpler case, I find the bug in my logic. I've been fortunate that heisenbugs have been rare. Once or twice, I have ended up with something to report to the devs. Generally, those were libraries (probably from sourceforge/github) with only a few hundred or less users that did not get a lot of testing.
- ChrisArchitect 2y ago(2004) Title is: David A. Wheeler's Review of Debugging by David J. Agans
- Tepix 2y agoMy first rule for debugging debutants: Don't be too embarassed to scatter debug logmessages in the code. It helps. My second rule: Don't forget to remove them when you're done.
- lanstin 2y agoMy rule for a long time has been anytime I add a print or log, except for the first time I am writing some new cide with tricky logic, which I try not to do, never delete it. Lower it to the lowest possible debug or trace level but if it was useful once it will be useful again, even if only to document the flow thru the code on full debug. The nicest log package I had would always count the number of times a log msg was hit even if the debug level meant nothing happened. The C preprocessor made this easy, haven't been able to get a short way to do this counting efficiently in other languages.
- sitkack 2y agoI really like this.
- condour75 2y agoOne good timesaver: debug in the easiest environment that you can reproduce the bug in. For instance, if it’s an issue with a website on an iPad, first see if you reproduce in chrome using the responsive tools in web developer. If that doesn’t work, see if it reproduces in desktop safari. Then the iPad simulator, and only then the real hardware. Saves a lot of frustration and time, and each step towards the actual hardware eliminates a whole category of bugs.
- ChrisMarshallNY 2y ago#7 Check the plug: Question your assumptions, start at the beginning, and test the tool. I have found that 90% of network problems, are bad cables. That's not an exaggeration. Most IT folks I know, throw out ethernet cables immediately. They don't bother testing them. They just toss 'em in the trash, and break a new one out of the package.
- nickcw 2y agoI prefer to cut the connectors off with savage vengeance before tossing the faulty cable ;-)
- teleforce 2y agoThe tenth golden rule: 10) Enable frame pointers [1]. [1] The return of the frame pointers: https://news.ycombinator.com/item?id=39731824 https://news.ycombinator.com/item?id=39731824
- PhunkyPhil 2y agoI would almost change 4 into "Binary search". Wheeler gets close to it by suggesting to locate which side of the bug you're on, but often I find myself doing this recursively until I locate it.
- ajuc 2y agoYeah people say use git bisect but that's another dimension (which change introduced the bug). Bisecting is just as useful when searching for the layer of application which has the bug (including external libraries, OS, hardware, etc.) or data ranges that trigger the bug. There's just no handy tools like git bisect for that. So this amounts to writing down what you tested and removing the possibilities that you excluded with each test.
- analog31 2y agoOne I learned on Friday: Check your solder connections under a microscope before hacking the firmware.
- InitialLastName 2y agoThe worst is when it works only when your oscilloscope probe is pushing down on the pin.
- GuB-42 2y agoRule 0: Don't panic Really, that's important. You need to think clearly, deadlines and angry customers are a distraction. That's also when having a good manager who can trust you is important, his job is to shield you from all that so that you can devote all of your attention to solving the problem.
- Cerium 2y agoSlow is smooth and smooth is fast. If you don't have time to do it right, what makes you think there is time to do it twice?
- adamc 2y agoI had a boss who used to say that her job was to be a crap umbrella, so that the engineers under her could focus on their actual jobs.
- dazzawazza 2y agoIdeally it's crap umbrellas all the way down. Everyone should be shielding everyone below them from the crap slithering its way down.
- generic92034 2y agoThat must be the trickle-down effect everyone talking about in the 80ies. ;)
- saghm 2y ago
- BlueUmarell 2y agoPost: "9 rules of debugging" Each comment: "..and this is my 10th rule: <insert witty rule>" Total number of rules when reaching the end of the post: 9 + n + n * m, with n being number of users commenting, m being the number of users not posting but still mentally commenting on the other users' comments.
- deleted 2y ago[deleted]
- fedeb95 2y agorule -1: don't trust the bug issuer
- nextlevelwizard 2y agobad rule for bad programmers. trust, but verify.
- fedeb95 2y agoyes, that was the point.
- reverendsteveii 2y agoReview was good enough to make me snag the entire book. I'm taking a break from algorithmic content for a bit and this will help. Besides, I've got an OOM bug at work and it will be fun to formalize the steps of troubleshooting it. Thanks, OP!
- sumtechguy 2y agoI recommend this book to all Jr. devs. Many feel very overwhelmed by the process. Putting it into nice interesting stories and how to be methodical is a good lesson for everyone.
- deleted 2y ago[deleted]
- omkar-foss 2y agoFor folks who love to read books, here's an excerpt from the Debugging book's accompanying website (https://debuggingrules.com/ https://debuggingrules.com/): "Dave was asked as the author of Debugging to create a list of 5 books he would recommend to fans, and came up with this. https://shepherd.com/best-books/to-give-engineers-new-perspectives https://shepherd.com/best-books/to-give-engineers-new-perspe..."
- burrish 2y agothanks for the link
- omkar-foss 2y agoMost welcome :)
- shahzaibmushtaq 2y agoI can't comment further on David A. Wheeler's review because his words were from 2004 (He said everything true), and I can't comment on the book either because I haven't read it yet. Thank you for introducing me to this book. One of my favorite rules of debugging is to read the code in plain language. If the words don't make sense somewhere, you have found the problem or part of it.
- BWStearns 2y ago> Check the plug I just spent a whole day trying to figure out what was going on with a radio. Turns out I had tx/rx swapped. When I went to check tx/rx alignment I misread the documentation in the same way as the first. So, I would even add "try switching things anyways" to the list. If you have solid (but wrong) reasoning for why you did something then you won't see the error later even if it's right in front of you.
- SoftTalker 2y agoYes the human brain can really be blind when its a priori assumptions turn out to be wrong.
- jgrahamc 2y agoWasn't Bryan Cantrill writing a book about debugging? I'd love to read that.
- bcantrill 2y agoI was! (Along with co-author Dave Pacheco.) And I still have the dream that we'll finish it one day: we had written probably a third of it, but then life intervened in various dimensions. And indeed, as part of our preparation to write our book (which we titled The Joy of Debugging), we read Wheeler's Debugging. On the one hand, I think it's great to have anything written about debugging, as it's a subject that has not been treated with the weight that it deserves. But on the other, the "methodology" here is really more of a collection of aphorisms; if folks find it helpful, great -- but I came away from Debugging thinking that the canonical book on debugging has yet to be written. Fortunately, my efforts with Dave weren't for naught: as part of testing our own ideas on the subject, I gave a series of presentations from ~2015 to ~2017 that described our thinking. A talk that pulls many of these together is my GOTO Chicago talk in 2017, on debugging production systems.[0] That talk doesn't incorporate all of our thinking, but I think it gets to a lot of it -- and I do think it stands at a bit of contrast to Wheeler's work. [0] https://www.youtube.com/watch?v=30jNsCVLpAE https://www.youtube.com/watch?v=30jNsCVLpAE
- gregthelaw 2y agoIt's a great talk! I have stolen your "if you smell smoke, find the source" advice and put it in some of my own talks on the subject.
- Zolomon 2y agoI have been bitten more than once thinking that my initial assumption was correct, diving deeper and deeper - only to realize I had to ascend and look outside of the rabbit hole to find the actual issue. > Assumption is the mother of all screwups.
- sitkack 2y agoThis is how I view debugging, aligning my mental model with how the system actually works. Assumptions are bugs in the mental model. The problem is conflating what is knowledge with what is an assumption.
- astrobe_ 2y agoI've once heard from an RTS game caster (IIRC it was Day9 about Starcraft) "Assuming... Is killing you".
- kmoser 2y agoWhen playing Captain Obvious (i.e. the human rubber duck) with other devs, every time they state something to be true my response is, "prove it!" It's amazing how quickly you find bugs when somebody else is making you question your assumptions.
- ianmcgowan 2y agoI used to manage a team that supported an online banking platform and gave a copy of this book to each new team member. If nothing else, it helped create a shared vocabulary. It's useful to get the poster and make sure everyone knows the rules. https://debuggingrules.com/download-the-poster/ https://debuggingrules.com/download-the-poster/
- nottorp 2y agoI’d add “a logging module done today will save you a lot of overtime next year”.
- goshx 2y ago> Quit thinking and look (get data first, don't just do complicated repairs based on guessing) From my experience, this is the single most important part of the process. Once you keep in mind that nothing paranormal ever happens in systems and everything has an explanation, it is your job to find the reason for things, not guess them. I tell my team: just put your brain aside and start following the flow of events checking the data and eventually you will find where things mismatch.
- pbalau 2y agoAcquire a rubber duck. Teach the duck how the system works.
- throwawayfks 2y agoI worked at a place once where the process was "Quit thinking, and have a meeting where everyone speculates about what it might be." "Everyone" included all the nontechnical staff to whom the computer might as well be magic, and all the engineers who were sitting there guessing and as a consequence not at a keyboard looking. I don't miss working there.
- drivers99 2y agoThere's a book I love and always talk about called "Stop Guessing: The 9 Behaviors of Great Problem Solvers" by Nat Greene. It's coincidental, I guess, that they both have 9 steps. Some of the steps are similar so I think the two books would be complementary, so I'm going to check out "Debugging" as well.
- __MatrixMan__ 2y ago> Check that it's really fixed, check that it's really your fix that fixed it, know that it never just goes away by itself I wish this were true, and maybe it was in 2004, but when you've got noise coming in from the cloud provider and noise coming in from all of your vendors I think it's actually quite likely that you'll see a failure once and never again. I know I've fixed things for people without without asking if they ever noticed it was broken, and I'm sure people are doing that to me also.
- jwpapi 2y agoI’m not sure that doesn’t sit well with me. Rule 1 should be: Reproduce with most minimal setup. 99% you’ll already have found the bug. 1% for me was a font that couldn’t do a combination of letters in a row. life ft, just didn’t work and thats why it made mistakes in the PDF. No way I could’ve ever known that if I wouldn’t have reproduced it down to the letter. Just split code in half till you find what’s the exact part that goes wrong.
- physicles 2y agoRelated: decrease your iterating time as much as possible. If you can test your fix in 30 seconds vs 5 minutes, you’ll fix it in hours instead of days.
- 101011 2y agoRule 4 is divide and conquer, which is the 'splitting code in half' you reference. I'd argue that you can't effectively split something in half unless you first understand the system. The book itself really is wonderful - the author is quite approachable and anything but dogmatic.
- gnufx 2y agoThen, after successful debugging your job isn't finished. The outline of "Three Questions About Each Bug You Find" <http://www.multicians.org/thvv/threeq.html http://www.multicians.org/thvv/threeq.html> is: 1. Is this mistake somewhere else also? 2. What next bug is hidden behind this one? 3. What should I do to prevent bugs like this?
- sandbar 2y agoTake the time to speed up my iteration cycles has always been incredibly valuable. It can be really painful because its not directly contributing to determining/fixing the bug (which could be exacerbated if there is external pressure), but its always been worth it. Of course, this only applies to instances where it takes ~4+ minutes to run a single 'experiment' (test, startup etc). I find when I do just try to push through with long running tests I'll often forget the exact variable I tweaked during the course of the run. Further, these tweaks can be very nuanced and require you to maintain a lot of the larger system in your head.
- astrobe_ 2y agoAlso sometimes: the bug is not in the code, its in the data. A few times I looked for a bug like "something is not happening when it should" or "This is not the expected result", when the issue was with some config file, database records, or thing sent by a server. For instance, particularly nasty are non-printable characters in text files that you don't see when you open the file. "simulate the failure" is sometimes useful, actually. Ask yourself "how would I implement this behavior", maybe even do it. Also: never reason on the absence of a specific log line. The logs can be wrong (bugged) too, sometimes. If you printf-debugging a problem around a conditional for instance, log both branches.
- hughdbrown 2y agoIn my experience, the most pernicious temptation is to take the buggy, non-working code you have now and to try to modify it with "fixes" until the code works. In my experience, you often cannot get broken code to become working code because there are too many possible changes to make. In my view, it is much easier to break working code than it is to fix broken code. Suppose you have a complete chain of N Christmas lights and they do not work when turned on. The temptation is to go through all the lights and to substitute in a single working light until you identify the non-working light. But suppose there are multiple non-working lights? You'll never find the error with this approach. Instead, you need to start with the minimal working approach -- possibly just a single light (if your Christmas lights work that way), adding more lights until you hit an error. In fact, the best case is if you have a broken string of lights and a similar but working string of lights! Then you can easily swap a test bulb out of the broken string and into the working chain until you find all the bad bulbs in the broken string. Starting with a minimal working example is the best way to fix a bug I have found. And you will find you resist this because you believe that you are close and it is too time-consuming to start from scratch. In practice, it tends to be a real time-saver, not the opposite.
- randerson 2y agoThe quickest solution, assuming learning from the problem isn't the priority, might be to replace the entire chain of lights without testing any of them. I've been part of some elusive production issues where eventually 1-2 team members attempted a rewrite of the offending routine while everyone else debugged it, and the rewrite "won" and shipped to production before we found the bug. Heresy I know. In at least one case we never found the bug, because we could only dedicate a finite amount of time to a "fixed" issue.
- hughdbrown 2y ago> The quickest solution, assuming learning from the problem isn't the priority, might be to replace the entire chain of lights without testing any of them. So as a metaphor for software debugging, this is "throw away the code, buy a working solution from somewhere else." It may be a way to run a business, but it does not explain how to debug software.
- __mharrison__ 2y agoGo on a walk or take a shower...
- kazinator 2y agoI've had trouble keeping the audit trail. It can distract from the flow of debugging, and there can be lots of details to it, many of which end up being irrelevant; i.e. all the blind rabbit holes that were not on the maze path to the bug. Unless you're a consultant who needs to account for the hours, or a teller of engaging debugging war stories, the red herrings and blind alleys are not that useful later.
- sitkack 2y agoIf folks want to instill this mindset in their kids, themselves or others I would recommend at least The Martian by Andy Weir https://en.wikipedia.org/wiki/The_Martian_(Weir_novel) https://en.wikipedia.org/wiki/The_Martian_(Weir_novel) https://en.wikipedia.org/wiki/Zen_and_the_Art_of_Motorcycle_Maintenance https://en.wikipedia.org/wiki/Zen_and_the_Art_of_Motorcycle_... https://en.wikipedia.org/wiki/The_Three-Body_Problem_(novel) https://en.wikipedia.org/wiki/The_Three-Body_Problem_(novel) To Engineer Is Human - The Role of Failure in Successful Design By Henry Petroski https://pressbooks.bccampus.ca/engineeringinsociety/front-matter/henry-petroski-and-to-engineer-is-human-the-role-of-failure-in-successful-design/ https://pressbooks.bccampus.ca/engineeringinsociety/front-ma... https://en.wikipedia.org/wiki/Surely_You%27re_Joking,_Mr._Feynman https://en.wikipedia.org/wiki/Surely_You%27re_Joking,_Mr._Fe...!
- deepspace 2y agoAgree. I think Zen and the Art of Motorcycle Maintenance encapsulates the art of troubleshooting the best. Especially the concept of "gumption traps", "What you have to do, if you get caught in this gumption trap of value rigidity, is slow down...you're going to have to slow down anyway whether you want to or not...but slow down deliberately and go over ground that you've been over before to see if the things you thought were important were really important and to -- well -- just stare at the machine. There's nothing wrong with that. Just live with it for a while. Watch it the way you watch a line when fishing and before long, as sure as you live, you'll get a little nibble, a little fact asking in a timid, humble way if you're interested in it. That's the way the world keeps on happening. Be interested in it." Words to live by
- sitkack 2y agoYeah, it is probably what I would start with and all the messages the book is sending you will resurface in the others. You have to cultivate that debugging mind, but once it starts to grow, it can't be stopped.
- fud101 2y agoi'm slogging through zen, it's a bit trite so far (opening pages). im struggling to continue. when will it stop talking about the climate and blackbirds and start saying something interesting?
- nox101 2y ago> #1 Understand the system: Read the manual, read everything in depth, know the fundamentals, know the road map, understand your tools, and look up the details. Maybe I'm mis-understand but "Read the manual, read everything in depth" sounds like. Oh, I have bug in my code, first read the entire manual of the library I'm using, all 700 pages, then read 7 books on the library details, now that a month or two has passed, go look at the bug. I'd be curious if there's a single programmer that follows this advice.
- deleted 2y ago[deleted]
- scudsworth 2y agogreat strawman, guy that refuses to read documentation
- mox1 2y agoI mean he has a point. Things are incredibly complex now adays, I don't think most people have time to "understand the system." I would be much more interested in rules that don't start with that... Like "Rules for debugging when you don't have the capacity to fully understand every part of the system." Bisecting is a great example here. If you are Bisecting, by definition you don't fully understand the system (or you would know which change caused the problem!)
- adolph 2y agoThis was written in 2004, the year of Google's IPO. Atwood and Spolsky didn't found Stack Overflow until 2008. [0] People knew things as the "Camel book" [1] and generally just knew things. 0. https://stackoverflow.blog/2021/12/14/podcast-400-an-oral-history-of-stack-overflow-told-by-its-founding-team/ https://stackoverflow.blog/2021/12/14/podcast-400-an-oral-hi... 1. https://www.perl.com/article/extracting-the-list-of-o-reilly-animals/ https://www.perl.com/article/extracting-the-list-of-o-reilly...
- feoren 2y agoEssentially yes, that's correct. Your mistake is thinking that the outcome of those months of work is being able to kinda-probably fix one single bug. No: the point of all that effort is to truly fix all the bugs of that kind (or as close to "all" as is feasible), and to stop writing them in the first place. The alternative is paradropping into an unknown system with a weird bug, messing randomly with things you don't understand until the tests turn green, and then submitting a PR and hoping you didn't just make everything even worse. It's never really knowing whether your system actually works or not. While I understand that is sometimes how it goes, doing that regularly is my nightmare. P.S. if the manual of a library you're using is 700 pages, you're probably using the wrong library.
- gregthelaw 2y agoI love the "if you didn't fix it, it ain't fixed". It's too easy to convince yourself something is fixed when you haven't fully root-caused it. If you don't understand exactly how the thing your seeing manifested, papering over the cracks will only cause more pain later on. As someone who has been working on a debugging tool (https://undo.io https://undo.io) for close to two decades now, I totally agree that it's just weird how little attention debugging as a whole gets. I'm somewhat encouraged to see this topic staying near the top of hacker news for as long as it has.
- bch 2y ago> If you didn't fix it, it ain't fixed AKA: “Problems that go away by themselves come back by themselves.”
- andypi_swfc 2y agoI found this book so helpful I created a worksheet based on it - might be helpful for some: https://andypi.co.uk/2024/01/26/concise-guide-to-debugging-anything-cheat-sheet/ https://andypi.co.uk/2024/01/26/concise-guide-to-debugging-a...
- pcblues 2y agoOver twenty five odd years, I have found the path to a general debugging prowess can best be achieved by doing it. I'd recommend taking the list/buying the book, using https://up-for-grabs.net https://up-for-grabs.net to find bugs on github/bugzilla, etc. and doing the following: 1. set up the dev environment 2. fork/clone the code 3. create a new branch to make changes and tests 4. use the list to try to find the root cause 5. create a pull request if you think you have fixed the bug And use Rule 0 from GuB-42: Don't panic (edited for line breaks)
- samsquire 2y agoOne thing I have been doing is to create a directory called "debug" from the software and write lots of different files when the main program has executed to add debugging information but only write files outside of hot loops for debugging and then visually inspect the logs when the program is exited. For intermediate representations this is better than printf to stdout
- worldhistory 2y agoGreat book +1
- _madmax_ 2y agoI had the incredible luck to stumble upon this book early in my career and it helped me tremendously in so many ways. If I could name only one it would be that it helped me get over the sentiment of being helpless in front of a difficult situation. This book brought me to peace with imperfection and me being an artisan of imperfection.
- khana 2y ago[dead]
- jagged-chisel 2y ago> Ask for fresh insights (just explaining the problem to a mannequin may help!) You can’t trust a thing this person says if they’re not recommending a duck.
- dalton_zk 2y agoFirst time hearing about these 9 rules, but I learning most of them by experience with many years resolving or trying to resolved bugs. Only thing that I dont agree is the book cost US$ 4.291,04 on Amazon
- dalton_zk 2y agobtw the hardcover its this price
- coldtea 2y ago>Rule 1: Understand the system: Read the manual, read everything in depth (...) Yeah, ain't nobody got time for that. If e.g. debugging a compile issue meant we read the compiler manual, we'd get nothing done...
- manhnt 2y ago> Make it fail: Do it again, start at the beginning, stimulate the failure, don't simulate the failure, find the uncontrolled condition that makes it intermittent, record everything and find the signature of intermittent bugs Unfortunately, I found many times this is actually the most difficult step. I've lost count of how many times our QA reported an intermittent bug in their env, only to never be able to reproduce it again in the lab. Until it hits 1 or 2 customer in the field, but then when we try to take a look at customer's env, it's gone and we don't know when it could come back again.
- fasten 2y agoNice classic that sticks to timeless pricniples. the nine rules are practical with war stories that make them stick. but agree that "don't panic" should be added
- 01308106991 2y agoHalo
- pbertrand429 2y ago[dead]