11 ms·
Losing faith in testing
- nickd2001 3y agoThere are good and bad tests.... ;) Good = actually tests something, precisely, properly, doesn't take an age to run. Enough of these in CI, and a deploy is far less likely to break prod. Bad = author didn't really know what they were trying to test, mocks fall out of date with reality (if they ever reflected it in the first place), other problems such as repetitive code and anything which tempts people to short circuit rather than keeping tests up to date, e:g tests that take far longer to write than the code they test. Shout out for pytest-cases as a library that can reduce repetition and save time. I've been in this game long enough to remember the days before massive automated test suites. Quality software did indeed get shipped in those days too. It took longer, with less regular releases. It took an army of skilled QA people with thorough test plans. You can do good or bad testing whether its today's TDD or yesterday's manual QA.
- ivanjermakov 3y agoI believe that the feeling when tests should be added to asset a certain feature comes only with experience. You just what logic is not trivial enough and where is makes sense to test it.
- tracerbulletx 3y agoAgreed, writing a core system component that isn't going to change and be called from thousands of different places? By all means heavily test and fuzz and cover the entire API surface. Writing some product feature that is probably going to be changed 1000 times and is at the edge of the system, waste of time. You'd be better off just having really good metrics and alerting for system degradation that measure actual business metrics, staggered canary roll outs, and easy rollbacks.
- gustavpaul 3y agoWrite the tests you need to sleep at night and avoid you and others you care about burning out when everything starts falling apart and even the smallest changes inspire dread.
- hvis 3y agoSpeaking of text editors and tools like that, you can often avoid having tests (or postpone adding them for a long time), if the logic is on the main execution path, meaning you'll execute it every time you run the program, and whatever failures that can happen, are reasonably easy to pinpoint (i.e. the program shows error backtraces or somehow traces problems otherwise). This is from my experience hacking on Emacs, naturally. At the same time, projects that you might ship for an employer or a client, are more critical to check for correctness before deploying, and are often more complex to run and check manually on the regular than writing at least one "happy path" integration test at least for the main scenario (or several).
- jonahx 3y ago1. There is no substitute for simplification and good design. You cannot test yourself out of a mess. 2. What you actually want is confidence in the code, and your ability to make changes. Sometimes tests are a good tool to achieve that, sometimes they are not. 3. The previous two points can and will be used to justify sloppy code by bad programmers, but it doesn't make them less true.
- rsyring 3y ago> You cannot test yourself out of a mess. My dev shop has taken over two dumpster fire web app projects, both many years old, passing through many unskilled/inexperienced hands, atrocious architecture and implementation, and no tests. But, actively being used and a full rewrite not being in the cards for some time. The very first thing we did for both was start writing integration tests at the http (wsgi client) layer to cover the majority of client facing functionality. We learned a lot about the apps in the process, fixed bugs where we could, documented and fixed lots of security issues, but generally didn't refactor anything that wasn't very broke. Once that was done, we could start refactoring which, for one project, included a migration off Mongo to tabled Postgres. The refactoring included significant unit and other testing so that we had very good and helpful test coverage when we were finished. I do believe there are times when testing yourself out of a mess is the only reasonable option.
- jonahx 3y agoFair enough, and I have had similar experiences. I would note a few things: - Sounds like a major undertaking. - You note this was done in service of refactoring. So you didn't test yourself out of the mess so much as use testing to enable the simplification that got you out of the mess. - I would argue that what is really happening here is that, by spending the time to create these tests and refactor, the current team is creating a shared mental model of the messy codebase -- they are making it understandable, at least in large part. So, you might amend my original statement "there is not substitute for simplification and good design" to "there is not substitute for making the code comprehensible". While a simple, good design is the best way to do this... it can also be achieved by having everyone expend more effort to understand a bad design well enough to work with it. And this latter strategy is, in practice, the far more common one.
- valcron1000 3y ago> In both codebases I’ve merged PRs without any tests and frequently see others do the same. I would never accept such PR. This kind of policies then end up biting you in the long run (experience talking)
- wilkystyle 3y agoI'm not going to comment on unit tests (of which we have many), as that seems to be where the biggest divide is, but I will say that our integration tests are worth their weight in gold.
- samatman 3y agoEditorializing the title to something grammatically-incorrect? Please don't. I consider it significant that both of the examples he cited are interactive programs. Thorough testing is low-payoff for those, for two reasons: it's common to tweak behaviors a bit, and if something is broken, you'll notice in the process of using it. Not no tests, but fewer tests, makes sense. At the opposite end, I'm working on a VM, and you better believe it's got tests. Not enough, it needs more, it always needs more, but they're a godsend. When I add an optimization to the compiler, I want confidence that it hasn't broken other behaviors, which frequently it does at first. They've enabled several refactors, with at least one more big one on the roadmap. Change one end of the pipeline, change the middle, change the end: half the tests fail, figure out why, the tests are green, all is well. It's almost a tautology but: write tests for a reason. Good reasons change over the lifecycle of a program. Early on, fixing simple invariants and preventing regression is a good motive, but with the recognition that things are going to change. Dogmatic TDD can lock in a design too early, if literally everything has a test right from the beginning the effort of every change is multiplied. For a mature systems program designed to be robust and load-bearing, complete coverage could be a good goal. For something like a text editor or terminal emulator, that's probably overkill outside of the core components. Tests aren't free, but no tests can get pretty expensive too.
- vegetablepotpie 3y agoThere’s a flip side. A culture that thinks end-to-end testing is the only legitimate way to test a system and that TDD is an unnecessary expense that is neither necessary nor sufficient for success. That culture is right, but it’s also wrong [1]. Although it is true that you can spend most of your time testing, providing little value, it’s equally true that with any software system, that by developing it, you will break it in subtle ways; ways in which you could spend weeks fixing bugs you introduced and fixed before, therefore creating little value. Unfortunately there is no substitute for critical thinking when engineering. This piece says we need to practice critical thinking with tests and that having a test that moves a mouse and clicks to confirm functionality, with suites that take 40 minutes to run, is going too far. Ball cites two examples of projects that strike a balance in testing speed and coverage. But how do we achieve this? I used “culture” in a specific way to describe working environments. Culture, not process, politics, or incentives, drives how you do testing. A culture is an environment that has preferences. A culture that will generate good tests is a culture that values technical rigor. Rigor is important for forging tests, but it’s also important for removing tests. To make these good test suites, we need to be comfortable with removing tests, but more importantly understand why a test is needed and remove it when we can’t explain it. [1] I’m not going to say there’s a “balance” to be struck, because that’s the language of people who say we should not write any tests.
- deleted 3y ago[deleted]
- userbinator 3y agoAn interesting correlation that I have observed over the years is that the more one is "religious" about writing tests, the less actual understanding of the code one seems to have. "Beware of bugs in the above code; I have only proved it correct, not tried it." - Donald Knuth
- BurningFrog 3y agoIf you have good tests, you don't need to know the code as well.
- eddd-ddde 3y agoThe more I program the more I feel this way. Do I really care how the code looks like or does something? Not really, all I know is this set of specifications are held true as I make changes, as long as I'm happy with those specs, I'm happy with the code.
- Jach 3y agoOn the small I wholly agree, like how much energy do some teams or even companies still waste on style document types of disagreements? But other aspects, I do often care how the code "looks" or more precisely "does" something. Like, it shouldn't be needlessly wasteful of resources, but I don't want to pigeonhole myself as "the performance guy". It shouldn't be using under-educated idioms that increase the likelihood that someone's gotta come back to this later to fix something stupid like a null pointer exception. At the same time certain things that sometimes get derided as ivory tower complex (or just "clever") code constructions, I don't think should necessarily be avoided all the time, but should be tastefully balanced with an aim towards broader understanding and not showing off cleverness for the sake of cleverness. I've replaced so many hundreds of lines of code with some relatively simple tens of lines type theory constructions in plain old Java 8, just because a dev who stopped learning around Java 1.4/5 doesn't understand doesn't make them "complex" or "clever". I don't even particularly like static typing, but a tool is a tool. But often I just don't care about even those details, and sometimes feel guilty about it. Carmack says to fill your products with give-a-damn, but I'm sorry, for so many things, a lot of the time I just don't/didn't. Work must be done anyway though, and not just by me. So practices that are broader, like enforcing tests, do help deliver a good enough product even under the guidance of devs and management and management's management etc. who seem to care even less than I do and don't even feel bad about it. It's particularly crazy when you talk to a customer who is gushing about something you know could have been even better with slightly different prioritization and tradeoffs; an important lesson is that many people are stoked merely that something exists. It's sad when tests catch something that really should have been caught, if not during code authoring time by an author who actually cares a bit more than just doing the job however and going home, then by code review time, but in large companies you sometimes have to just accept things for long periods and at least with more tests we can be more likely to catch things at all. (Not to mention they're pretty valuable when the original author who best understood the code is long gone, they help you make minor tweaks without having to tradeoff new feature development time with time getting a deep enough understanding of that old code that still mostly works most of the time.)
- peteforde 3y agoI'm shocked by how many developers check in code that passes the tests but they have not actually tested to make sure it works. Also, I'm tired of people not factoring in (and not honestly reporting) the time and frustration spent yak shaving to keep testing infrastructures working. I believe that it's because folks convince themselves that to acknowledge time lost to the ritual preparation for testing is a kind of weakness because other people aren't having those problems because surely you'd hear more about it. In reality, you can only write tests to cover the cases you anticipate. Correlating test coverage with reliability can be deadly; instead of losing sleep, periodically make sure that you can restore your backups to a production state and maybe even run some drills to see how your team responds when an unanticipated problem arises.
- Gigachad 3y agoTests are unbelievably useful for updating libraries. Every time I update rails I see a ton of specs fail all over the app highlighting breaking changes not mentioned in the docs. Stuff that is impossible to anticipate otherwise.
- bschwindHN 3y agoThis is where a nice type system and compilation checks in CI are very useful to have.
- Gigachad 3y agoI agree a type system does wipe out 80% of the tests you need, but I still feel like the 20% is useful. You can just write tests that verify the output looks right without having to run every single line of code to make sure there are no typos or type issues.
- reissbaker 3y agoYeah, I think that's one of the driving forces away from high code coverage / dogmatic TDD: dynamically typed languages used to be a lot more popular, and you really do need incredible amounts of testing to keep those working reliably in the long run. Now that typed languages are more popular — while some of the tricky logic parts are still worth testing — you can wipe out a lot of the rest because it simply won't compile when, say, you passed a nullable variable to a function that didn't expect it.
- norir 3y agoI have maintained a widely used (250-500k+ unique ip dl/month consistently over the last 6 years) interactive terminal program that had essentially no tests of its interactive behavior. Building and restarting the program took a minimum 15 seconds no matter how trivial the change. It became an absolute nightmare to work with and I eventually stopped contributing because it was so frustrating to work with. Writing good automated tests for interactive programs has to be done from the beginning or it will be almost impossible to effectively add later. To test an interactive program, you need to be able to simulate input and analyze the output stream. This can be done efficiently and effectively but again only if the program is designed this way from the beginning. At some point, if they continue on this trajectory, I would expect zed will end up in a similar state to the program I described above. It will become extremely difficult to reason about the effects of ui changes and very painful to debug and troubleshoot. Regressions will happen because relatively common cases (say 1-5% of users) that the developers themselves don't regularly use will not be noticed. Bug fixing becomes a game of whack a mole as manual changes to fix one case break another without your noticing. I hope they're able to avoid this fate because it was hellish for me.
- jonathankoren 3y ago>Building and restarting the program took a minimum 15 seconds no matter how trivial the change. It became an absolute nightmare to work with and I eventually stopped contributing because it was so frustrating to work with. Wait. You’re complaining about a 15 second compilation and startup loop? Don’t take this the wrong way, but I don’t think compiled languages are for you.
- preommr 3y agoThat's very unhelpful. I abhor long compile times and I exclusively use staticly typed, compiled/transpiled languages. The solution isn't to just shrug it off, but seriously evaluate how difficult it would be to refactor to smaller modules and if the benefits would be worth it. Sometimes, it's not worth the hassle. But if it's getting to a point where velocity is a concern, and a potential major version is on the horizon, a refactor to reduce code debt, and make things modular can be a real possibility if explained correctly to the right stakeholders.
- crdrost 3y agoRich Hickey once said something in the midst of my “TypeScript, Haskell, la la la la la” phase: > I like to ask this question, what's true of every bug ever found in the field? (It got written?) Pff. It got written, yes. What's a more interesting fact? (pause) It passed the type checker! What else did it do? (the tests?) It passed all the tests. So now what do you do? The basic problem is that we really want tests to somehow specify the contract of the software, but when you are writing “given, when, then” the “givens” pin down too much of how it is done and the “when/then” is scoped to distinct sub-contracts of different parts so it pins down how the work was broken-down... It hits like a syllogism, right? All tests are software, all software has scope creep, scope creep in testing is contract creep.
- fyrn_ 3y agoBugs all passed the tests sounds like survior bias. Like how all the fighter planes that returned with damage had damage to the wings. Point is, damage anywhere else was fatal, the takeaway should not have been to armor the wings more, which was the first thing which was tried. Likewise thinking about bugs this way discounts potential bugs that _were_ stopped by tests.
- WesolyKubeczek 3y agoIt's not like the planes that have returned. A bug reported in the field could well have been a crash. A costly, all-hands-on-deck incident kind of crash. And that bug still passed your own tests, compiled well, went into the production. Occam's razor tells me that it's more indicative of test harness being not comprehensive enough. Fix it to reproduce the bug/crash, then fix the code, then ship the fix, then tentatively pat yourself on the back for potential bugs that were stopped by the test. Until the next crash.
- fyrn_ 3y agoOh, I agree with that. I thought you were arguing that tests were less valuable because all production bugs passed the tests.
- andrewl 3y agoTests are valuable, but of course they’re never perfect. Certain kinds of programs require them more than others. And no development approach, method, or tool is the best in all cases. The program with the most tests that I’m aware of is SQLite. I have a feeling Hipp and his team know their code extremely well and that they do a lot of thinking before they change anything. Then they run their massive test suite, which they built because they’re good enough to know that they’re not good enough to think everything through without making errors. Tests have saved me from a few blunders over the years.
- wavemode 3y ago> Enter Ghostty and Zed. You're taking two desktop applications, written in Zig and Rust respectively, and extrapolating their lack of testing rigor as evidence that "there’s no correlation between software quality and tests"? Really? Have you tried maintaining quality in enterprise software without tests? If you ever manage to do so, I'd be very interested to read about it.
- misternugget 3y agoHey, author here. Yeah, I worked at Sourcegraph and I do think we built some high-quality stuff and we did write tests. I also think a lot of them were necessary. But, like I wrote here, I think that maybe we/I sometimes overdid it with tests and I'm not so sure about the use of /some/ of them anymore.
- osigurdson 3y ago>> Maybe the tests are only a symptom. A symptom of something else that causes the quality. Love this comment!
- gavmor 3y ago> Both are among the highest-quality software I have ever used and hacked on. > Both have less tests than I expected. Interesting bit of context: Zed's Nathan Sobo and Max Brunsfeld are both alumni of Pivotal Labs, a firm _notorious_ for its zealous adherence to, among other things, TDD. I don't, therefore, find it surprising that "neither codebase has tests, for example, that take a long-ass time to run," because the more stringently one test-drives, the less patiently one tolerates slow test suites. Besides that, after a serious investment in test-driving, one starts to learn which sorts of tests have paltry or even negative ROI; "tests that click through the UI and screenshot and compare" were right there at the top of the chopping block for most Pivotal devs, and tests that "hit the network" were explicitly taboo! I think that Kent Beck quote is great, but it's good advice for people who test too much, rather than devs in general or, god forbid, junior developers! The way I think about TDD, it's just like how after a while, one gets tired of copy+pasting code into the terminal, and reluctantly writes it to a file. One gets tired of emailing files, and checks them into version control. One gets tired of manipulating state--of clicking through the UI, of newing up a bunch of collaborators in the REPL, of smashing tab while blanking on the name of the method one has only just written--and writes a test. And maybe one who wakes up tired writes the test first. ;)
- seer 3y agoI dunno, in my career I’ve found that UI e2e tests were the ones that actually found bugs. All the unit tests that I’ve written ware mostly ceremony, to increase code coverage, etc, and were the first to need refactoring after some code change. A lot of refactoring. The tests that stayed true were the UI tests that almost always pointed to real problems with real code. It took a while to figure out how to write them though - trying to rely as little as possible on the internal ids / html / css and write them with what the user sees - e.g. instead of clicking on the button with “testid=login” we would “click on the button with the text login in it”. Identifying fields by the labels to them or tables by their column headers. It was inspired by rails’ capibara testing lib. And making sure the tests were not flaky, fast to execute and isolated took some time. But it was so worth the investment - it’s surprising how little those tests would change, as it allowed us to fearlessly refactor stuff without changing any tests. They felt a lot more like a friendly QA helping out rather than an annoyance you had to deal with after you’ve finished writing your code. And writing them was actually fun, since you didn’t have to understand how the app worked, fiddle with brittle css identifiers etc, you just wrote the steps you thought the user should do and saved it into a file. Being UI tests kinda meant they tested the whole system with the various micro services involved, databases and other infra. And I think this is where most problems in software arise, at the edges of systems when they try to interact with each other. Static types, immutability automatic API schemas and validators usually make sure the code one writes executes reasonably well, its where the code one writes starts interacting with all the other systems where people usually can’t anticipate things. And thats where the integration / e2e / UI tests help the most I think.
- shcheklein 3y agoThere was an interesting discussion recently on (somewhat) opposite perspective https://antithesis.com/blog/is_something_bugging_you/ https://antithesis.com/blog/is_something_bugging_you/ - how good testing systems makes everything way faster
- osigurdson 3y agoI don't know. When working in an unfamiliar codebase, tests are very helpful (even when some percentage of them will always be annoying). The problem is, the industry has focused on having lots of tests instead of having good code.
- TylerE 3y agoOne problem I don't see talked about enough - to the point where I don't even have a name for it - goes something like this, You have a giant, long lived system, like decade plus. You've got huge number of both unit and integration tests. You're cruising along, implementing a new feature. Everything's looking good from your local manual tests and the tests you wrote over the new behavior. Then you ship the branch up to the test runner to let the full suite cook, and you get it back, and there's like 9 random test fails that aren't in master. Has your change had unintended consequences? Is the test just out of date? If the test is out of date, how do we update it such that it both passes and we are at least reasonably confident that the test is still actually testing what it was intended to test? Where it gets real obnoxious is when the test isn't testing anything even vaguely related to your change so it can turn into a real snipe hunt. Five minutes here, ten minutes there, it adds up. What's frustrating is that the number of times this has found an actual issue that I can personally remember.. I won't swear it's zero, but I'm struggling to recall a specific example.
- osigurdson 3y agoI guess it all depends on the value of developer time vs the cost of releasing with an undiscovered bug. In any case, I'm glad the industry is starting to realize that tests are code, code can have bugs and code can be stupid.
- TylerE 3y agoWhat’s kinda funny is that there have been a few where we’ve had old code tests written sort of in the form of “try this thing which this system should reject as impossible” and then you realize that after your change that in fact the business rules actually handle this corner case gracefully instead of exposing… I’ll admit this trade off is probably helped by our clients being mostly state govt, so they’re often grateful to get anything at all functional. One never wants to ship bugs, we’ve actually had plenty of positive client interaction along the lines of “Hey, we’re just impressed you caught it and are the process of pushing a fix before we even noticed”.
- zachmu 3y agoOver the lifetime of a code base, most of the value of a test is not in verifying functionality. It's keeping you or someone else from accidentally breaking it when you change something.
- throwaway74432 3y ago>I get paid for code that works, not for tests A blog post could be written about just this statement and how it contributes to a low trust workplace where those who cut corners are favored by stakeholders and everyone else is left scrambling to clean up the messes left in their wake. If you're writing code for yourself, sure, be targeted and conservative with your tests. But when you're working with others, for goodness sake, put the safety nets in place for the next poor soul that has to work on your code.
- makeitdouble 3y agoThat quote is totally true though. Ultimately, tests are there to make sure code works, not for tests' sake. The rest of the sentence you're quoting being "so my philosophy is to test as little as possible to reach a given level of confidence" Overall the approach in the OP looks to me like a decently balanced take, trying to aim for enough tests without excess.
- godelski 3y agoIt's impossible to prove that code works. But tests are a strong indicator and at least put bounds on where the program does work. If you're paid to write code that works, you're paid to write tests. This is fairly standard in every other engineering field.
- makeitdouble 3y agoThis is the kind of shortcut that gets easily forgotten after a while IMHO. Why you write tests is important, and for instance coverage numbers are not that. Most automated coverage assessments still won't guarantee you're testing all the critical patterns (you just need enough to touch all the paths) and a low number doesn't always mean it's not enough. I understand the use as an heuristic's, but as it gets widely adopted it also becomes more and more useless. I mean, today we see people eyeing at LLMs to boost their coverage numbers automatically, and that trend of writing low effort tests has been going all for a while IMO.
- yolovoe 3y ago[flagged]
- xlii 3y agoI was never a huge TDD believer and I’m not one today. It’s not about tests but about managing complexity. When working with highly dynamic languages like JavaScript, Python, Ruby (at least couple years ago) tests were the only tools we had to handle it. Today some of the most common issues are caught by popular static typed languages (Rust, TypeScript). There are some very smart tools like prop tests and if someone is really deep into modeling it’s also possible to test concepts with TLA+ (fun if you need to explore infinite possibilities of reality bending scenarios). Also qualities of certain languages also make code easier and stabler in domains - e.g. Erlang/Elixir, Clojure or Haskell (and there are much more but those are in my monkey zone). But in the end for me testing is just that: Managing ever growing complexity. And since tracking and maintaining change due to distributed development effort is hard the smallest common denominator I know is… write it twice.
- throwaway2037 3y agoThis is a great post. I feel the same on many topics. > I was never a huge TDD believer and I’m not one today. I'll never forget this savage takedown of TDD by Cedric Beust: https://www.beust.com/weblog/the-pitfalls-of-test-driven-development/ https://www.beust.com/weblog/the-pitfalls-of-test-driven-dev... The best rebuttal to TDD that I ever heard from a developer that I respected: "I don't care if you use TDD or whatever. When you commit code, tests need to be included." That's it. And yet, the TDD evangelicals are like the vegans of diet or calisthenicians of fitness -- always annoying, no matter what they say. (Side note: I was a vegan for many years, and I still thought the loud ones were annoying!) > write it twice This is what I hate so much about unit tests -- you chisel the statue once from stone, then, by writing unit tests, you essentially chisel the inverse to fit your new statue. So exhausting.
- xlii 3y agoAnd yet this is one of the best methods knows to man (See double entry accounting which is used with great success in finance for hundreds of years). An interesting approach I’ve seen was - instead of writing tests - describing step by step scenario for fellow engineer to run on review. When it broke it was due to implicit assumptions or unclear instructions. I believe it was right solution as system input was triple digits of free form documents versus hundreds of configurations and rulesets. I couldn’t imagine dataset for those. Effect was more concise, stable test suite and a rather boring system.
- mellutussa 3y ago> was 100% sure that I know how the code works and that this can’t happen again. No, no, no and nope.
- mellutussa 3y agoThere was this sailor who always did extra knots and shit on his ropes and lines. Because he wanted it to be extra safe. Then in a storm the ship sank because he couldn't undo that shit quickly enough. I guess you can tank your software project too by too much or wrong testing.
- kookamamie 3y ago> maybe there’s no correlation between software quality and tests Bingo. The underlying assumption that tests are some god-given faultless spec is flawed. In fact, the tests themselves shouldn't be considered to be of any higher quality than the poor code they test. Good teams produce good software, regardless of the approach or ideology.
- SPBS 3y agoTests are very much a "if you liked it then you shoulda put a test on it" thing. If some end-user property is desirable to you, make sure a test covers it so that you can be sure it still works that way. This is a godsend when doing extensive refactors on the codebase, which lets you move fast. https://twitter.com/simonw/status/1701764953114546664 https://twitter.com/simonw/status/1701764953114546664: "The single biggest productivity enhancement I've ever found for my own personal projects is writing comprehensive tests for them. I don't mean TDD - I rarely write tests first - I mean trying to never land a feature or fix a bug without a test that proves that it works" > No tests that click through the UI and screenshot and compare and hit the network. That's fair, I think testing UI is just hard in general and it's easier to rely on users to submit bug reports especially if the UI rarely breaks.
- pydry 3y agoThe whole article can be boiled down to tests are an investment and you should make sure your investments have a return. It sounds obvious but it's a valid point. It's more common for people to have a dogma based attitude to testing to an investing approach.
- ThalesX 3y agotl;dr; be careful with your testing strategy, it might break your company. There's some aspects that stand to gain from automated testing, it's important that developers test their work, but in my opinion manual testing is super valuable especially for an early stage start-up. I was working for a startup that first had the strategy of testing everything, at a point when we didn’t even start working on the product. We spent a horrible ammount of time getting the test system up and running (microservices, browser add-ons), with the UI testing being the most challenging. But then they also wanted to "move fast and break things", which we did, so we spent an annoying ammount of time fixing breaking tests. This is when the “let’s just delete the test” expression started popping up. So we ended up with an unmaintained testing system that no one cared about anymore, keeping our build on red because we always had the testing system in the backlog so it was going to be done at one point. Then, the quality of the codebase started degrading to the point where we’d wake up with features that have been broken for a while and no one noticed it, even though each individual developer was testing their own flows. We had no one / nothing testing the entire system. At this point, I suggested we hire 1 – 2 manual testers, as our testing strategies are obviously failing, the product is suffering in terms of quality and we could get them for relatively cheap compared to dev-time. I’ve had great success working with manual testers for very complex products with real world repercussions, compared to this tiny start up in dev tooling. They refused. So they decided we’d do cross functional testing and then test the entire system whenever we’d do a merge. So developer velocity fell from a cliff because we became the manual testers. We still had no customers at this point. Runway got shorter and shorter. And the start-up became a statistic.
- ninetyninenine 3y agoThe counterfactual may have been the same. There is a cost to having tons of tests too.
- xiwenc 3y agoPerhaps the key with testing strategy is to define what to test. As pointed out, code coverage is not a good metric when followed blindly. I believe there are at minimum 2 levels that need testing. First is unit testing. It should ensure complex functions work as intended as a unit. Second, for your core functionalities, have e2e test cases. This ensures your product actually works for the end user. Unit testing should make up the biggest part of the test suites. E2e should be kept at minimum but yet satisfies the quality requirements. Few years back i wrote a bit about this based on the AAA method: https://cinaq.com/blog/2019/05/05/simple-high-value-tests-with-python-flask/ https://cinaq.com/blog/2019/05/05/simple-high-value-tests-wi... The core idea is to determine what are high value tests.
- rpigab 3y agoAt home, I only test parts of code where I think I need it, like regexes. At work, as a web developer in Java Spring, I am required to have at least 80% coverage, enforced by Jacoco and tracked in Sonar. This means that if I write a method in an adapter that accepts a single parameter and only makes a single call to another method on some repository passing the parameter to it without transformation and returning the result, I must write a test for it, with Mockito. At this point I'm not testing my code, I'm testing the Java API itself and the JVM it's running on, which I find hilarious. I'm hoping one day this kind of test fails, which will mean that Java does not garantee that reaching a method call will result in a method being called.
- CuriouslyC 3y agoTest coverage requirements are such a band-aid on bad engineering culture. Just instruct reviewers to look at the coverage during code review and flag uncovered important business logic, *BOOM* problem solved. If you can't trust your code reviewers to do this, that's a sign they're either lazy or incompetent.
- rpigab 3y agoAlso, Goodhart's law.
- forgotusername6 3y agoI don't do TDD. Tests help me write code. I write the code, iterate on it, then write tests to confirm the assumption I made in the code. The act of writing then uncovers misunderstandings I had, careless bugs etc. When the tests are complete, they protect my code from someone else changing the code such that the things I wanted the code to do no longer happen. Is every test I've ever written useful? No, absolutely not. But I don't have time/am not best placed to determine which would be useful at the point of writing, so I write them all.
- ChrisMarshallNY 3y ago> people who give a damn. That’s the money shot, in my own opinion. I write a lot of tests, but I am also constantly testing my work, even after shipping (in my latest app, it has been out since January, and is already at 1.2.0, mostly due to fixes for corner cases and UI improvements). I don’t think there’s any substitute for Giving A Damn.
- blackbrokkoli 3y agoMan I don't love these unspecified, context-free takes on "building software". The author lists a lot of tech stack, but no scope, no context, no size estimates about the software he's built. Now we have a thread where the gal writing FORTRAN for legacy nuclear plants is telling the guy doing IT for a mom and pop how he's approaching tests all wrong. Meanwhile, the third person working in an expanding startup who has "ninja" in his job description understands neither and looks down on both. Not very productive.
- roland35 3y agoI think everyone has something in common though - they all think testing is a waste of time until all the sudden things start not working at all and it isn't!
- 2-718-281-828 3y agowhich should be an experience almost any software developer makes within one year of developing software
- ninetyninenine 3y agoI've worked in start ups where they don't have automated tests. The code quality is not as good. But that imo is not a symptom of lack of tests. It's a symptom of being a start up and the desire to move fast. As for mysterious error or bugs that occur the rate at which we get them in the code is no larger than code with tests. It turns out that manual testing and static checking is mostly enough. I think for most stuff testing is an illusion. It's one of those faith based mantras developers follow without any basis in science. There are a few applications where I feel automated tests are required but they are not the majority of software projects.
- rightbyte 3y ago> It's a symptom of being a start up and the desire to move fast. My take is that you can only move fast when you are not in a hurry. If you are in a hurry, you need to work slower, or things will be messed up with no time to fix it. Thus, usually, there is no need to move fast, since if you can, you are not in a hurry anyway.
- begueradj 3y agoNot convinced at all. Clearly a quick and poorly written article. Short and not elaborated. Clickbait.
- keybored 3y ago> First, my credentials. More than half of all the code I wrote in my life is test code. My name is attached to hundreds of pages of TDD. Am I supposed to trust you more or less from this introduction? > But neither codebase has tests, for example, that take a long-ass time to run. Is this for real? Any hint of tongue-in-cheek? The two projects have tests. But they don’t have long-ass-running tests? It sounds like they might test the core data structures and core logic. But maybe they don’t “integrate” by setting up and tearing down all application state and friends. So maybe what is ostensibly missing are those integration tests to complement the lean core-logic tests. And it’s surprising that the heavy integration tests leave little return? Why? They are long-ass-running, they are hard and messy to set up, and in the worst case they just work as a smoke test, which you can do manually once in a while anyway. Losing faith on testing for what? They already do testing.
- tdudhhu 3y agoI have more faith in 'fail fast/early'. It's impossible to test everything. But if you fail fast at least you know something is broken.
- sethammons 3y agoOur paying customers don't appreciate bugs. A layer of testing _is_ fail fast/early because it prevents an issue before it gets to the customer. One gig I had: several times per week the site would go down and nobody could log in. People were just pushing code and yolo'ing into prod. Paying customers were churning. The solution? A single cypress test on a canary node that logged in and edited one setting. 2 or 3 minutes of barrier to production and now the site doesn't go down periodically with deploys.
- tdudhhu 3y agoI am not saying you should not test. To me it seems obvious that you test what you create. But it is impossible to test everything. When you fail fast you are quicker to notice it yourself. And when it was deployed without you nothing the bug it will still surface very early because it breaks your app. Broken gets fixed, shitty lasts forever.
- sethammons 3y ago> Broken gets fixed, shitty lasts forever. Love it. More succinct than "devs need to feel the pain to fix it."
- ninetyninenine 3y agoThis blog post makes a valid point. The problem is testing is a religion. People believe in it via faith. Also like religion, such faith is almost impossible to fully remove. Not without hard evidence. The actual way forward is data. Data driven metrics that can show the cost of tests vs. No tests.
- EVa5I7bHFq9mnYK 3y agoTDD was born when people moved from proper typed languages to javascript, where compiler doesn't check anything and a test suite compensates for that. Some MBA educated managers probably believed 100% coverage = no bugs. Real 100% coverage means you make sure every possible input produces an expected output. Which is impossible to do, except for the simplest systems.
- tanepiper 3y agoIn the team I lead, we haven't really shipped tests in over a year. We've also been in discovery and proof-of-concept phase, and where I do agree with the author is it's not worth it to spend tireless hours on having green pipelines at this stage. However we are now starting to put our platform towards a what will become a production state, and it's clear that without some level of testing, the confidence of stability is too low. (Thankfully our codebase is lightweight and mostly integration between systems, with some components for a third-party tool)
- ninetyninenine 3y agoI think most people believe in tests based off of anecdotal data and gut feelings. This post is saying that it is possible that much of those anecdotes and gut feelings may be erroneous. Similar to how believers of religion show faith in belief despite clear lack of evidence. All humans are capable for falling for faith-based tropes... as shown by basically almost every variation of religion in the world. Its quite possible that if we have some way to measure the quantitative cost of tests vs. No-tests that those results would form a definitive basis against testing. Think about it. Like a believer your automatic reaction may be to counter this argument. But think, does your belief in testing have any basis in science at all? Likely not.
- dhbradshaw 3y agoI like the way that testing as I go alters my code base, making each bit of code runnable in isolation and so easy to understand and fix. I hate flaky tests and I hate slow tests. Figuring out how to eliminate both takes care, but it gets rid of two of the three main negatives of testing. The third main negative for me is that testing can prematurely add friction around changing things like function signatures. Mitigating this is addressed well in my current favorite essay on effective testing: https://matklad.github.io/2021/05/31/how-to-test.html https://matklad.github.io/2021/05/31/how-to-test.html
- taylodl 3y agoTo me, TDD is about software architecture and design, not testing. Bob Martin said it best - TDD forces you to create a system that's testable. You're not going to get to the end and scratch your head and wonder how are we ever going to test this thing? That's not gonna happen. What TDD isn't about, or at least shouldn't be anyway, is testing every little last minute detail. Like Kent Beck said, if there are classes of errors I simply don't make then why am I creating a test to see if I made that error? That's why I coined the term Maintenance Driven Development (https://taylodl.wordpress.com/2012/07/21/maintenance-driven-development/ https://taylodl.wordpress.com/2012/07/21/maintenance-driven-... - wow, it's been 12 years now!). Test for the types of mistakes you and your team typically make. Make your tests productive. When bugs arise then create a test to re-create the bug, and then fix it. Your test will prove you've fixed the bug. Overtime your test cases will grow and they'll be concentrated around the areas in your software where you're having actual problems. This is what makes testing effective.
- nojvek 3y agoTesting is wonderful- if one starts with the goal of delivering highest quality per unit time spent. If you start with a goal of coverage or every little thing having X unit test, it becomes a pain on everyone. Integration tests are the hardest. They test a huge surface area. They are slow, flakey and hard to debug when something is wrong. But they offer the biggest bang for buck since they test it end to end from the user’s perspective. Unit tests are wonderful for libraries with well crafted and stable interfaces. The bottom line is always making a pragmatic decision - for the unit time spent, where can I add guards with highest bang. SQLite is sqlite because they are obsessed with making it battle ready. Most great quality software has a team that takes the quality bar seriously and has tests to hold that bar. Somethings don’t need that quality bar because - it’s not proven anyone will use it.
- lamontcg 3y ago> In both codebases I’ve merged PRs without any tests and frequently see others do the same. And the world didn’t end and no one shed any tears and the products are still some of best I’ve ever used and the codebases contain some of the most elegant code I’ve ever read. Yeah if someone actually has 10 years of experience in the product and can tell that the code is obviously correct and tests would be annoying for anyone to write, and the contributor will go radio silent if you ask for the tests, and that it makes the product better and it is unlikely to bite you in the ass, then you can ship it. It is vital to have that well-built mental model of what in the codebase is likely to bite you in the ass though. This doesn't necessarily lead to a slippery slope where every PR is full of missing or shitty tests and slide right into game development. The slippery slope argument is a logical fallacy. Similarly, the perfect process isn't one where every PR is cut up until 100 line or less PRs, with unit tests and two developers sit down and do code review to nitpick every other line and then both of them flip the switches like they're sitting in a minuteman silos starting WWIII.
- fireflash38 3y agoThe problem with automated testing is that the people who would best benefit from having a very detailed and comprehensive test suite to verify their work are precisely the same people who cannot write that test suite effectively. Someone who half asses the business code is absolutely going to less than half ass the test code, and now you have no real verification that what they did was right.... And now likely have some shitty tests to take care of too!
- crq-yml 3y agoI think our options basically come down to types, tests, and syntax. Syntax is the least explored because it's expensive to iterate on and teams can't agree on it, so we sit around hoping that the next generation of languages will let us consume a better syntax instead of daring to build one. But a syntactical approach leads towards writing the specification in terms of the high level description, folding up lower level pieces of the problem into bits of syntax, and then having them meet in the middle with "enough" compilation. You end up with a program that is right because it literally can't be written wrong. Types do some specification, but they work in a more bureaucratic sense of "you have to do what you mean" by checking to see if you filled out the right form. Of course teams love doing that, because they need a little bit of bureaucracy. Everyone seems to agree on some amount of type discipline being needed at scale. Tests exist at the far end of specification, where the program is just "what it's been tested to be." This is most useful when you don't have a clear specification and you're mostly looking at the integration side of things, not deeply examining the behavior. Teams can add lots of tests, but the tests are a scattershot and don't always do anything specific. They do have a tendency to tell you how a change will break an existing specification, though.