6 ms·
Most tests should test only one aspect of the system and that aspect is most often correctness. Dirty databases make the tests non-deterministic. You can use a
by AtNightWeCode 5y ago
Most tests should test only one aspect of the system and that aspect is most often correctness. Dirty databases make the tests non-deterministic.
You can use a pre-seeded database for testing. SQLite can be used sometimes for instance.
Create load tests to test the performance.
Data can and sometimes should be reused for a specific set of tests. This is needed for browser tests for instance.
As a person who worked on systems with thousands of tests. I cannot stress enough that there should never be any randomness added to the automatic tests. Most times it is better to add more tests then to make a test do many things.
- ryanbrunner 5y ago> Dirty databases make the tests non-deterministic. If this is the case, is the test really testing something useful? Tests should verify that the code works as it's intended to work. If pre-existing (but presumably valid) data can break the test, then surely there's states where whatever behaviour is being tested won't work in production? Provided the test is creating and verifying its own data, it's not really non-deterministic except in a very technical sense.
- bsaul 5y agoI personnally use tests for two things : - in tdd, to make sure my code works in expected scenarios - in CI to make sure new code don't break old features. In both case, i need to control my testing environment. Randomizing inputs could be used in a third scenario for extensive testing, a bit like fuzzing. But it's yet another different usage.
- ryanbrunner 5y agoI guess the disconnect is that I don't think of the state of the database as inputs, per se, but more the environment the test is running in. I wouldn't for instance necessarily need to set my clock to a fixed time (provided I'm not specifically testing code where that is relevant), or wipe out the hard drive that the code is running on for a test.
- deleted 5y ago[deleted]
- michaelt 5y ago> If pre-existing (but presumably valid) data can break the test, then surely there's states where whatever behaviour is being tested won't work in production? Imagine I'm testing a web page that shows a list of all customers who have placed at least one order. If I set up the test by wiping the database and adding ten customers, I can assert that the list of customers should contain those ten and no others, giving it ten rows. If I don't wipe the database, the web page may show more than ten customers - and I can't tell if the extra entries indicate a fault in my code or not.
- ryanbrunner 5y agoSure, but in that case I wouldn't test that the list contains exactly the 10 customers I expect, since that's not what the code is supposed to do. That's a shortcut you introduced while testing. I usually approach things like this with two assertions - that the size of the list under test increased by X, and that the newly created record(s) exist in the list. After all, the code doesn't show the 10 most recent customers that purchased an item, it shows ALL customers. You could hide a pagination bug where only the first page of results functions if you wipe the database first.
- sseagull 5y agoBut then how do you deal with unique constraints or things like upserts? If I want to add 10 entries, then I need to know that all have been added and not pre-existing, or else the code for adding new entries doesn't actually get tested. Or you end up with spurious test errors if insertion fails due to duplicates.
- dragonwriter 5y ago> But then how do you deal with unique constraints or things like upserts? If I want to add 10 entries, then I need to know that all have been added and not pre-existing, or else the code for adding new entries doesn't actually get tested. Then your test code needs to take into account the current state of the database when creating input data for the code under test. It is more complicated than with a know start state where you can pick fixed data, sure, but it's doable.
- kazinator 5y agoWhat if the test thinks it is creating its own data, but that has failed, and the old data is mistaken for new data?
- someotherperson 5y ago> As a person who worked on systems with thousands of tests. I cannot stress enough that there should never be any randomness added to the automatic tests. To take a counter position, I'm a person who has also worked on many systems with many thousands of tests. Every single test suite I've seen pulling static fixtures as test data was brittle and near worthless. Introducing randomness to the data you test against is the only way to actually test your product with near-realistic inputs. To this end, packages like Faker[0] are worth their weight in gold. [0] https://github.com/joke2k/faker https://github.com/joke2k/faker
- AtNightWeCode 5y agoWith random in this case, I mean data that changes between test runs. I often use generated data sets. If you can find errors by feeding random data into your tests, the tests does not cover the code in the first place.
- michaelt 5y agoWell, the question is what you do when a test starts failing non-deterministically. If your team respond to a random test failure by finding the root cause and fixing it, uncertain test inputs could be right for you. On the other hand, if your team responds by hitting the 'rerun tests' button until it passes? In that case random failures are just slowing you down.
- goto11 5y agoThe correct response is to "freeze" the parameters of the failing test into a regular test which fails consistently.
- someotherperson 5y agoI'm not entirely following here. What is the point of even having tests if there are teams out there who disregard their outputs? Freezing a set of parameters that always pass, in case you don't run into the failing edge case, defeats the entire point of testing in the first place. If it's one test that is failing because the test itself is defective, and if that test is not for a crucial part of the product, then skip the test in the interim and come back to it later. But giving it a happy path for success sounds really strange to me.
- pmarreck 5y ago> Dirty databases make the tests non-deterministic. If this breaks your tests, you have a concurrency bug, pure and simple. Either in your code, or in your test assertions. Your test assertions (like your code) should never assume that what you just put into the database is the only thing in there which will match, or the last thing that was put in there. In my case (working in Elixir/Phoenix), it runs all the tests concurrently. This will cause any concurrency issue to crop up real quick, and that's exactly how I like it.
- opportune 5y agoSome tests simply cannot be deterministic because the underlying system is not deterministic (eg testing something that has multiple threads, perhaps some of which live in far-flung members of members, etc). A dirty database is what I would call “avoidable non determinism” because you can easily clean it and you probably do not want to test handling bad databases (maybe you do though). Thread scheduling is not something you would want to work around, because you want to test a wide variety of actual production scheduling patterns to make sure there is not something like a deadlock (which also means, to get full value from the test, you probably need to run it hundreds of times and/or turn on sanitizers). One perhaps better word you are looking for re determinism is hermeticism. A dirty database with junk from previous runs makes the test non-hermetic (not self contained). You can have a hermetic but non-deterministic test also, like telling a bunch of threads to print the ABC’s.
- mjr00 5y ago> As a person who worked on systems with thousands of tests. I cannot stress enough that there should never be any randomness added to the automatic tests. Most times it is better to add more tests then to make a test do many things. Completely agreed. To put a real cost to this: I worked at AWS on a major service. The test suite ran nightly (it required a shared environment and the tests took almost 10 hours; an argument for the importance of independence and test speed for sure, but that's a different discussion). It was also filled with spurious failures; on a good day, each team might "only" have to deal with 10 random failures, while a "bad" day might have seen 20-30. As such, a process was put into place: each team had a dedicated "QA on-call" resource that rotated every week. The QA on-call's role was to, every day, look at the test logs and investigate the failures assigned to their team, providing an explanation of why the test failed, and certifying that it wasn't actually a problem. Due to the nature of the system, this generally took all day regardless of whether it was a good or bad day, though on bad days it might even require pulling in an extra person on your team to help. So with 10 teams and let's call an average AWS engineer's salary $250k, you're looking at $2.5 million/year spent just due to flaky tests requiring manual investigation. In fairness, these tests were flaky because they were integration suites using real networking, hardware, and other AWS services, so flakiness was somewhat inevitable. But intentionally introducing nondeterminism and random failures in tests would lead to the same thing.