3 ms·
These examples show the system being very tightly coupled to the actual html. IME this leads to very brittle tests that fail due to restructoring/reorganizing/r
by Afton 7y ago
These examples show the system being very tightly coupled to the actual html. IME this leads to very brittle tests that fail due to restructoring/reorganizing/redesigning.
My two cents: I've never seen automated visual regression testing that wasn't terrible to work with, and where the most common result (by a large margin) of a test failing was someone would update the expected file/image so it would pass with the new visuals. It's a hard problem, and one that I've personally decided isn't worth doing for the customer facing software that I've been involved with.
- swampthing 7y agoHave you tried Percy? There certainly can be false positives at times, but you can just approve the diffs so it's not a huge pain in the end.
- sbergot 7y agoBeing coupled to the HTML is more stable than with a screenshot. GUI testing takes a lot of effort but it can absolutely be worth it. I work in a SaaS B2B app and we have been saved a few times by them. They also allow us to engage in big refactorings for reasonable costs.
- gnagatomo 7y ago>the most common result (by a large margin) of a test failing was someone would update the expected file/image so it would pass with the new visuals Previous solutions that my team developed ended up like this. We found the problem can be split in two: 1. Figuring out if some change is actually wanted/expected We found that doing a d-hash + hamming distance (levenshtein) is a very good way to handle tolerance, as it will ignore most text and subtle layout changes, but unlike most examples the screenshot is scaled down to 200px wide (this number is arbitrary, we're still figuring out how much tolerance it does provides). 2. How to handle this changes This is the hard part. For now we are only selecting the images that failed in the previous step and by using some jenkins plugins they are presented during the pipeline side by side (old, diff, new) and are reviewed manually, but the versioning process is automatized and the results are stored in the changelog.
- webdiff 7y agoI'm working on a product that focuses on doing visual regression testing on the html & css level, and the beta will be happening in one or two months. If anyone is interested in it drop your email here! https://webdiff.io https://webdiff.io
- Dowwie 7y agoThat's the point of testing. The test identifies a break. A person investigates and either updates the test or fixes a problem.
- 2rsf 7y agoThat's not how it should work, you should include updating of the tests in the scope of any task and make sure enough time and resources are assigned. Ideally the tests should be maintained by the developers themselves so they know how and where to quickly change them to accommodate code and behavior changes.
- Afton 7y agoI feel like you may not have read my comment very generously. I am very familiar with the point of a test. Let me try again to explain. There is a cost to a brittle test. UI testing suffers from this more than other kinds because there are many 'plausible' UI arrangements, and as the product shifts and changes, you need to distinguish 1. "The dialog moved slightly to the left"/"We refactored the HTML, but it still looks the same" from 2. "The dialog is now underneath another element". Suppose it takes 10 minutes to find, fix, get reviewed, and push the fix, deploy the fix, and validate the fix. If the number of failures that are more like the former are RADICALLY more than the latter type of failure, then it is easy enough to come to the conclusion that the test is not giving you a reasonable ROI. Perhaps the cost of releasing a latter-type-of-failure is not that bad, if you can fix it and get it to production quickly. It might be cheaper overall then the ongoing maintenance cost on a test that would prevent this failure. Also, and this is culture and product dependent, but we're talking about this like it's a single test, when it's usually a suite of tests (or multiple suites). If they have a 5% failure rate, and 95% of the time its really an 'update the test, this is the new expected', people will stop trusting the tests, and will take shortcuts. So you may find that instead of spending 100% of the maintenance costs for the suite of brittle tests, you're spending 70%, but only getting 20% of the benefit because people become accustomed to the failures , and once a test is failing, no one will notice that the failure changed from a 'benign failure' to a "customer-can't use" failure. (Note: numbers are imaginary, but not crazy). At one company I worked for, it was so bad, that when we tried to introduce testing/checkin rigor, multiple developers pulled me over to make me explain "Why this dumb test is failing on my checkin attempt", and we would look at the logs and other artifacts to uncover that "It's failing because you changed something without updating the relevant tests". It took quite a bit of time to re-train developers used to brittle tests, to respect and maintain non-brittle ones. And that is why I am against automated UI testing in general. :)