5 ms·
I work in this space. We manage thousands of e2e tests. The pain has never been in writing the tests. Frameworks like Playwright are great at the UX. And having
by msoad 2y ago
I work in this space. We manage thousands of e2e tests. The pain has never been in writing the tests. Frameworks like Playwright are great at the UX. And having code editors like Cursor makes it even easier to write the tests. Now, if I could show Cursor the browser, it would be even better, but that doesn’t work today since most multimodal models are too slow to understand screenshots.
It used to be that the frontend was very fragile. XVFB, Selenium, ChromeDriver, etc., used to be the cause of pain, but recently the frontend frameworks and browser automation have been solid. Headless Chrome hardly lets us down.
The biggest pain in e2e testing is that tests fail for reasons that are hard to understand and debug. This is a very, very difficult thing to automate and requires AGI-level intelligence to really build a system that can go read the logs of some random service deep in our service mesh to understand why an e2e test fails. When an e2e test flakes, in a lot of cases we ignore it. I have been in other orgs where this is the case too. I wish there was a system that would follow up and generate a report that says, “This e2e test failed because service XYZ had a null pointer exception in this line,” but that doesn’t exist today. In most of the companies I’ve been at, we had complex enough infra that the error message never makes it to the frontend so we can see it in the logs. OpenTelemetry and other tools are promising, but again, I’ve never seen good enough infra that puts that all together.
Writing tests is not a pain point worth buying a solution for, in my case.
My 2c. Hopefully it’s helpful and not too cynical.
- TechDebtDevin 2y agoI doubt that screenshot methods are the bottleneck considering that's the method Microsoft and Anthropic are using.
- tomatohs 2y agoIt's absolutely not the bottleneck. OpenAI can process a full resolution screenshot in about 4 seconds.
- fullstackchris 2y agoTo be fair, this is NOT the case with native mobile apps. There are some projects like detox that are trying to make e2e tests easier, but the tests themselves can be painful, run fairly slow on emulators, etc. Maybe someday the tooling for mobile will be as good as headless chrome is for web :) Agreed though that the followup debugging of a failed test could be hard to automate in some cases.
- edelans 2y agoI think we can claim that at Waldo. Check for yourself: I've just recorded this [1] scripted test on the wikipedia mobile app, and it yields this [2] Replay. In less than a minute we spin up a fresh virtual device, install your app on it, execute the 8 steps of the script. As a result, you get the Replay of the session : video synchronized with interaction timeline, device & network logs, so you can debug in full context. [1]: https://github.com/waldoapp/waldo-programmatic-samples/blob/main/ios/test/onboarding.ts https://github.com/waldoapp/waldo-programmatic-samples/blob/... [2]: https://share.waldo.com/7a45b5bd364edbf17c578070ce8bde22024024 https://share.waldo.com/7a45b5bd364edbf17c578070ce8bde220240...
- egeek 2y agoDo you have any pricing info available? All I can see is get started for free, but no info on what it might cost later
- hn_throwaway_99 2y agoWhile I agree with your primary pain point, I would argue that that really isn't specific to tests at all. It sounds like what you're really saying is that when something goes wrong, it's really difficult to determine which component in a complex system is responsible. I mean, from what you've described (and from what I've experienced as well), you would have the same if not harder problem if a user experienced a bug on the front end and then you had to find the root cause. That is, I don't think a framework focused on front end testing should really be where the solution for your problem is implemented. You say "This is a very, very difficult thing to automate and requires AGI-level intelligence to really build a system that can go read the logs of some random service deep in our service mesh to understand why an e2e test fails." - I would argue what you really need is better log aggregation and system tracing. And I'm not saying this to be snarky (at scale with a bunch of different teams managing different components I've seen that it can be difficult to get everyone on the same aggregation/tracing framework and practices), but that's where I'd focus, as you'll get the dividends not only in testing but in runtime observability as well.
- Lienetic 2y agoAgreed. Is there a good tool you'd recommend for this?
- hn_throwaway_99 2y agoIt's been quite some time but New Relic is a popular observability tool whose primary goal (at least the original primary goal I'd say) is being able to tie together lots of distributed systems to make it easier to do request tracing and root cause analysis. I was a big fan of New Relic when I last used it, but if memory serves me correctly it was quite expensive.
- krainboltgreene 2y ago"OpenTelemetry and other tools are promising, but again, I’ve never seen good enough infra that puts that all together." It's a two paragraph comment and you somehow missed it.
- cschiller 2y agoThanks for your thoughtful response! Agree that digging into the root cause of a failure, especially in complex microservice setups, can be incredibly time-consuming. Regarding writing robust e2e tests, I think it really depends on the team's experience and the organization’s setup. We’ve found that in some organizations—particularly those with large, fast-moving engineering teams—test creation and maintenance can still be a bottleneck due to the flakiness of their e2e tests. For example, we’ve seen an e-commerce team with 150+ mobile engineers struggle to keep their functional tests up-to-date while the company was running copy and marketing experiments. Another team in the food delivery space faced issues where unrelated changes in webviews caused their e2e tests to fail, making it impossible to run tests in a production-like system. Our goal is to help free up that time so that teams can focus on solving bigger challenges, like the debugging problems you’ve mentioned.
- Terretta 2y agoIntegrate with https://www.honeycomb.io https://www.honeycomb.io
- rafaelmn 2y agoI think either you're overselling the maturity of the ecosystem or I've been unfortunate enough to get stuck with the worst option out there - Cypress. I run into tooling limitations and issues regularly, only to eventually find an open GitHub issue with no solution or some such.
- ec109685 2y agoThere are silly things that trip up e2e tests like a cookie pop up or network failures and whatnot. An AI can plow through these in a way that a purely coded test can’t. Those types of transient issues aren’t something that you would want to fail a test for given it still would let the human get the job done if it happened in the field. This seems like the most useful part of adding AI to e2e tests. The world is not deterministic, which AI handles well. Uber takes this approach here: https://www.uber.com/blog/generative-ai-for-high-quality-mobile-testing/ https://www.uber.com/blog/generative-ai-for-high-quality-mob...
- tomatohs 2y agoI predict an all out war over deterministic vs non-deterministic testing, or at least a new buzzword for fuzzy testing. Product people understand that a cookie banner "shouldn't" prevent the test from passing, but an engineer would entirely disagree (see the rest of the convos below). Engineers struggle with non-deterministic output. It removes the control and "truth" that engineering is founded upon. It's going to take a lot of work (or again, a toung-in-cheek buzzword like "chaos testing") to get engineers to accept the non-deterministic behavior.
- tomatohs 2y agoYou're totally right here, but "debugging failed tests" is a mature problem that assumes you have working tests and people to write them. Most companies don't have the resources to dedicate full engineer time to QA, and if they do nobody maintains the test. Debugging failed test is a "first world problem"
- AdieuToLogic 2y ago> ... "debugging failed tests" is a mature problem that assumes you have working tests and people to write them. I am reminded of an old s/w engineering law: Developers can test their solution or Customers will. Either way, the system will be tested.
- codedokode 2y agoSorry if it is a stupid idea, but cannot you log all messages to a separate file for each test (or attach test id to the messages)? Then if the test fails, you can see where the error occured.
- msoad 2y agoWhere I work there are 1,500 microservices. How do I get a log of all of those services -- only related to my test's requests in a file? I know there are solutions for this, but in the real world I have not seen it properly implemented.
- ergeysay 2y agoAs you said, OpenTelemetry and friends can help. I had great success with these. I am curious, what were implementation issues you have encountered?
- antonvs 2y agoThis works easily enough in the major cloud environments, since logging tends to be automatic and centralized. The only thing you need to do is make sure that a common request id or similar propagates to all the services, which is not that difficult.