Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Visweshyc
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
Visweshyc
7mo ago
We see this as different from review. The system generates tests to catch second-order effects and executes them against the live application to expose bugs
2.
▲
by
Visweshyc
7mo ago
Fair feedback. Will make that clearer. Appreciate it
3.
▲
by
Visweshyc
7mo ago
Good point. To keep the regression tests reliable as the app evolves, we run a reliability cascade. First, we generate and execute deterministic Playwright from the codebase. If execution fails then we fall back to DOM and aria tree. If tha
4.
▲
by
Visweshyc
7mo ago
We evaluated test generation using Claude code and our purpose built harness and measured the quality of tests in catching the unknown unknowns. We noticed Claude Code misses the second order effects that actually break applications. You a
5.
▲
by
Visweshyc
7mo ago
The system focuses on going beyond the happy path and generating edge case tests that try to break the application. For example, a Grafana PR added visual drag feedback to query cards. The system came up with an edge case like - does drag f
6.
▲
by
Visweshyc
7mo ago
Thanks! To execute these tests reliably you would need custom browser fleets, ephemeral environments, data seeding and device farms
7.
▲
by
Visweshyc
7mo ago
Yes we currently support web apps but plan to extend the foundation to test mobile applications on device emulators
8.
▲
by
Visweshyc
7mo ago
Thanks! We believe executing the scenarios and showing what actually broke closes the loop
9.
▲
by
Visweshyc
7mo ago
Thanks for the feedback! - Agreed that the form factor can be condensed with a link to detailed information - With the codebase understanding, backend is where we are looking to expand and provide value - The intelligence of the models doe
10.
▲
Launch HN: Canary (YC W26) – AI QA that understands your code
58 points
by
Visweshyc
7mo ago
|
26 comments