5 ms·
Show HN: Finalrun – Spec-driven testing using English and vision for mobile apps
I wanted to test mobile apps in plain English instead of relying on brittle selectors like XPath or accessibility IDs.
With a vision-based agent, that part actually works well. It can look at the screen, understand intent, and perform actions across Android and iOS.
The bigger problem showed up around how tests are defined and maintained.
When test flows are kept outside the codebase (written manually or generated from PRDs), they quickly go out of sync with the app. Keeping them updated becomes a lot of effort, and they lose reliability over time.
I then tried generating tests directly from the codebase (via MCP). That improved sync, but introduced high token usage and slower generation.
The shift for me was realizing test generation shouldn’t be a one-off step. Tests need to live alongside the codebase so they stay in sync and have more context.
I kept the execution vision-based (no brittle selectors), but moved test generation closer to the repo.
I’ve open sourced the core pieces:
1. generate tests from codebase context
2. YAML-based test flows
3. Vision-based execution across Android and iOS
Repo: https://github.com/final-run/finalrun-agent https://github.com/final-run/finalrun-agent
Demo: https://youtu.be/rJCw3p0PHr4 https://youtu.be/rJCw3p0PHr4
In the Demo video, you’ll see the "post-development hand-off." An AI builds a feature in an IDE, and Finalrun immediately generates and executes a vision-based test for it verifying the feature developed by AI.
- deleted 6mo ago[deleted]
- arnold_laishram 6mo agoLooks pretty cool. How does your agent understand plain english?
- ashish004 6mo agoWe have built a QA agent that can understand your plain english intent and uses vision to reason and navigate the app to test your intent. You can check our benchmark here https://finalrun.app/benchmark/ https://finalrun.app/benchmark/ and how we architected our agent for the benchmark https://github.com/final-run/finalrun-android-world-benchmark https://github.com/final-run/finalrun-android-world-benchmar.... Its all open source
- arbaaz 6mo ago[dead]
- sahilahuja 6mo agoAgentic testing. Kudos to your decision to open-source it!
- ashish004 6mo ago[dead]
- avikaa 6mo agoThis solves a massive headache. The drift between externally generated tests and an active codebase is a brutal problem to maintain. Using vision-based execution instead of brittle XPaths is a great baseline, but moving the test definitions to live directly alongside the repo context is definitely the real win here. Did you find that generating the YAML from the codebase context entirely eliminated the "stale test" issue, or do developers still need to manually tweak the generated YAML when mobile UI layouts change drastically? Great project!
- ashish004 6mo agoHi Avikaa, finalrun provides skills that you can integrate with any IDE of your choice. You can just ask the finalrun-generate-test skill to update all the test for your new feature.
- rootally7 6mo ago[dead]
- gavinray 6mo ago> The shift for me was realizing test generation shouldn’t be a one-off step. Tests need to live alongside the codebase so they stay in sync and have more context. Does the actual test code generated by the agent get persisted to project? If not, you have kicked the proverbial can down the road.
- ashish004 6mo agoYes gavinray, It gets persisted to the project. Its lives alongside the codebase. So that any test generated has the best context of what is being shipped. which makes the AI models use the best context to test any feature more accurately and consistently.
- ashish004 6mo agoJust updated README.md, it's lot simpler and addresses on the core. Thanks for the feedback, please checkout
- srinidhigs829 6mo agoI just ran my first test. Thanks team :)
- ashish004 6mo agoDo share us feedback.
- usual_engineer 6mo agoVerification of AI generated code right would be dope. We do something similar in our company for web with playwright but facing a lot of flaky tests. Will check this out
- ashish004 6mo agoHey, thats true, verification of ai generated code needs proof with video of each action, console logs and network logs. Would love to know how you are solving for web. would be a great learning for me too.
- clyp 6mo ago[dead]
- l2s0 6mo ago[dead]
- Mousumi-Dhar 6mo agoLove the simplicity of it.
- cadamsdotcom 6mo agoSorry to ask a dumb question, but, why not move the tests into the repo? Monorepos have many benefits chiefly being able to commit atomically reduces incidental complexity from drift. It’s good enough for Google and Facebook!
- ashish004 5mo agoTrue. Thats exactly what we have done. Once finalrun is installed, it ships all the skills to generate and run the test from the same repo.