3 ms·
Few questions: 1. Are you seeing a lower number of production hot fixes? 2. Have you tried other models? Or thought about using Output of different
by vladdoster 1y ago
Few questions:
1. Are you seeing a lower number of production hot fixes?
2. Have you tried other models? Or thought about using Output of different models and using combining the output if they have a delta.
3. Other than time & cost, what benchmarks in terms of software are considered (i.e., less hot fixes, etc.)?
This is really cool, btw
- jampa 1y agoThanks! > Are you seeing a lower number of production hot fixes? Yes, with E2E tests in general: They are more effective at stopping incidents than other tests, but they require more effort to write. In my estimation, we prevent about 2-3 critical bugs per month from being merged into main (and consequently deployed). For this project specifically: I think the critical bugs would have been caught in our overnight full E2E run anyway. The biggest gain was that E2E tests took too much time in the pipeline, and finding the root cause of bugs in nightly tests took even more time. When a test fails in the PR, we can quickly fix it before merging. > Have you tried other models? Or thought about using output from different models and combining the results when they differ? Not yet, but I think we need to start experimenting. Claude went offline for 30 minutes over the last 2 days, and engineers were blocked from merging because of it. I'm planning to add claude-code-router as a fallback.