3 ms·
> The question is, why are you unit testing ML functions? ML is not software engineering, and unit tests not a good way to broadly evaluate them, outside some “
by george3d6 6y ago
> The question is, why are you unit testing ML functions? ML is not software engineering, and unit tests not a good way to broadly evaluate them, outside some “third rail” type tests for insanely important cases.
In practice, I agree with this, and that's how I usually write my tests when dealing with ML code. Though honestly I think this perspective could well extend beyond ML code to any code with very broad input and output spaces.
> You’d typically use statistical testing (e.g. f1 scores) to determine whether your models are performing as expected, but they are ultimately probabilistic in nature and highly unlikely to satisfy all of the intended cases simultaneously.
Yes, but the solution to that (in practice) is to basically benchmark against your own (previous) code and alternative implementations. Based on the idea that "if 10x tried it, and mine is better, than I must be doing something correctly" and/or "if the previous version did worst, well, then something must be improving".
Again, I think this is equally true for testing a ray tracer or compilation times as it is for testing an ML algorithm.
> There is of course the philosophical question of how you determine the quality of your tests, which are themselves software. The practical answer is to keep your testing software as trivially simple, obvious, and boring as possible; tests can then be validated with the mark I eyeball.
But this results in a situation where "good" tests are extremely subjective, and replacing a team of engineers with another (or even just a member of that team) would result in the need to replace the tests.
> Testing is obviously not impossible, so let’s set that aside for a moment.
> None of this has any bearing on the usefulness or otherwise of automated testing.
To me it seems to have baring in that, unless I am able to have at least some heuristic based on which I can define "good" or "sufficient" testing the question of what to test and how much time to allocate to it becomes impossible to answer. And that is a very useful question I'd want to have an answer to that is better than just "based on my intuition and what I think is sufficient as to not have my bosses/customer scream at me when stuff breaks, this amount of testing is good".