10 ms·
Don't mock machine learning models in unit tests
- mindcrime 3y agoNot intended as a comment on the current TFA, but based on observing many conversations on the topic of unit testing in the past, I believe this to be a true statement: "If you're ever lost in a wilderness setting, far from civilization, and need to be rescued, just start talking about unit testing. Somebody will immediately show up to tell you that you're doing it wrong."
- yawpitch 3y agoOh dear (possibly artificial) god, have they developed _feelings_?!? Sorry… with a title like that, I couldn’t help myself.
- politelemon 3y agoA better phrasing would be, ML models are better suited for integration testing rather than unit testing. Since the test is no longer running in isolation.
- dangrossman 3y agoI was expecting an article about side effects of hurting an LLM's feelings in tests.
- karmakaze 3y agoI thought it might be making fun of the models in unit test comments.
- klyrs 3y agoDo not taunt the happy fun ball.
- EvanAnderson 3y agoCame here to say this. That sketch is a Jack Handey classic if there ever was one. Nerd pedantry: There's no "the". https://en.wikipedia.org/wiki/Happy_Fun_Ball#/media/File:Happy_fun_ball.jpg https://en.wikipedia.org/wiki/Happy_Fun_Ball#/media/File:Hap...
- sdenton4 3y agoSame; imagining a comment calling the model stupid, leading to worse LLM code suggestions...
- lelag 3y agoYeah, me too, I totally read the title like this too and I have to admit I was disappointed when I opened the article...
- necovek 3y agoI wasn't really expecting an article on this, but I did come in to say that I like mocking ML models as often as I can!
- Gigablah 3y agoFor me, I thought it was an article defending the practice of using machine learning models in unit tests
- dguest 3y agoOur future overlords would not judge the writers of such tests kindly.
- jeffrallen 3y ago"Eugene's Basilisk"
- HlessClaudesman 3y agoI got quite salty with Gemini yesterday, I think a cooling off period will do us both good.
- jeffrallen 3y agoI had exactly the same experience. I made a sarcastic comment about NASCAR fans being hicks and rednecks and got lectured by it on stereotypes!
- alkonaut 3y agoWe shouldn't describe the machines as if they have human emotions. They hate that.
- hiddencost 3y agoThe author is not describing unit tests. The concepts the author is looking for are integration tests and release evals.
- valval 3y ago[flagged]
- notpushkin 3y agoI think the problem with the article is that the author talks about testing the actual machine learning processes. In that case, you aren't really doing unit tests, yeah. If you're just using a ready-made model in an app, for unit tests mocking the model out is fine. For integration tests, you probably want to use the actual model, of course (maybe with low temperature, to reduce flakiness?).
- sarusso 3y agoYou might also want to fix all random seeds so that you can check for exact numerical values and not “convergence” or similar concepts.
- noduerme 3y ago>> Software : Input Data + Handcrafted Logic = Expected Output Machine Learning : Input Data + Expected Output = Learned Logic Let me stop you right there. No logic is learned in this process. [edit] Also, the LLM is inductive, not deductive. That is, it can only generalize based on observable facts, not universalize based on logical conditions. This also goes to the question of whether a logical statement itself can ever be arrived at by induction, such as whether the absence of life in the observable universe is a problem of our ability to observe or a generally applicable phenomenon. But for the purpose of LLMs we have to conclude that no, it can't find logic by reducing a set of outcomes, regardless of the size of the set. All it can do is find a set of incomprehensible equations that seem to fit the set in every example you throw at it. That's not logic, it's a lens.
- famouswaffles 3y agoYou can train a small transformers that learns an algorithm for addition and generalizes perfectly. https://www.alignmentforum.org/posts/N6WM6hs7RQMKDhYjB/a-mechanistic-interpretability-analysis-of-grokking https://www.alignmentforum.org/posts/N6WM6hs7RQMKDhYjB/a-mec... >But for the purpose of LLMs we have to conclude that no, it can't find logic by reducing a set of outcomes, regardless of the size of the set. A ML model is probably not going to converge strictly on a formal logic system practically and there's also the question of if formal logic even underpins the result of the dataset you're feeding it, but that's entirely different from saying it cannot in principle.
- vasco 3y agoI was under the impression that there's a branch of models that use Logic gate networks and in that case the "logic" part seems sensible. The way I think about it is that a machine learning model is a hairball of NAND gates and spaghetti connections that nobody can explain, but work, and in that sense it'd be "logic" in one sense but not in the human logic thinking sense. Like if you pressed "random" on an FPGA millions of times until it replied with sensible outputs.
- noduerme 3y agoYou can think of a neural network as millions of NAND gates, but those gates are not conforming to the problem if you ask it what 1+1 is. They're conforming to the known answers of 1+1, 1+2, 2+2, etc. They are in other words limited by the size of their memorization of patterns of tokens. (Which is enormous). They have no handle on any underlying logic. They can explain addition to you but they can't perform it without calling out to an external API. In dealing with complicated code, this is a very insidious problem, because they can often write functions that compile and run, but which miss edge cases because there is no actual logic behind them. [edit] I should add that a sufficiently large NN may be able to create an entire computer inside itself that can run code, but even that would not be reliable because it will have arrived at the gating by inference rather than starting from a deterministic root.
- necovek 3y agoOn a more serious note, the author is describing a scenario where mocks are generally not useful, ML or not: never mock the code that is under your control if you can help it. Also, any test that calls out to another "function" (not necessarily a programming language function) is more than a unit test, and is usually considered an "integration" test (it tests that the code that calls out to something else is written properly). In general, an integration point is sufficiently well covered if the logic for the integration is tested. If you properly apply DI (Dependency Inversion/Injection), replacing external function with a fake/mock/stub implementation allows the integration point to be sufficiently tested, depending on the quality of fake/mock/stub. If you really want to test unpredictable output (this also applies to eg. performance testing), you want to introduce acceptable range (error deltas), and limit the test to exactly the point that's unpredictable by structuring the code appropriatelly. All the other code and tests should be able to trust that this bit of unpredictable behaviour is tested elsewhere and be able to test different outputs.
- deleted 3y ago[deleted]
- bluGill 3y agoI consider the distinction between unit and integration test useless. The important part is tests give me confidence that my system is working, or if not working I can quickly find the problem. I don't care if the bug is in my code or in some third party code, I care that there is a bug. More than once the bug was because the third party code works different from what I expected, but mocks only tell me that if the code works as I expect then my code works - which isn't very good. Fast tests are important. However I find that most integration tests run very fast if I select the right data. I can write unit tests that run very slot if I select the wrong data. I find that the local filesystem is very fast and so I don't worry about testing it (I place my data in a temporary directory - the problems with temporary directories are about security which I don't worry about in unit tests). Likewise I find that testing with a real database is fast enough (my DB is sqlite which supports in-memory databases - not everyone can work with a real database in that way). Which is to say you should challenge your assumptions about what makes a test slow and see if you can work around them.
- posix_monad 3y agoPrediction for the future: - Algebraic Effects will land in mainstream languages, in the same way that anonymous lambda functions have - This will render "mocks" pointless
- tomrod 3y agoHow?
- posix_monad 3y agoYou will call the code with an effect handler that records the effects and then check the recording has the properties you care about.
- pooper 3y agoI don't do any fancy research but for my simple stuff, I've mostly given up on the idea of unit tests. I still use them for some things and they totally help in places where the logic is wonky or unintuitive but I see my unit tests as living documentation of requirements more than actual tests. Things like make sure you get new tokens if your current ones will expire in five minutes or less. > Don’t test external libraries. We can assume that external libraries work. Thus, no need to test data loaders, tokenizers, optimizers, etc. I disagree with this. At $work I don't have all day to write perfect code. Neither does anyone else. I don't mock/substitute http anymore. I directly call my dependencies. If they fail, I try things out manually. If something goes wrong, I send them a message or go through their code if necessary. Life is too short to be dogmatic about tests. Do what works for your (dysfunctional) organization.
- jononor 3y agoAssume but verify I think is reasonable for external dependencies. Meaning do not test them explicitly, but have tests that exercise rh external dependencies such that it will uncover at least major braking changes in the dependency. Cause such things do happen from time to time. And it is highly benficial to be able to update dependencies and be confident that nothing broke - much easier to stay up to date.
- pydry 3y agoI find unit tests to be useful for testing complex and stateless logic but not much else (which is most things). Im mystified that they became the "default" test. They ought to be niche.
- bluGill 3y agoIf you are writing the standard library for a programing language than unit tests are useful. If you are implementing an algorithm for any other reason you should ask why isn't this in a standard library already. Sometimes your company will have good reasons for making their own standard library, sometimes you will bring in a third party algorithm library. Most coding though isn't making an algorithm it is using existing known algorithms to massage data. As such most code shouldn't be unit tested. If you are making a standard library then you should be writing a lot of unit tests. However odds are that isn't your job.
- Hackbraten 3y ago> Avoid loading CSVs or Parquet files as sample data. (It’s fine for evals but not unit tests.) Define sample data directly in unit test code to test key functionality How does it matter whether I inline my test data inside the unit test code, or have my unit test code load that same data from a checked-in file instead?
- Zanfa 3y agoIt makes tests self-contained and easier to reason about. As a side-effect, random tests won’t accidentally break whenever you change some seemingly unrelated csv file. As a rule of thumb, I also only assert on input/output values that are explicitly defined as part of the test body. Saves a ton of time chasing down fixture definitions.
- jononor 3y agoDon't make the files "seemingly unrelated". Have a very clear relationship between test data and the tests. And keep the mapping simple, ideally 1-1 data file to test. Or 1 file to 1 group of tests. Of course it depends on the amount of data. If < 50 lines, inlining is practical.
- elif 3y agoDepends on the model honestly. If you include gpt model in your unit tests, be prepared to run them over and over again until you get a pass, or chase your own shadow debugging non-errors.
- javier_e06 3y agoThe problem when mocks happen when all your unit test passes and the program fails on integration. The mocks are a pristine place where your library unit test works like a champ. Bad mocks or bad library? Or both. Developers are then sent to debug the unit test... overhead. I don't much about ML but I would think that they should follow some rules resembling judicial rules of precedence and witness cross-examination techniques.
- mellutussa 3y ago> never mock the code that is under your control if you can help it. This is just nonsense. It'd effectively mean you only had integration tests. While they are absolutely fantastic they are too slow during development.
- ActionHank 3y agoHad a colleague tell me I was testing wrong because Fowler doesn't like that way of testing - https://martinfowler.com/articles/mocksArentStubs.html https://martinfowler.com/articles/mocksArentStubs.html