4 ms·
To be fair, you can get really very far with non-machine-learning automated techniques (I did good-enough algorithms for an in-situ fluorescent gene sequencing
by itamarst 6y ago
To be fair, you can get really very far with non-machine-learning automated techniques (I did good-enough algorithms for an in-situ fluorescent gene sequencing image processing pipeline at one job). I suspect any form of good automated processing, regardless of whether it's AI, would be welcome to biologists.
If you'd ever like to chat about the automation parts, I'd be interested in hearing how you're approaching it; the niche of "scientific computing, but repeatable" is quite different than traditional scientific software, and it seems like people are still in early stages of figuring out how to do it.
Would also be interested in hearing how you approach correctness. The best approach I've discovered is metamorphic testing. Basically you modify real inputs, and then ensure the output matches. E.g. you say, "OK, I have this finished algorithm that segments cells, `f(image) -> cells`. Now, if I double brightness on everything, I would expect the same results, so let's see what `f(brighter_image)` is, I would expect same output." Or like "If I merge nuclei-looking splotches that cross what original segmentation boundary was, that should result in fewer cells." Unfortunately only discovered this technique after I left the image processing job, so haven't had chance to try it.
- mike210 6y agoDefinitely! There are some applications where traditional methods are just simply good enough. However, when they aren't, it can be incredibly frustrating. From our conversations with scientists, this kind of data (3D, histology, difficult tissues, new assays) is increasing in volume. As for correctness, we've only done mAP scores and traditional accuracy metrics so far to compare with other algorithms, but we also have our own internal metrics and a test set we're building out in-house to cover many edge cases, many of which cover some of the things you are talking about. One thing we're always trying to be sensitive of is fairness. We want to make sure that we're not biasing the test to our algorithm, which would make us look better than we are.
- itamarst 6y agoI guess when I say correctness, I mean "how do I know it _continues_ to be correct on data we've never seen before". That's where metamorphic testing can be valuable, because it lets you at least find incorrectness on real-world data that hasn't been hand-tagged.
- mike210 6y agoAh, yes. We're even looking to use some generative models in order to even do variations based on data and then compare that we do similarly well between cases. I guess the point I was making was that we want to make sure we don't then use this generated or modified data in order to test other algorithms in the space and say we're better. Simply put, it would be unfair for us to make changes to perform better on a hurdle and then put other algorithms through those hurdles. But for internal use, it's definitely great!
- mike210 6y agoAlso would love to chat if you ping me about the automation part - would love to get some feedback there. michael at biodock dot ai