3 ms·
How easy is it to check the results of cell annotations for mistakes? Is it easy for a person to do, and this will save them a bunch of time getting a baseline
by ibash 2y ago
How easy is it to check the results of cell annotations for mistakes?
Is it easy for a person to do, and this will save them a bunch of time getting a baseline? Or could this lead to a bunch of mislabeled data?
- celltalk 2y agoIt still not 100% accurate but it should be useful for baseline annotations.
- gww 2y agoUsers of these kinds of tools should check that their marker genes are associated with the labelled cell types. There are known markers for many cell types across multiple organisms.
- f6v 2y agoI’ve been doing this for the past three years, it’s very challenging. I think most of these tools do well on very broad cell types like seen on the GitHub page. But the thing is that if you’re working with e.g. immune cell types then you can effortlessly label high-level clusters yourself. The real challenge is identifying fine cell subsets, like different types of CD8 T cells: naïve, central memory, effector memory, Temra, etc. I don’t think it’s a problem that can be solved by a tool though. One issue is that “classical” cell type definitions are based on flow cytometry which uses antibodies to define cell types. These definitions don’t translate that well to scRNA-seq as it’s a completely different protocol. For example, naïve and central memory cells are separated based on CD45 isoforms and this information isn’t available in single cell gene expression with 10x Genomics(most popular protocol). Another issue is that people use ad-hoc cell type definitions. It’s common to come up with a random gene as a cell state marker. Which means the definitions aren’t comparable between studies and a lot of manual curation is required. Which makes some sense because “cell type” is an abstraction. In real world, the cell types and states are much more complex and are often continuous rather than discrete. Taken together, building cell type classifier is a very difficult task that depends much more on the data quality, context(which tissue data comes from), and training labels. You can build a very decent classifier with a regression model if you have good data.
- celltalk 2y agoYou’re right in most of your comments however a fine-tuned model will be able to pick up nuances which are missed by a general model. For instance, AVP is used for HSCs, yet not related to flow cytometry at all. If you fine-tune an LLM with experts by your side one time it will be able to give you the grained cell types such as T-cell subsets. Plus, a regression model won’t give you the reasoning behind the given cell type annotation.