6 ms·
An ambitious project to map all the cells in the human body
- untilHellbanned 9y agoThis will be useful, but similar to hype machine behind react or node.js, molecular biology is jerked around by new technologies that confer unclear value over existing approaches. In this case, it’s single cell rna seq. I’d argue we never got very far with bulk measurement RNA analysis because it’s not a functional technique, rather than us not having single cell resolution. Look at the nobel prizes. So many of them have simple genetics or biochemistry at their core. It’s because those experiments were functional. And before you tell me we haven’t yet had enough time for gene expression studies to deliver a nobel prize, I’d say well it’s been almost 25 years. Yamanka 4 factors is less than 10. CRISPR, also less than 10 should come soon too.
- nonbel 9y agoWhile I like this project, it seems too ambitious for the current state of affairs. Start with counting how many cells of each type are present for various tissues. Do this using cadavers of various ages, etc. Pretty much all we have right now for such estimates are totally back of the napkin.
- vanderZwan 9y ago> Start with counting how many cells of each type are present for various tissues. I'm confused: what do you think this project is trying to achieve if not exactly that?
- nonbel 9y agoThey are planning on sequencing, etc. That is all great but honestly I think it is jumping the gun. I am saying to simply count the cells, so we have the most basic of information. Work on this only really began in 2013...[1] From (quickly) reading their whitepaper it isn't clear to me whether they will even get the count data. [1] https://www.ncbi.nlm.nih.gov/pubmed/23829164 https://www.ncbi.nlm.nih.gov/pubmed/23829164
- vanderZwan 9y agoI'm sorry, maybe I'm missing something but what you're saying is not making any sense to me. Furthermore it gives me the impression you completely misunderstand the biological science behind this project. I thought you meant figure out the number of different types of cells, but you're actually saying counting the numbers of cells of each type? You're putting the cart before the horse, since we haven identified each cell type yet. That's what this project is doing! What does counting cells in tissues even mean here? Which cells? If you can't even tell the cell types apart, what exactly would you be counting? Do you think we already know every kind of cell type and what does? Because we don't. And even ignoring that: say that we figure out the liver has N hundred million cells. On average, because tissue size wildly varies per person. But hey, sure, from a Pure Science perspective we could figure that out. What does that information then tell us? Is it information in any way meaningful? How does counting the number of cells give us a better understanding of what that tissue does? Would you suggest we try to make sense of how countries are organised and what kind of culture they contain by grouping them by population size? Apparently Ghana and Nepal are identical according to this logic, as are Cameroon and Taiwan, and Niger and Sri Lanka[0]. By comparison, determining cell types via gene expression is like measuring education levels and which percentage of the population is in which profession. I don't know about you, but when it comes to making sense of how the body works, my money is on that. [0] https://en.wikipedia.org/wiki/List_of_countries_and_dependencies_by_population https://en.wikipedia.org/wiki/List_of_countries_and_dependen...
- nonbel 9y ago"Type" of cell is a human construct. There is no actual thing as a type of cell, so the list is always complete. People could decide to have more or less types depending on what they are working on, but there is no sense in trying to delineate all the "types". Currently we already have a classification system for cell types that is used to distinguish between different types of cancer. They should use exactly that same system. Of course the numbers will vary by individual. What we want to get is basic upper/lower bounds. It would probably be best to see how well # of cells correlates with weight, skin area, etc and split the data into groups based on whatever easily measured physiological parameter would work best. Once you have number of cells at various ages, then any biological theory that requires the collective activity of many cells will need to be consistent with these values. This can produce a major constraint on much theorizing, especially in areas like cancer. Say you have a theory that cancer is caused by multiple mutations (eg, n = 7) accumulating in a cell lineage. You have other data about mutation rates per basepair per division (eg, p = 1e-8). If this theory is correct you would need a certain number of divisions to explain the curve of cancer incidence by age for that tissue. That number of divisions can be estimated from the difference in number of cells at various ages (or at least we can get some bounds on it by ignoring cell death, etc) and compared to the value required by the cancer theory.
- vanderZwan 9y agoIt's likely that most people here are not up to date with just how quickly the field of biology has changed in the last decade. Last year I somehow bumbled my way into programming for Sten Linnarsson's group[0], via a HN who's hiring thread, no less! That group turned out to be Kind Of A Big Deal in this field (I mean, I was just impressed by how well-written Sten's code was, especially for a professor, and liked the idea of programming for scientists). The stuff I learned about this field is pretty mindblowing. Think of the wall of text below as context for this article, as explained by someone not inhibited by proper knowledge of what he's talking about. We start with a tiny biology refresher. Every cell in your body is essentially a clone: barring mutations, they all have the same genetic information encoded in their DNA. We often think of DNA as nature's code. To build on this metaphor, think of your stem cells as freshly installed PCs with all possible software you could need, pre-loaded on a gigantic HDD: your genome. And just like a program on your disk, genes don't do anything until they're activated. To "run" the code in a gene, the cell makes active copies of it: RNA. In oversimplified terms, a copies of RNA represents loading a genetic program into RAM and running it. The rest of the cell is basically the wetware required to run that code, which is important too: a PC is kinda useless without IO (just look at the struggles early nerds had with simply finding a use for the Altair 8800 if you don't believe me[1]). Anyway, going from stem cell to a specific cell-type is like setting up those identical PCs for the different things that we use PCs for by opening different programs. Unlike PCs our wetware is massively parallel: to get more performance in a specific task our cell just creates more RNA copies, "loads many copies of the same program". Because of this, we can measure the activity of a gene by measuring the number of RNA copies of it in a cell. Now imagine we're trying to reverse engineer an alien computer, and unlike Independence Day[2], we're not dealing with Mac-compatible hardware here. In programming, when you try to reverse engineer what someone else's code does, a disassembly probably won't cut it. Similarly, knowing the genome will not be enough: to really make sense of all it all, to "debug" the DNA, we need to see that "running code" in context, in relation to the whole organism. What molecular biologists have been doing in recent years is measure the gene expression (the number of RNA molecules) of individual genes in individual cells, across as many cells as possible, at various stages of cell development. Then they compare the gene expression levels with various algorithms to organise them. Combined with extra meta-data, like what tissue the cells came from, what cell-type we know the cell belonged to based on morphology (in plain English: its visible shape in a microscope), and at which point in development cell was harvested, we can then start painting a picture of cell development and cell types. They take cell tissues, separate it into individual cells, then measure the expression of each gene in each cell. The techniques they have figured out techniques, like droplet-based approaches[3], to do this in bulk are mind-blowing. When given the budget, and in the hands of appropriately skilled molecular biologists, you can now "easily" measure the expression of all genes in hundreds of thousands of cells. Doing this you can figure out the different cell types, which genes are involved them, and which stage of development cell types start (dis)appearing. But remember that we separated all of the cells, so we lost the context of the original tissue. To fix this, you can take new tissue samples, and apply techniques like smFISH or MERFISH[4] to to attach fluorescent markers to individual copies of RNA of a specific gene. Gene expression is measured just by counting the number fluorescent dots on microscope slide, each representing an individual RNA molecule. Yes, biologists actually do this (with software help of course), and it works. You end up with a map of gene expression in the tissue. Also the pictures look beautiful. This is, in a nutshell, single-cell transcriptomics[5]. Or at least the part I am exposed to through the research group I work for. It's mind-boggling and awesome to see what has become possible in the span of just a few years, and the field hasn't plateaued yet. Thanks to all of this data we can discover all kinds of wonderful things that were invisible before. A simple example: last year it was discovered that one type of neuron in the sympathetic nervous system turned out to have seven sub-types from the genetic point of view[6]. The hypothesis is that having many sub-types makes sure they "wire up" correctly; you wouldn't want to have goosebumps whenever your heart beats. Of course, the part of the article that went viral was that we have a type of neuron solely responsible for goose bumps and nipple erections... Anyway, the speed and size at which data is being gathered is rapidly growing. The article even mentions this explicitly. New techniques for dealing with this growing mountain of data are developed rapidly too. For example, the research group I work for just co-published a paper with Peter Kharchenko's group, describing a technique to estimate in which direction a cell's RNA expression is moving[7]. So instead of just having scalar tSNE scatter plots, we get tSNE vector fields, showing us the direction in which cell development is moving. And the coolest part is that it can be applied to existing data sets, because it relies of various types of measurements that were already being done anyway. I'm just standing at the sidelines of all this, being incredibly impressed and humbled that I get to try to contribute a little bit to all of this. I'm just a programmer/interaction designer. Sten Linnarsson hired me to develop a user-friendly, flexible and fast web-based data browser, the Loom Viewer, for his new Loom data format. But this post is long enough as is, to I'll reply to myself with a comment explaining that. Anyway, this is an exciting, ambitious project, that probably will have an enormous impact on biology and medical research, on a much shorter time-scale than you think. [0] http://linnarssonlab.org/ http://linnarssonlab.org/ [1] https://en.wikipedia.org/wiki/Altair_8800 https://en.wikipedia.org/wiki/Altair_8800, https://www.youtube.com/watch?v=1FDigtF0dRQ https://www.youtube.com/watch?v=1FDigtF0dRQ [2] https://scifi.stackexchange.com/questions/15141/how-did-the-computer-virus-get-uploaded-into-the-mothership-in-independence-day#15143 https://scifi.stackexchange.com/questions/15141/how-did-the-... [3] http://mccarrolllab.com/dropseq/ http://mccarrolllab.com/dropseq/, https://directorsblog.nih.gov/2015/06/02/single-cell-analysis-powerful-drops-in-the-bucket/#more-4718 https://directorsblog.nih.gov/2015/06/02/single-cell-analysi... [4] http://thenode.biologists.com/fishing-fish-2/resources/ http://thenode.biologists.com/fishing-fish-2/resources/, https://www.youtube.com/watch?v=-jIZ3bH-rAE https://www.youtube.com/watch?v=-jIZ3bH-rAE [5] https://en.wikipedia.org/wiki/Single-cell_transcriptomics https://en.wikipedia.org/wiki/Single-cell_transcriptomics [6] http://ki.se/en/news/special-nerve-cells-cause-goose-bumps-and-nipple-erection http://ki.se/en/news/special-nerve-cells-cause-goose-bumps-a..., http://www.nature.com/neuro/journal/v19/n10/full/nn.4376.html http://www.nature.com/neuro/journal/v19/n10/full/nn.4376.htm... [7] https://www.biorxiv.org/content/early/2017/10/19/206052 https://www.biorxiv.org/content/early/2017/10/19/206052, http://velocyto.org/ http://velocyto.org/
- xvilka 9y agoBy the way, the finished or in progress previous great mapping efforts, like human genome or human brain connectome - are those datasets available to download somewhere?
- vamin 9y agoNot sure about the brain connectome project, but for the human genome project (and many other genome projects) whole genomes are available for browsing and download here: https://genome.ucsc.edu/cgi-bin/hgGateway https://genome.ucsc.edu/cgi-bin/hgGateway.