4 ms·
Junk DNA is, AFAIK, not actively expressed (used to create proteins). It's important, though, in the sense that spacing between gene expression sites is a contr
by aaaronic 5y ago
Junk DNA is, AFAIK, not actively expressed (used to create proteins). It's important, though, in the sense that spacing between gene expression sites is a control on which genes get expressed under which conditions (so the junk adds necessary spacing between important genes).
I did _some_ research on epigenetics during my MS degree. Spacing between sites was an important factor in our modeling of gene expression.
- grishka 5y agoI remember reading how some of the "junk" DNA turned out to be important because while it doesn't make proteins, the "non-coding" RNA it gets transcribed into regulates something.
- ShroudedNight 5y ago> spacing between gene expression sites is a control on which genes get expressed under which conditions This makes it sound like it represents control flow rather than data. If its presence does / can make a material difference on the output encoding, it strikes my non-expert ears as actively perilous to label such DNA 'junk'
- chaxor 5y agoWhat research have you seen on modeling gene expression? I'm genuinely curious, as I haven't really seen many convincingab initio studies towards this. I could see finding certain features like this spacing as predictive of perhaps some other feature, but I haven't seen any research that really tackles generation of gene expression data from first principles and input of DNA sequence. It's my understanding that modeling the kinetics is difficult, as we really haven't tried making the full network of differential equations. Does anyone have a project that points to the 'final solution' to this? I know recently there was a paper in cell that modeled the cell with the smallest viable genome to predict cell division, but that's a bit further away from complete modeling of our 30k genes' (much less isoforms) dynamics.
- jashephe 5y agoGlobal models of gene expression for an entire cell are fairly distant at this point, but there is quite a bit of work into modeling transcriptional activity from sequence. If you're interested in reading more, a relevant technology to search for would be the "Massively Parallel Reporter Assay", or MPRA, which couples pools of 10⁴–10⁵+ synthetic DNA sequences with RNA sequencing to measure transcriptional output. Data from MPRA experiments is being used to train models, although these models are not anywhere near a point where you could model the gene expression of all regulatory elements in a cell; they are usually focused on a specific factor or regulatory sequence.
- sooheon 5y agoNotably DeepMind had a recent paper on using transformers to predict long range interactions in gene expression: https://www.nature.com/articles/s41592-021-01252-x https://www.nature.com/articles/s41592-021-01252-x
- chaxor 5y agoThe "train models" or ML portion is what I'm disappointed with unfortunately. I make ML models to predict things from genetic information somewhat regularly, but we all are aware of the enormous issues with that. I am more interested in the ab initio methods, as I have seen them be spectacularly useful in other fields - like Bethe salpeter equations in condensed matter physics.
- kkylin 5y agoNot just spacing. The sequence also matters as they serve as binding sites for enzymes that can promote or repress the expression of downstream genes. As just one (relative simple) example of how complicated genetic circuitry can be, I really enjoyed & recommend Mark Ptashne's A Genetic Switch: Phage Lambda for anyone who doesn't mind doing some slightly technical reading. Disclaimer: not a biologist, and would be interested in hearing from someone more knowledgeble than I, both about the Ptashne book and about recommended reading.
- toper-centage 5y agoDo junk DNA is like code styling and comments in programming.
- meowkit 5y agoIts closer to a config file / internal functions that modify the state variables of a system instead of generating objects. The junk DNA doesn't explicitly get read, but it interacts in nonlinear ways with the executable "text" portion of the DNA. Also disclaimer: My only knowledge of this is from Nessa Carey's The Epigenetics Revolution and some additional online reading.
- pas 5y agoOr, more like Makefiles, and an enormous amount of cache for the runtime and build system, but this cache never really gets invalidated. Basically it's like using a "buggy filesystem" (like ext2) that only shows the first ~16000 files in a directory. So having too much junk can hide stuff, having too little of it can uncover files that evolution carefully hid, and so on.
- robwwilliams 5y agoNo, this is the wrong metaphor. You want the right metaphor: one brilliant coder, 10 total newbie coders, and five cats walking all over the key boards. And no delete function at all!
- sterlind 5y ago
- robwwilliams 5y agoNope, this is not a claim that is supported well. Yes sure there will be a handful of examples, mostly in promotor and enhancer regions, but insertions or deletions in introns and repetitive non-coding DNA will generally have very modest impact on phenotypes and these mutations generally cannot be extirpated by natural selection. That is why the term “junk” DNA is actually not that bad or incorrect in many cases. If you prefer think of it as “spandrel” DNA in the sense used by SJ Gould.
- asdff 5y agoWhat do you mean by gene expression sites? If you mean enhancers, afaik you can excise and insert them anywhere in the chromosome and they will still effect function in cis.