15 ms·
I'm sure it's more complex than I grasp as a layperson, but I'm utterly amazed at how simple this _appears_. I get the feeling that this is something I have a b
by drtz 6y ago
I'm sure it's more complex than I grasp as a layperson, but I'm utterly amazed at how simple this _appears_. I get the feeling that this is something I have a better chance of understanding than the average SaaS Terms and Conditions.
I expected to have to scroll through pages upon pages of indecipherable text. Instead it's no bigger than a large paragraph of text, and I can easily fit it on my screen.
- flobosg 6y agoSequencing technologies have improved immensely over the last decade and a half. And, in this particular case, getting the sample RNA is incredibly easy, since its purity and integrity in the vial is quite high.
- shellfishgene 6y agoI didn't look at the details of how they sequenced it, but given that there are chemically modified bases in the mRNA vaccines there is a chance the normal methods for sequencing (and the first step of translating to DNA) don't work. Well, I guess in practice they did.
- flobosg 6y agoWhile not completely equal to the naturally occuring bases, the modified bases in the vaccine mRNA need to be able to complement to the non-modified ones present in tRNA anticodons during translation. If they can pair to their corresponding natural bases, then the chemically modified RNA can be also used as a template by the reverse transcriptase to generate the complementary DNA needed for the sequencing reaction.
- shellfishgene 6y agoGenerally I agree, but it could be the case that the modified bases work just well enough for tRNA matching in the ribosome, but not with the reverse transcriptase.
- flobosg 6y agoThe mechanism of base complementarity is identical in both cases. If a modified uracil complements an adenine in tRNA, it will complement an adenine in the RT primer or an adenine being added to it.
- lettergram 6y agoIt really is that “simple.” Getting it designed and building it is more difficult. At its core, it’s a piece of mRNA that creates a protein. That code gets transcribed into a protein (often those are relatively short). That protein then triggers your bodies immune response, which trains it to attack covid19. Inject this mRNA into a cell and it’ll create the protein. Anything can be injected at this point once the mechanism for injection is developed
- wombatpm 6y agoWhich makes me wonder. Could you place the entire virus genome in these liposomes and get them to hijack the machinery to make an entire virus? Like plasmid but for viral structures?
- lettergram 6y agoYes, that's one of the concerns many have about this technology. Literally, anything can be injected and done at this point. Not sure where the technology is exactly at, but I suspect we're no more than 5 years from major incident related to this. Even this vaccine, we really don't know the long-term impacts or risks involved with this. For instance, this vaccine does appear more risky than the standard flu vaccine: https://wonder.cdc.gov/controller/saved/D8/D137F981 https://wonder.cdc.gov/controller/saved/D8/D137F981 Presumably this is due to increased inflammation. It's not hard to imagine that we'll be doing genetic editing soon enough with this (if we aren't already).
- sterlind 6y agoHow do you edit genes with an mRNA vaccine? You'd need DNA, enzymes (maybe requiring post-translational modification) to splice them in, etc. Also, you might not even be able to print full viruses with this platform. manufacturing mRNA is different from manufacturing all the random types of RNA in a virus, isn't it?
- lettergram 6y ago
- hfjfktmtkrn 6y agoIt's not really that simple. Only two companies in the world succeeded, the French company Sanofi which also tried making a mRNA vaccine failed.
- fermienrico 6y agoIt’s like looking at the binary file and saying “that’s pretty simple” while ignoring the massive amount of machinery that allows us to run that file and use it (CPUs, Motherboards, computers, etc). I presume a whole bunch goes into making vaccine and this is just the top of the iceberg.
- WheelsAtLarge 6y agoTrue, most pharmaceuticals can't do it now but given the right knowledge, which is known, it can be done relatively fast. I suspect in the next few years there will be many companies that will be able to replicate and advance the process.
- puzzlingcaptcha 6y agoHere is a breakdown https://berthub.eu/articles/posts/reverse-engineering-source-code-of-the-biontech-pfizer-vaccine/ https://berthub.eu/articles/posts/reverse-engineering-source... discussed previously https://news.ycombinator.com/item?id=25538820 https://news.ycombinator.com/item?id=25538820
- tablespoon 6y ago> This is somewhat of a problem for our vaccine - it needs to sneak past our immune system. Over many years of experimentation, it was found that if the U in RNA is replaced by a slightly modified molecule, our immune system loses interest. For real. > So in the BioNTech/Pfizer vaccine, every U has been replaced by 1-methyl-3’-pseudouridylyl, denoted by Ψ. The really clever bit is that although this replacement Ψ placates (calms) our immune system, it is accepted as a normal U by relevant parts of the cell. Neat.
- seiferteric 6y agoUmm, isn't that kind of scary? Like could you create a virus with this Ψ that our immune system can't fight at all?
- carlmr 6y agoMaybe I'm thinking too simplified here, but wouldn't this only work on the first iteration? After all the virus would replicate with Us in your cells and then the replicas wouldn't have the advantage anymore.
- bobthebuilders 6y agoBy that point the cell would be producing the associated protein though. Getting it inside the cell is the goal here from what I've read.
- carlmr 6y ago
- azernik 6y agoThe protein they're trying to manufacture is indeed quite simple - AFAIU both BioNTech and Moderna put together their sequences in a weekend. (Though there was a more involved process of winnowing down the sequences for the most effective ones.) The technically challenging parts are: - delivery mechanism: you need to take a very unstable molecule, protect it from the environment - both external, and when inside the patient - and insert it into a human cell. (This is called the "platform", and is usually developed independently from the specific payload.) - manufacturing: both producing the mRNA itself at a large scale, and inserting it into the delivery mechanism, at a large scale and in low-temperature conditions - testing: the newly-developed payload and the existing platform were integrated at small scales within weeks, but testing the thing for safety and efficacy took months EDIT: As schoen pointed out, this was not actually released by Moderna, but reverse engineered by third-party researchers. Original text was: "Hence they feel safe releasing this. Their moat is not the gene sequence, their moat is everything else."
- schoen 6y ago> Hence they feel safe releasing this. Their moat is not the gene sequence, their moat is everything else. One or more of the vaccine developers may have released such details, but this particular file is a reverse engineering effort by unaffiliated scientists based on analyzing the dregs of used vaccine vials (!). Edit: See https://news.ycombinator.com/item?id=26628594 https://news.ycombinator.com/item?id=26628594 for more substantive discussion about this.
- azernik 6y agoAh - thanks for pointing this out! Edited to make sure readers see it.
- wespiser_2018 6y agofrom what I've gathered, the rate limiting step for production as of yet, is creating the lipid vesicles and getting the RNA inside of them. Only a few companies have a process for this, and the supply chain for the precursors is limited as well.
- purple_ferret 6y agoAny individual protein doesn't seem that complex since it's just a combination of some 20 amino acids, but the variations are endless: "Since each of the 20 amino acids is chemically distinct and each can, in principle, occur at any position in a protein chain, there are 20 × 20 × 20 × 20 = 160,000 different possible polypeptide chains four amino acids long, or 20n different possible polypeptide chains n amino acids long. For a typical protein length of about 300 amino acids, more than 10^390 (20^300) different polypeptide chains could theoretically be made. This is such an enormous number that to produce just one molecule of each kind would require many more atoms than exist in the universe."
- MauranKilom 6y agoThe exponentiation signs got lost in your quote. Would you mind adding them back in?
- abfan1127 6y agoproteins are also unique in that not just their sequence matters, but also their physical shape. 2 proteins can have the same sequence but a different physical shape, and therefore have different impacts on the body's chemistry. I started a PhD researching DSP methods for matching protein sequences and locations of amino acids. Fun stuff.
- Gatsky 6y agoThen there are also post translational modifications, like addition of acetyl or phosphate groups, and sugars to the protein (glycoproteins). I mean, I can understand how an eye or a brain can evolve by natural selection, but I’m still stunned by abiogenesis. I guess we’ll never know for sure how it all started.
- mattnewton 6y agoI think it's a bit like a private key- the difficulty is in finding some combination that works in an absolutely massive space of possible proteins, not necessarily in the length of the protein.
- airstrike 6y ago> and I can easily fit it on my screen. ...with GATACCA right in the middle, but unfortunately with no GATTACA that I could find.
- staplung 6y agoHeh. Technically, there isn't even GATTACA in there since it's RNA and hence all the T's are actually U's. It's just convention to use the T's. GAUUACA doesn't have the same ring to it.
- softwaredoug 6y agoGreat quote from Maurice Hilleman, creator of many (most?) of our childhood vaccines goes something like “Don’t be smart. Instead be careful and accurate” Lots of these things aren’t complicated. It’s the careful systematic testing and public trust building that’s the hard part.
- gerdesj 6y ago"but I'm utterly amazed at how simple this _appears_." Biology is a funny old thing. You can look at that concise description - the orange and so on blocks of a few letters and a few short groupings. Now ATCG are basic building blocks but they consist of quite a lot of stuff. I think it's a bit more complex than that because this is RNA not DNA so ATCG might not be quite right. Each of those bases are horrifically complicated depending on scale. Search "ATCG" - this is a good start: https://en.wikipedia.org/wiki/Nucleobase https://en.wikipedia.org/wiki/Nucleobase Now dive into one of those bases and decompose it to its constituent atoms. Now look at the maths around this stuff. It gets quite complicated, quite quickly. That said, the fact that a bloody complicated thingie can be described so concisely is absolutely amazing and as you say it looks so simple.
- flemhans 6y agoIt'd be cool to make an easy-to-use interface, still.
- xjlin0 6y ago"but I'm utterly amazed at how simple this _appears_." Remind me the joke of the consultant engineer knows where to make X by the chalk. LOL
- amluto 6y agoIt appears simple, but a whole lot of work went in to producing that string even pte-COVID. Some of it is generic in the sense that it might apply to any mRNA vaccine. Some is quite specific: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5584442/ https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5584442/ There’s also (IIRC, no citation right now) prior work suggesting that coronavirus vaccines against the spike are likely to be effective and that vaccines against the N protein might be counterproductive.
- jldugger 6y agoLiken it to the 4kb demoscene: it's amazing what can be done with a little bit of information, as long as you don't have to describe the machine running it. Or the distribution method, or even really invent the thing, since you're mostly just copying someone else's work. Plus it doesn't have to even do anything. In fact, doing anything might be a problem, so best to just sit there and look menacing (and spikey).
- GuB-42 6y ago> Liken it to the 4kb demoscene Coincidentally, the mRNA sequences for both vaccines are about 4kb (kilobase) long.
- fnord77 6y agoeach one of those letters represents a ~15 atom molecule, so in a way it is a compressed representation
- Black101 6y agoso, explain it to me ?
- DecoPerson 6y agoCheck out this video by The Thought Emporium to see how far we’ve come in these matters: https://youtu.be/J3FcbFqSoQY https://youtu.be/J3FcbFqSoQY This should hopefully provide you with some useful perspective.
- anxrn 6y agoNot Moderna, but this [1] was a very useful primer on grokking how the Pfizer vaccine works, especially for computer programmers. [1] https://berthub.eu/articles/posts/reverse-engineering-source-code-of-the-biontech-pfizer-vaccine/ https://berthub.eu/articles/posts/reverse-engineering-source...
- ceratin6 6y ago> I'm sure it's more complex than I grasp as a layperson, but I'm utterly amazed at how simple this _appears The genetic code is simple. But possibly not the nanobot that I felt like attached to my optical nerve with a slightly painful prick behind my right eye. I’m not kidding. I never have eye pain and within a day or two of my vaccination, I had a prick feeling once in the back of my right eye, later accompanied by light coming from within my eye as it was closed in the dark for only a monent for a small part in or close to the middle of my visual field. I wish I was making this shit up. It was really the only weird side effect, coincidental or not, of the vaccine for me. The good part is that my vision in that eye has been slightly better since. I’m not recommending not to get the vaccine. You should get it. But this experience was weird enough to share. I had a fleeting thought of buying an eyepatch for privacy.
- andagainagain 6y agoI'm estimating roughly 90-ish characters in a row, roughly 40 rows encoding the spike protein. So about 3600 base pairs. There are 3 base pairs per amino acid, so That's 1200 amino acids. For comparison, the smallest chain that they technically call a protein is 100 amino acids that's an arbitrary limit to separate proteins from enzymes. So this thing isn't tiny tiny. But Titin (also called connectin), a giant protein responsible for passive elasticity in mucles, is ~27,000-35,000 amino acids. So this thing isn't even close to the biggest proteins out there.
- flobosg 6y ago> that's an arbitrary limit to separate proteins from enzymes Do you mean “to separate polypeptides from proteins”? Enzymatic activity has nothing to do with size. For example, one of the smallest enzymes in humans has 62 amino acid residues. And, under certain conditions, even single amino acids can be catalytic. But yeah, the polypeptide-protein threshold can get fuzzy, especially with the recent advances in miniprotein characterization.
- andagainagain 6y agoyes, that is what I meant. It's been a long time since I've used that info. The story I remember was that Insulin was the first protein that was sequenced, which is funny because it was before they made the distinction. It's actually too small to be considered a protein now.
- devenvdev 6y agoI remember a lot of features and especially bug fixes where I had to change one line of code, it took hours to figure out how exactly though. I guess this is kinda similar?
- pyinstallwoes 6y agoSee RadVac: https://radvac.org/ https://radvac.org/ Make your own, open-source. Really cool. A user on lesswrong made their own (with no prior experience): https://www.lesswrong.com/posts/niQ3heWwF6SydhS7R/making-vaccine https://www.lesswrong.com/posts/niQ3heWwF6SydhS7R/making-vac...
- weinzierl 6y ago> "Instead it's no bigger than a large paragraph of text, and I can easily fit it on my screen." When I saw it, I thought that it could almost fit in a tweet, so I just did it: https://twitter.com/weinzierl/status/1376807707957719041?s=21 https://twitter.com/weinzierl/status/1376807707957719041?s=2... The sequence takes 16 tweets, 15 if you don't split at line endings and remove spaces (4175 nucleobases / 280 nucleobases/tweet ~ 14.9 tweets).
- lifthrasiir 6y agoOr you can use base2048 [1] to compress it down to 3 tweets (4175 nucleobases * 2 bits per nucleobase / 3080 bits per base2048 tweet = 2.7 tweets). [1] https://github.com/qntm/base2048/ https://github.com/qntm/base2048/
- learnstats2 6y agoThe genetic code itself is reasonably comparable to ASCII in complexity - every 6 bits is the code for one amino acid in a string, which will fold itself into the required protein.
- phreeza 6y agoPeople tend to think of genetic code as a sort of assembly language which is very verbose, but I wonder if the correct way to view it is in fact a very terse domain-specific language, because it actually depends on the entire complex machinery of the cell to be present in order to work, which in itself contains a lot of information?
- Tuna-Fish 6y agoYes. Another important metaphor is that the common idea of DNA as blueprints is entirely wrong. It's not blueprints, it's a recipe. A blueprint describes what something is. A recipe describes the steps needed to make something, making use of a lot of complex existing machinery and parts with only a reference to them.
- wubbfindel 6y agoInteresting reasoning. But isn't it true to say that the "complex existing machinery and parts" which interprets the DNA was itself put together from instructions found in other DNA? I suppose that metaphors are rarely entirely comparable.
- Tuna-Fish 6y agoSome of that machinery and parts isn't directly represented by DNA. As an example, DNA codes some proteins that help extend cell walls, but those only work if you already have cell walls. If you have only the full DNA for a cell, and no other knowledge, you cannot build that cell out of that.
- mattkrause 6y agoThe nucleotide sequence is obviously important, but people also sometimes forget that DNA and RNA are real things with 3D structure too. That matters too: it’s as if builders make errors where the blueprint rolls up or pages stick together. The whole thing is absolutely fascinating and wild.
- jldugger 6y ago
- nraynaud 6y agothe way I see it we're just at the beginning, and we're mainly copy/pasting a lot of code, we understand some small parts, and generally in the teenage years of genome programming. I don't know how long it will be before we get a bit more serious with it, but geneticists have a big obstacle in their understanding, any change might needs a thousand strong lifelong population study to be understood. That's way crappier than dumping the assembly or only having the documentation in Chinese. I will add that moreover the developers might have been even more conservative in their code because they knew it was going for large scale deployment, they probably avoided the cutting edge as much as they could.
- fiftyfifty 6y agoThe New York Times published an article last year with the entire genome of the SARS-Cov-2 virus, with a breakdown of different sections to explain what protein the RNA codes for and what that protein does. Like you said it was amazing that it all fit within an [albeit long] newspaper article. It doesn't surprise me that the RNA for the vaccine, which only targets a single protein, is even smaller than that. Here's the NY Times article I was referring too: https://www.nytimes.com/interactive/2020/04/03/science/coronavirus-genome-bad-news-wrapped-in-protein.html https://www.nytimes.com/interactive/2020/04/03/science/coron...
- gremlinsinc 6y agoThe way it reads like source code, truly makes me circle back to the idea we're all living in a simulation.
- ohmyzee 6y agoBravo! Nice execution of the tips from yesterday's article! https://www.cs.purdue.edu/homes/dec/essay.criticize.html https://www.cs.purdue.edu/homes/dec/essay.criticize.html
- biolurker1 6y agoMathematical truths about abstract notions of string theory fit in a line.