3 ms·
> Conventional biostatistical analysis indicates that the probability of this sequence randomly being present in a 30,000-nucleotide viral genome is 3.21 ×10−11
by mizzack 5y ago
> Conventional biostatistical analysis indicates that the probability of this sequence randomly being present in a 30,000-nucleotide viral genome is 3.21 ×10−11
That seems like a pretty small number.
- fabian2k 5y agoI don't find that calculation convincing as it ignores that these viruses already have a furin cleavage site that by definition must be pretty similar to the sequence here. So the probablity would have to be calculated as "how likely is it that a virus with a related furin cleavage site accumulates the number of mutations necessary to arrive at this particular optimized one". And even then as this sequence provides a fitness advantage a naive calculation could be seriously off as well.
- XorNot 5y agoThe probability is also some argument by increduality anyway: 10-11, or you know, an occurrence rate of 10 after a trillion attempts. In an "average" COVID-19 infection course in an adult human, it is estimated that at peak infection a person has 10^9 to 10^11 virion particles in their body alone. Multiply by the all the people infected, + all the animals, + parallel gene transfer with other viruses in the ecosystem... [1] https://www.pnas.org/content/118/25/e2024815118 https://www.pnas.org/content/118/25/e2024815118
- PragmaticPulp 5y agoEach virus copy isn’t a different random sequence though. They have to share most of their code to work.
- XorNot 5y agoOf course, but if you're going to make arguments about probability in nature, they only mean something compared to the attempt space.
- uxp100 5y agoYou know, I really know nothing about this stuff, but this use of statistics (across many fields we see the odds of this occurring are X) kinda feels like nonsense. The odds of this occurring is 1. Because it did, and we found it occurred, and then went back and concocted these numbers. It’s like a poker hand you draw, and wow, the odds of getting this hand are miniscule. But the odds of getting a set of cards, if you draw them, is 1. Going back and finding the “odds” of an event that already happens is mostly meaningless. Extremely unlikely things happen constantly, see any particular game of poker. What I think would have been at stat that would be interesting is what are the odds of ANY patented sequence appearing in this genome.
- PragmaticPulp 5y agoIt’s misleading. Imagine this same analogy for computer programs. Instead of genetic code, considers bits. Some author without understanding of how code works finds a small matching bit sequence in two different executables, then implies that the shared sequence is evidence of some conspiracy. As a programmer you’d know that the bits of a program aren’t random because they actually do something specific for the programs execution, so it’s not surprising that different programs would share short bit sequences. But someone who doesn’t understand programming might assume it’s legitimate evidence of a conspiracy. Same thing is happening here. Genetic code isn’t random because it corresponds to specific functions. It’s not surprising that different viruses would end up sharing bits of genetic sequence if they have similar functions. The author is misusing statistics for random distributions on something that isn’t random. Don’t take them seriously.
- salawat 5y agoIn the world of StackOverflow driven software authorship, a particular sequence of source code instructions appearing in source code, on the presence of HTTP requests to that article, coupled with people doing work in that area? Yeah... I don't need statistics for that one. The bloody tooling speaks, and the chart of human agency/interest speaks for itself. To be honest, ai find it amazing the backflips people are doing to justify that there's no possible way that an interested grad student with the right tools would have tried something like this. I'd have. Of course I avoid doing science in those types of areas because Murphy finds a way.
- numpad0 5y agoI'm a total layperson, but looks like virus sheddings are measured as 10^x copies/ml in serum, and 10^5.5 copies is a typical amount for COVID-19. 10^6ml is 1kL, or 1m^3(3x3x3ft). Doesn't look all that small to me.
- contravariant 5y agoIt is, though I'm not certain it means much. In fact it's misstated since it's the probability that this exact sequences appears in 30,000 uniformly random base pairs and it appears in a library of 24712 sequences of sequences of 3300 uniformly random base pairs (see Figure 2). Which to me doesn't say much. It's like pointing out it's very unlikely it would have been typed by a monkey mashing on a typewriter. A somewhat more honest question would be to look at what they did, which is look at a length 12 sequence (and it's complement) that they apparently found interesting and see if it showed up in a library of 24712 sequences with an average length of 3300. You can estimate this probability by calculating the average number of matches you'd expect to find. This can be done by taking the number of positions a match could be in, approximately 24712*3300 - 24712*11, times 2 to account for its complement showing up, and multiply this by the chance two length 12 sequences will match, estimated by them at (1/4)^12 (in reality this probability should probably be higher since base pairs aren't uniformly random). Since this is an expected value we can ignore that matches at adjacent positions aren't independent. This suggests you'd expect to find about 10 matches. So I'm not sure if they simply picked the longest one or if there was just the 1 match. You do get a lowish number of 0.000059 of matching sequences of length 19, that said this chance would be significantly higher when you know there exists a match of length 12, especially if the sequences aren't simply uniformly random.