11 ms·
AlphaGenome: AI for better understanding the genome
- dekhn 1y agoWhen I went to work at Google in 2008 I immediately advocated for spending significant resources on the biological sciences (this was well before DM started working on biology). I reasoned that Google had the data mangling and ML capabilities required to demonstrate world-leading results (and hopefully guide the way so other biologists could reproduce their techniques). We made some progress- we used exacycle to demonstrate some exciting results in protein folding and design, and later launched Cloud Genomics to store and process large datasets for analytics. I parted ways with Google a while ago (sundar is a really uninspiring leader), and was never able to transfer into DeepMind, but I have to say that they are executing on my goals far better than I ever could have. It's nice to see ideas that I had germinating for decades finally playing out, and I hope these advances lead to great discoveries in biology. It will take some time for the community to absorb this most recent work. I skimmed the paper and it's a monster, there's just so much going on.
- bitpush 1y ago[flagged]
- pinoy420 1y ago[dead]
- CGMthrowaway 1y agoYeah it comes off as braggy, but it’s only natural to be proud of your foresight
- deleted 1y ago[deleted]
- liamwire 1y agoNatural? Sure. Deserved? Not really, not unless we’re also forthcoming in our lack of foresight and the times we plainly got it wrong.
- shadowgovt 1y agoFWIW, I interpreted more as "This is something I wanted to see happen, and I'm glad to see it happening even if I'm not involved in it."
- plemer 1y agoCould be either. Nevertheless, while tone is tricky in text, the writer is responsible for relieving ambiguity.
- spongebobstoes 1y agoeliminating ambiguity is impossible. the reader should work to find the strongest interpretation of the writer's words
- coderatlarge 1y agothat’s a lot to expect of readers… good writing needs to give readers every opportunity to find the good in it.
- shadowgovt 1y agoIt is a lot to expect of readers... It's also explicitly asked of us in this forum. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html. "Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize. Assume good faith."
- plemer 1y agofair point
- coderatlarge 1y agoit’s fine for a forum to try to have different expectations than the local cafe - that’s kind of like a host asking their guests to remove shoes before walking into their home. but it doesn’t really change a priori basic facts about good writing. perhaps this is the appropriate forum to reference pg https://paulgraham.com/writing44.html https://paulgraham.com/writing44.html https://paulgraham.com/essay.htm https://paulgraham.com/essay.htm
- dvaun 1y agoA charitable view is that they intended "ideas that I had germinating for decades" to be from their own perspective, and not necessarily spurred inside Google by their initiative. I think that what they stated prior to this conflated the two, so it may come across as bragging. I don't think they were trying to brag.
- alfanick 1y agoI don't find it rude or pretentious. Sometimes it's really hard to express yourself in hmm acceptable neutral way when you worked on truly cool stuff. It may look like bragging, but that's probably not the intention. I often face this myself, especially when talking to non-tech people - how the heck do I explain what I work on without giving a primer on computer science!? Often "whenever you visit any website, it eventually uses my code" is good enough answer (worked on aws ec2 hypervisor, and well, whenever you visit any website, some dependency of it eventually hits aws ec2)
- camjw 1y ago100% but in this case they uh… didn’t work on it, it seems?
- varelse 1y ago[dead]
- project2501a 1y agoFrom Marx to Zizek to Fukuyama^1, 200 years of Leftist thinking nobody has ever came close to say "we can fix capitalism". What makes you think that LLMs can do it? [1] relapsed capitalist, at best, check the recent Doomscroll interview
- deepdarkforest 1y ago> Sundar is a really uninspiring leader I understand, but he made google a cash machine. Last quarter BEFORE he was CEO in 2015, google made a quarterly profit of around 3B. Q1 2025 was 35B. a 10x profit growth at this scale well, its unprecedented, the numbers are inspiring themselves, that's his job. He made mistakes sure, but he stuck to google's big gun, ads, and it paid off. The transition to AI started late but gemini is super competitive overall. Deepmind has been doing great as well. Sundar is not a hypeman like Sam or Cook, but he delivers. He is very underrated imo.
- modeless 1y agoLike Ballmer, he was set up for success by his predecessor(s), and didn't derail strong growth in existing businesses but made huge fumbles elsewhere. The question is, who is Google's Satya Nadella? Demis?
- bitpush 1y agoSince we're on the topic of Microsoft, I'm sure you'd agree that Satya has done a phenomenal job. If you look objectively, what is Satya's accomplishments? One word - Azure. Azure is #2, behind AWS because Satya's effective and strategic decisions. But that's it. The "vibes" for Microsoft has changed, but MS hasnt innovated at all. Satya looked like a genius last year with OpenAI partnership, but it is becoming increasingly clear that MS has no strategy. Nobody is using Github Copilot (pioneer) or MS Copilot (a joke). They dont have any foundational models, nor a consumer product. Bing is still.. bing, and has barely gained any market share.
- modeless 1y agoMicrosoft has become a lot more friendly to open source under Satya. VSCode, GitHub, and WSL happened during his tenure, and probably wouldn't have happened under Ballmer. Turning the ship from a focus on protecting platform lock-in to meeting developers where they are is a huge accomplishment IMO.
- 1y ago
- spankalee 1y agoDid you ride the Santa Cruz shuttle, by any chance? We might have had conversations about this a long while ago. It sounded so exciting then, and still does with AlphaGenome.
- VirusNewbie 1y agoGoogler here ---^ I have incredibly mixed feelings on Sundar. Where I can give him credit is really investing in AI early on, even if they were late to productize it, they were not late to invest in the infra and tooling to capitalize on it. I also think people are giving maybe a little too much credit to Demis and not enough to Jeff Dean for the massive amount of AI progress they've made.
- semiinfinitely 1y agoNice wow 20% of the credit goes to you for thinking of this years ago. Kudos
- eleveriven 1y agoIt's easy to forget how early some of these ideas were being pushed internally
- nextos 1y agoI found it disappointing that they ignored one of the biggest problems in the field, i.e. distinguishing between causal and non-causal variants among highly correlated DNA loci. In genetics jargon, this is called fine mapping. Perhaps, this is something for the next version, but it is really important to design effective drugs that target key regulatory regions. One interesting example of such a problem and why it is important to solve it was recently published in Nature and has led to interesting drug candidates for modulating macrophage function in autoimmunity: https://www.nature.com/articles/s41586-024-07501-1 https://www.nature.com/articles/s41586-024-07501-1
- rattlesnakedave 1y agoDoes this get us closer? Pretty uninformed but seems that better functional predictions make it easier to pick out which variants actually matter versus the ones just along for the ride. Step 2 probably is integrating this with proper statistical fine mapping methods?
- nextos 1y agoYes, but it's not dramatically different from what is out there already. There is a concerning gap between prediction and causality. In problems, like this one, where lots of variables are highly correlated, prediction methods that only have an implicit notion of causality don't perform well. Right now, SOTA seems to use huge population data to infer causality within each linkage block of interest in the genome. These types of methods are quite close to Pearl's notion of causal graphs.
- ejstronge 1y ago> SOTA seems to use huge population data to infer causality within each linkage block of interest in the genome. This has existed for at least a decade, maybe two. > There is a concerning gap between prediction and causality. Which can be bridged with protein prediction (alphafold) and non-coding regulatory predictions (alphagenome) amongst all the other tools that exist. What is it that does not exist that you "found it disappointing that they ignored"?
- Scaevolus 1y agoNaturally, the (AI-generated?) hero image doesn't properly render the major and minor grooves. :-)
- solarwindy 1y agoFor anyone wondering: https://www.mun.ca/biology/scarr/MGA2_02-07.html https://www.mun.ca/biology/scarr/MGA2_02-07.html
- jeffbee 1y agoAnd yet still manages to be 4MB over the wire.
- smokel 1y agoThat's only on high-resolution screens. On lower resolution screens it can go as low as 178,820 bytes. Amazing.
- nh23423fefe 1y agowhen a human does it, its style! when ai does it, you cry about your job.
- AntiqueFig 1y agoMaybe they were depicting RNA? (probably not)
- dekhn 1y agoNo; what they drew doesn't look like real DNA or (duplex double stranded) RNA. Both have differently sized/spaced grooves (see https://www.researchgate.net/profile/Matthew-Dunn-11/publication/286936714/figure/fig3/AS:307313886941186@1450280748491/Geometric-parameters-of-natural-A-and-B-form-helices-A-Natural-genetic-polymers.png https://www.researchgate.net/profile/Matthew-Dunn-11/publica...). At least they got the handedness right.
- jeffhwang 1y agoWhen I was restudying biology a few years ago, it was making me a little crazy trying to understand the structural geometry that gives rise to the major and minor grooves of DNA. I looked through several of the standard textbooks and relevant papers. I certainly didn't find any good diagrams or animations. So out of my own frustration, I drew this. It's a cross-section of a single base pair, as if you are looking straight down the double helix. Aka, picture a double-strand of DNA as an earthworm. If one of the earthworms segments is a base-pair, and you cut the earthworm in half, and turn it 90 degrees, and look into the body of the worm, you'd see this cross-sectional perspective. Apologies for overly detailed explanation; it's for non-bio and non-chem people. :) https://www.instagram.com/p/CWSH5qslm27/ https://www.instagram.com/p/CWSH5qslm27/ Anyway, I think the way base pairs bond forces this major and minor grove structure observed in B-DNA.
- mountainriver 1y agoWith the huge jump in RNA prediction seems like it could be a boon for the wave of mRNA labs
- iandanforth 1y agoThose outside the US at least ...
- TechDebtDevin 1y agoI've been saying we need a rebranding of mRNA in the USA its coming.
- divbzero 1y ago“in situ therapeutics”
- seydor 1y agothis is such an interesting problem. Imagine expanding the input size to 3.2Gbp, the size of human genome. I wonder if previously unimaginable interactions would occur. Also interesting how everything revolves around U-nets and transformers these days.
- teaearlgraycold 1y ago> Also interesting how everything revolves around U-nets and transformers these days. To a man with a hammer…
- SV_BubbleTime 1y agoSoon we’ll be able to get the whole genome up on the blockchain. (I thought the /s was obvious)
- TeMPOraL 1y agoOr to a man with a wheel and some magnets and copper wire... There are technologies applicable broadly, across all business segments. Heat engines. Electricity. Liquid fuels. Gears. Glass. Plastics. Digital computers. And yes, transformers.
- pfisherman 1y agoYou would not need much more than 2 megabases. The genome is not one contiguous sequence. It is organized (physically segregated) into chromosomes and topologically associated domains. IIRC 2 megabases is like the 3 sd threshold for interactions between cis regulatory elements / variants and their effector genes.
- eleveriven 1y agoEven just modeling 3D genome organization or ultra-long-range enhancers more realistically could open up new insights
- jebarker 1y agoI don't think DM is the only lab doing high-impact AI applications research, but they really seem to punch above their weight in it. Why is that or is it just that they have better technical marketing for their work?
- 331c8c71 1y agoThis one seems like well done research but in no way revolutionary. People have been doing similar stuff for a while...
- Gethsemane 1y agoAgreed, there’s been some interesting developments in this space recently (e.g. AgroNT). Very excited for it, particularly as genome sequencing gets cheaper and cheaper! I’d pitch this paper as a very solid demonstration of the approach, and im sure it will lead to some pretty rapid developments (similar to what Rosettafold/alphafold did)
- nextos 1y agoIn biology, Arc Institute is doing great novel things. Some pharmas like Genentech or GSK also have excellent AI groups.
- 331c8c71 1y agoArc have just released a perturbation model btw. If it reliably beats linear benchmarks as claimed it is a big step https://arcinstitute.org/news/virtual-cell-model-state https://arcinstitute.org/news/virtual-cell-model-state
- tim333 1y agoThey have been at it for a long time and have a lot of resources courtesy of Google. Asking perplexity it says the alphafold 2 database took "several million GPU hours".
- kridsdale3 1y ago
- twothreeone 1y agoMaybe "Release" requires a bit more context, as it clearly means different things to different people: > AlphaGenome will be available for non-commercial use via an online API at http://deepmind.google.com/science/alphagenome http://deepmind.google.com/science/alphagenome So, essentially the paper is a sales pitch for a new Google service.
- RivieraKid 1y agoI wish there's some breakthrough in cell simulation that would allow us to create simulations that are similarly useful to molecular dynamics but feasible on modern supercomputers. Not being able to see what's happening inside cells seems like the main blocker to biological research.
- m3kw9 1y agoI believe this is where quantum computing comes in but could be a decade out, but AI acceleration is hard to predict
- andrewchoi 1y agoThe folks at Arc are trying to build this! https://arcinstitute.org/news/virtual-cell-model-state https://arcinstitute.org/news/virtual-cell-model-state
- dekhn 1y agoSTATE is not a simulation. It's a trained graphical model that does property prediction as a result of a perturbation. There is no physical model of a cell. Personally, I think arc's approach is more likely to produce usable scientific results in a reasonable amount of time. You would have to make a very coarse model of the cell to get any reasonable amount of sampling and you would probably spend huge amounts of time computing things which are not relevant to the properties you care amount. An embedding and graphical model seems well-suited to problems like this, as long as the underlying data is representative and comprehensive.
- noduerme 1y agoI wish there were more interest in general in building true deterministic simulations than black boxes that hallucinate and can't show their work.
- t_serpico 1y ago'Seeing' inside cells/tissues/organs/organisms is pretty much most modern biological research.
- 1y ago
- lcfcjs6 1y ago[dead]
- cwmoore 1y agoJust add startofficial intel.
- mattigames 1y agoI bet the internal pitch is that genome will help deliver better advertisement, like if you are at risk of colon cancer they sell you "colon supplements", its likely they will be able to infer a bit about your personality just with your genome, "these genes are correlated with liking dark humor, use them to promote our new movie"
- LarsDu88 1y agoYou know the corporate screws are coming down hard, when the model (which can be run off a single A100) doesn't get a code release or a weight release, but instead sits behind an API, and the authors say fuck it and copy-paste the entirety of the model code in pseudocode on page 31 of the white paper. Please Google/Demis/Sergei, just release the darn weights. This thing ain't gonna be curing cancer sitting behind an API and it's not gonna generate that much GCloud revenue when the model is this tiny.
- twothreeone 1y agoThe bean counters rule. There is no corporate vision, no long-term plan. The numbers for the next quarter are driving everything.
- hotstickyballs 1y agoI can guarantee you that some smart person actually thinks that the opportunity size is measured as a fraction of the pharmaceutical industry market cap.
- Onawa 1y agoDoesn't matter if that person can't convince the bean counters and lawyers that it is in the company's long-term interest to release the "proprietary" data.
- wrsh07 1y agoThis is a strange take because this is consistent with what Google has been doing for a decade with AI. AlphaGo never had the weights released. Nor has any successor (not muzero, the StarCraft one, the protein folding alphafold, nor any other that could reasonably be claimed to be in the series afaik) You can state as a philosophical ideal that you prefer open source or open weights, but that's not something deepmind has prioritized ever. I think it's worth discussing: * What are the advantages or disadvantages of bestowing a select few with access? * What about having an API that can be called by anyone (although they may ban you)? * Vs finally releasing the weights But I think "behind locked down API where they can monitor usage" makes sense from many perspectives. It gives them more insight into how people use it (are there things people want to do that it fails at?), and it potentially gives them additional training data
- hyfgfh 1y agoCant await for people to use it for CRISPR an it hallucinate some weird mutation
- leoh 1y agoLet’s figure out introns pls
- xipho 1y ago"To ensure consistent data interpretation and enable robust aggregation across experiments, metadata were standardized using established ontologies." Can't emphasize enough about how DNA requires human data curation to make things work, even from day one alignments models were driven based on biological observations. Glad to see UBERON, which represents a massive amount of human insight and data curation of what is for all intents and purposes a semantic-web product (OWL based RDF at the heart) playing a significant role.
- deleted 1y ago[deleted]
- kat529770 1y ago[dead]
- eleveriven 1y agoCurious how it'll perform when people start fine-tuning on smaller, specialized datasets
- sahil_sharma0 1y ago[dead]
- another_twist 1y agoSo very similar approach to Conformer - convolution head for downsampling and transformer for time dependencies. Hmm, surprising that this idea works across application domains.
- i_am_not_groot 1y ago[flagged]
- Kalanos 1y agoThe functional predictions related to "non-coding" variants are big here. Non-coding regions, referred to as the dark genome, produce regulatory non-coding RNA's that determine the level of gene expression in a given cell type. There are more regulatory RNA's than there are genes. Something like 75% of expression by volume is ncRNA.
- wespiser_2018 1y agoIt's possible that the "functional" aspect of non-coding RNA exists on a time scale much larger that what we can assay in a lab. The sort of "junk DNA/RNA" hypothesis: the ncRNA part of the genome is material that increases fitness during relative rare events where it's repurposed into something else. On a millions or billions of year time frame, the organisms with the flexibility of ncRNA would have an advantage, but this is extremely hard to figure out with a "single point in time" view point. Anyway, that was the basic lesson I took from studying non-coding RNA 10 years ago. Projects like ENCODE definitely helped, but they really just exposed transcription of elements that are noisy, without providing the evidence that any of it is actually "functional". Therefore, I'm skeptical that more of the same approach will be helpful, but I'd be pleasantly surprised if wrong.
- cysteinechapel 1y agoSuch an advantage that is rare and across such long time scales would be so small on average that it would be effectively neutral. Natural selection can only really act on fitness advantages greater than on the order of the inverse of effective population size, which for large multicellular organisms such as animals, is low. Most of this is really just noisy transcription/binding/etc. For example, we don't keep transposons in general because they're useful, which are almost half of our genomes, and are a major source of disruptive variation. They persist because we're just not very good at preventing them from spreading, we have some suppressive mechanisms but they don't work all the time, and there's a bit of an arms race between transposons and host. Nonetheless, they can occasionally provide variation that is beneficial.
- dekhn 1y agoThere is a big long-running argument about what "functional" means in "non-coding" parts of the genome. The deeper I pushed into learning about the debate the less confident I became of my own understanding of genomics and evolution. See https://www.sciencedirect.com/science/article/pii/S0960982213002893 https://www.sciencedirect.com/science/article/pii/S096098221... for one perspective.
- insane_dreamer 1y agoThese are the type of advances in AI models that I'm excited about because of their potentially beneficial high impact for mankind. Not models that are a better (but less reliable) search engine or coding assistant and email writer. I wish more effort/money was going into this.
- kylehotchkiss 1y agoI'm somewhat a noob here, but does this model have good understanding of things like OvRFs, methylation, etc, or is it strictly a sequence pattern matching thingy?
- qoez 1y agoDemis to be the first to get 4 consecutive nobels
- b0a04gl 1y ago[dead]
- richardvc 1y agoUnderstanding the genome has always felt like trying to solve a massive puzzle with pieces constantly shifting. Tools like AlphaGenome are changing that—offering a more focused way to interpret complex genetic data. In the lab I worked in, precision was everything. We relied heavily on uv spectrophotometry for DNA quantification, and the systems from https://www.berthold.com/en/ https://www.berthold.com/en/ stood out for their consistency, even under demanding conditions. Their devices helped streamline processes where accuracy couldn’t be compromised, especially when dealing with fragile or low-concentration samples. Founded back in 1949, they’ve become a global reference for reliable measuring technology. From radiation detection to life sciences and industrial process control, their solutions cover diverse fields. For anyone navigating genomics or analytical research, choosing the right tools isn’t just about features—it’s about long-term dependability and clarity in results.