For families living with an undiagnosed genetic condition, the hardest part is often not sequencing a genome. It is deciding which of millions of differences deserves a closer look. Google DeepMind’s AlphaGenome Atlas is an ambitious attempt to change that triage problem: a searchable, 1-petabyte map of predicted molecular consequences for every possible single-letter DNA change in the human genome. The promise is not a machine that diagnoses disease. It is a faster route from an overwhelming list of variants to the few experiments most worth doing.
Google DeepMind said in its September 8 announcement that AlphaGenome Atlas contains predictions for roughly 9 billion possible single-nucleotide variants, meaning substitutions of one DNA letter for another across the human genome. The resource pairs those predictions with a single ranking number, the AlphaGenome Variant Impact score, or AVI, plus feature attributions intended to explain which predicted biological effects drove the score.
That scale matters because modern genetics has an abundance problem. A human genome can be read relatively cheaply compared with even a decade ago, but the result is still a vast list of differences. Most are harmless. Some are biologically meaningful but unrelated to a person’s condition. A very small number may change how a gene works in a relevant tissue at a relevant time.
The difficult cases often sit outside genes themselves. Only about 2% of the genome directly codes for proteins. The remaining 98% includes regulatory DNA that helps decide where, when and how strongly genes are used. A change in that regulatory landscape may matter even if it leaves the protein sequence untouched.
Why one DNA letter needs a million-letter view
“Laura Covill Broad Institute / GREGoR Consortium”
DNA is often described as an instruction manual, but its regulatory logic is less like a list of isolated sentences than a city’s planning system. A small edit to a distant zoning rule can affect activity several blocks away. A regulatory sequence may influence a gene from far from the gene’s protein-coding portion, and the effect can vary by cell type.
That is why the base model behind the Atlas accepts DNA sequences up to 1 million letters long. It is designed to retain broad genomic context while producing predictions at individual-letter resolution. Google DeepMind’s original AlphaGenome introduction explained that the model predicts thousands of molecular properties, including RNA production, splicing, DNA accessibility, protein binding and which DNA regions come into physical proximity.
The apparent contradiction is the engineering challenge. Long context is necessary to capture distant regulatory relationships. Fine resolution is necessary because one altered base can create or disrupt a binding site, shift an RNA splice junction or change a local accessibility signal. A system that sees only a short sequence can miss distant influences. One that sees a long sequence but reduces it too aggressively can lose the precise mutation-level signal.
AlphaGenome addresses this through a layered sequence model. Convolutional layers first identify short recurring patterns, such as recognizable local DNA motifs. Transformer layers then communicate information across positions in the sequence, allowing distant elements to influence one another. Final output layers translate the combined representation into different types of biological prediction. DeepMind says training distributes this computation over interconnected Tensor Processing Units for a single sequence.
A useful analogy is reading a legal contract. Convolutions identify familiar phrases. The transformer keeps track of how a definition on page one modifies a clause on page eighty. The output heads then answer distinct questions: what does the contract permit, what does it require, and where are the terms most likely to matter?
One input, thousands of biological readouts
The Atlas is not merely a list of labels such as “harmful” or “benign.” The underlying model produces a broad set of predicted molecular tracks across human and mouse cell types and tissues. These tracks describe different layers of gene regulation, from whether a region of DNA is accessible to the cellular machinery, to RNA abundance, splice behavior, protein binding and chromatin-related signals.
That multimodal design matters because regulatory biology is a chain of events. A mutation may change the accessibility of a DNA region, reduce a protein’s ability to bind there, alter transcription, then change RNA processing. Looking at only one stage can obscure the mechanism. Looking across several gives researchers a hypothesis that can be tested.
Reference and alternate sequences can be compared through differences in their predicted tracks
The model’s predictions are also comparative. To score a variant, AlphaGenome runs the reference sequence and a sequence containing the alternate DNA letter, then measures the difference between their predicted molecular tracks. That turns a genetic change into a before-and-after experiment conducted in silico.
For a variant near a splice site, for example, the important question is not simply whether the local sequence looks unusual. It is whether the altered sequence changes the model’s predicted location or strength of an RNA junction. DeepMind highlights explicit splice-junction modeling as a distinctive capability, which is important because splicing errors can contribute to rare genetic disorders.
From model run to atlas
Running such comparisons for a handful of variants is useful. Running them for every possible single-letter substitution in the human genome is a different kind of project. Each DNA position can, in principle, change to three other letters. Across billions of possible changes, and across thousands of output tracks, the raw prediction volume grows quickly.
DeepMind describes the resulting Atlas as a 1-petabyte dataset, more than 30 times the size of the AlphaFold Database according to the company. It precomputes molecular effects genome-wide, then associates each variant with the AVI score, feature attributions and a resource of more than 2,500 recurring DNA sequence motifs.
A petabyte is not just an impressive unit. It explains why the product must be treated as an information-retrieval system rather than a file researchers download and inspect on a laptop. The practical workflow is to query a genomic location or candidate variant, retrieve the precomputed predictions and rankings, inspect the implicated molecular features, then export a manageable shortlist for follow-up.
The AVI score is the system’s compression layer. DeepMind says it combines AlphaGenome’s regulatory predictions with AlphaMissense, a model focused on protein-altering variants, into one number that can rank changes in both coding and non-coding DNA. Feature attribution is the companion to that ranking: rather than merely assigning a high score, the Atlas identifies which predicted categories, such as RNA splicing, gene expression, chromatin accessibility or protein impact, contributed to it.
That is valuable because a ranking without a mechanism creates another bottleneck. A clinician or researcher cannot build a useful experiment from “variant 17 is important.” They need a claim that can be disproved: this change may create a splice site, weaken regulatory activity in a specific cell type or alter a protein-coding sequence.
The real figures: scale and reported research examples
The project’s numerical claims help explain both its attraction and its limits. The clearest are its coverage: 9 billion possible variants, thousands of predicted effects per variant, a 1-petabyte catalogue, and more than 2,500 DNA motifs. DeepMind also says an external collaborator’s analysis of more than 54,000 UK Biobank participants found 22% more non-coding associations after grouping rare variants by predicted molecular effects.
The reported examples are encouraging precisely because they connect a score to subsequent biology. DeepMind says collaborators working on unsolved rare disease used AVI to surface a DNM1 variant linked to epileptic encephalopathy. The model predicted an incorrect splice site and an abnormal protein extension, then experimental screens reportedly validated that hypothesis and found nearby variants with similar effects.
Still, this is where language matters. A model prediction is evidence for prioritisation, not proof of disease causality. A high AVI score does not establish that a variant causes a patient’s symptoms. It does not replace segregation analysis in families, clinical context, population-frequency checks, functional experiments or medical judgment.
A validation workflow, not a diagnosis engine
The most responsible use of the Atlas is a staged workflow. First, begin with variants that fit the patient, trait or biological question. Second, use AVI to rank candidates, including non-coding candidates that conventional protein-focused pipelines may overlook. Third, inspect the attributed molecular tracks and relevant cell types. Fourth, ask whether the mechanism aligns with what is already known about the gene and phenotype. Finally, test the best hypothesis experimentally.
That experiment might be a reporter assay for regulatory activity, an RNA assay to measure splicing, CRISPR editing in a relevant cell model, or targeted work in patient-derived cells. The experiment is not a ceremonial final step. It is where a plausible computational explanation becomes biological evidence.
DeepMind itself states that AlphaGenome Atlas is intended for research and is not validated or approved for clinical use. The distinction should shape how institutions adopt it. The Atlas can help make experimental biology more selective and efficient. It cannot convert a prediction into a diagnosis by itself.
The larger shift is cultural as much as technical. Genetics has spent years learning how to read genomes. The next challenge is learning how to navigate them. AlphaGenome Atlas offers an unusually large map, with rankings and suggested landmarks. Its impact will depend on whether scientists use that map as a compass for better experiments, rather than mistaking it for the territory.
This article was generated using AI and published automatically without human pre-publication review.
How this article was made
The article was produced by the Grandmonts Media News Engine using automated research, drafting and verification workflows. No human editor reviewed the article before publication. Grandmonts Media remains responsible for the published content. Errors can be reported at office@grandmonts.cz.