PAPER KEY: T658KG44
TITLE: Predicting expression-altering promoter mutations with deep learning
AUTHORS: Jaganathan, Kishore; Ersaro, Nicole; Novakovsky, Gherman; Wang, Yuchuan; James, Terena; Schwartzentruber, Jeremy; Fiziev, Petko; Kassam, Irfahan; Cao, Fan; Hawe, Johann; Cavanagh, Henry; Lim, Ashley; Png, Grace; McRae, Jeremy; Banerjee, Abhimanyu; Kumar, Arvind; Ulirsch, Jacob; Zhang, Yan; Aguet, Francois; Wainschtein, Pierrick; Sundaram, Laksshman; Salcedo, Adriana; Kyriazopoulou Panagiotopoulou, Sofia; Aghamirzaie, Delasa; Padhi, Evin; Weng, Ziming; Dong, Shan; Smedley, Damian; Caulfield, Mark; O’Donnell-Luria, Anne; Rehm, Heidi L.; Sanders, Stephan J.; Kundaje, Anshul; Montgomery, Stephen B.; Ross, Mark T.; Farh, Kyle Kai-How

Cite as: K. Jaganathan et al., Science 10.1126/science.ads7373 (2025).
RESEARCH ARTICLES
First release: 29 May 2025 science.org (Page numbers not final at time of first release) 1
The precise control of gene expression is broadly important across human health and development, but the mechanisms by which genomic sequence encodes these intricate programs remain incompletely understood. Central to gene regulation is the promoter, the site of transcriptional initiation, which integrates signals across multiple non-coding sequence elements to determine the proper cellular and temporal context for turning genes on and off. Experimental studies of promoters have shown that they can dramatically increase or decrease gene expression (1), suggesting that non-coding variants which fall within the promoters of clinically relevant genes may play key roles in rare genetic disorders and cancer (2, 3). However, clinical interest in promoter variants has been limited, due to the difficulty of distinguishing between functional non-coding variants that affect gene expression and those that are neutral, with only a handful of pathogenic non-coding variants in promoters having been identified to date (4–12). These challenges in finding pathogenic variants that lie outside protein-coding sequence continue to be a
major hindrance to realizing the potential of personalized genome sequencing. Deep learning models, known for their efficacy at recognizing patterns from large volumes of unstructured data, offer a promising path forward for distilling data collected from genome-wide sequencing and functional genomics experiments into models that can accurately predict the clinical impact of human genetic variation (13, 14). Recent examples that have seen adoption in clinical contexts include SpliceAI (15), which has been cited in expert guidelines for splice variant effect prediction in rare genetic disorders and oncology (16), and missense variant prediction based on protein language models and 3D crystal structures (PrimateAI-3D) (17). While deep learning models such as Evo2 (18), DeepSEA (19), Basenji (20), ExPecto (21), Enformer (22), PuffinD (23), ChromBPNet (24), and Borzoi (25) have been developed to infer the regulatory code directly from genomic sequence, predicting the effects of non-coding genetic variants remains an unmet challenge (26, 27).
Predicting expression-altering promoter mutations with
deep learning
Kishore Jaganathan1†, Nicole Ersaro1†, Gherman Novakovsky1†, Yuchuan Wang1, Terena James1, Jeremy Schwartzentruber1, Petko Fiziev1, Irfahan Kassam1, Fan Cao1, Johann Hawe1, Henry Cavanagh1, Ashley Lim1, Grace Png1, Jeremy McRae1, Abhimanyu Banerjee1, Arvind Kumar1, Jacob Ulirsch1, Yan Zhang1, Francois Aguet1, Pierrick Wainschtein1, Laksshman Sundaram1, Adriana Salcedo1, Sofia Kyriazopoulou Panagiotopoulou1‡, Delasa Aghamirzaie1§, Evin Padhi2, Ziming Weng2, Shan Dong3,4, Damian Smedley5, Mark Caulfield5, Anne O’Donnell-Luria6,7,8, Heidi L. Rehm6,7, Stephan J. Sanders3,4,9, Anshul Kundaje10,11, Stephen B. Montgomery2,10,12, Mark T. Ross1, Kyle Kai-How Farh1*
1Illumina Artificial Intelligence Laboratory, Illumina, Inc., San Diego, CA, USA. 2Department of Pathology, Stanford University, Stanford, CA, USA. 3Department of Psychiatry and Behavioral Sciences, UCSF Weill Institute for Neurosciences, University of California San Francisco, San Francisco, CA, USA. 4Institute of Developmental and Regenerative Medicine, Department of Paediatrics, University of Oxford, Oxford, UK. 5William Harvey Research Institute, Queen Mary University of London, London, UK. 6Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA. 7Center for Genomic Medicine and Analytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA. 8Division of Genetics and Genomics, Boston Children’s Hospital, Harvard Medical School, Boston, MA, USA. 9New York Genome Center, New York, NY, USA. 10Department of Genetics, Stanford University, Stanford, CA, USA. 11Department of Computer Science, Stanford University, Stanford, CA, USA. 12Department of Biomedical Data Science, Stanford University, Stanford, CA, USA.
†These authors contributed equally to this work. ‡Present address: Arsenal Biosciences, Inc., South San Francisco, CA, USA. §Present address: DELFI Diagnostics, Inc., Palo Alto, CA, USA.
*Corresponding author. Email: kfarh@illumina.com
Only a minority of patients with rare genetic diseases are currently diagnosed by exome sequencing, suggesting that additional unrecognized pathogenic variants may reside in non-coding sequence. Here, we describe PromoterAI, a deep neural network that accurately identifies non-coding promoter variants which dysregulate gene expression. We show that promoter variants with predicted expression-altering consequences produce outlier expression at both RNA and protein levels in thousands of individuals, and that these variants experience strong negative selection in human populations. We observe that clinically relevant genes in rare disease patients are enriched for such variants and validate their functional impact through reporter assays. Our estimates suggest that promoter variation accounts for 6% of the genetic burden associated with rare diseases.
Downloaded from https://www.science.org at Cold Spring Harbor Laboratory on August 06, 2025


First release: 29 May 2025 science.org (Page numbers not final at time of first release) 2
Results
PromoterAI predicts the effects of promoter variants on gene expression
We introduce PromoterAI, a convolutional deep neural network model that uses ~20 kb of sequence context around a promoter variant to predict its expression consequence. We first trained the model to predict histone modifications, DNA accessibility, transcription factor (TF) occupancy, and strandspecific CAGE [Cap Analysis of Gene Expression (28, 29)] values around transcription start sites (TSS) at base-pair resolution. We subsequently fine-tuned the model using a curated set of rare non-coding promoter variants associated with unusually high or low gene expression to further improve its accuracy across extensive validation benchmarks, which we publish as a resource (Fig. 1A, fig. S1A, and table S1, A to G). To create this curated list of variants for fine-tuning, we analyzed the Genotype-Tissue Expression v8 (GTEx) cohort, consisting of paired whole genome sequencing (WGS) and RNA-seq data across 49 tissues in 838 individuals (30). Specifically, we cataloged a set of rare (allele frequency < 0.5%) variants in the promoters of genes (TSS +/− 500 bp) with outlier expression across multiple tissues. These multi-tissue outliers were identified via a t-test between expression zscores from all measured tissues in the carriers versus the same tissues in non-carriers. To maximize confidence in the resulting outliers, when generating the expression z-scores, we subtracted out the contribution of other factors that also affect gene expression: principal components from the expression and genotype matrices, cis-effects of common expression quantitative trait loci (eQTLs) on gene expression, and trans-effects of correlated patterns of gene expression (Fig. 1B). To perform the trans-expression correction, we predicted each gene’s expression using the expression of all genes residing on other chromosomes and then subtracted the predicted expression values from the observed values, since these generally reflect global patterns unrelated to the effects of promoter sequence. To assess the gain in specificity after each correction step, we compared the number of observed expression outliers to background expectation, which we estimated by shuffling the cohort with variants being randomly re-assigned between individuals. After initially correcting for genotype and expression principal components, we found 2,540 multi-tissue outliers at a p-value threshold that produced 1,000 outliers in the shuffled controls; this increased to 3,116 after correcting for cis-eQTLs, and finally to 4,030 following the trans-expression correction (Fig. 1B and table S2). Individuals whose observed gene expression deviated from their predicted gene expression were enriched up to 6-fold for carrying a rare promoter variant in that gene for under-expression, and 2.5-fold for overexpression (Fig. 1C). To identify the extent of the region surrounding the TSS where we can robustly identify multi-tissue expression
outliers, we calculated enrichments of variants with significant effects on gene expression (t-test, p < 1e-4) within 100 bp sliding windows at distances up to 5 kb from the TSS (Fig. 1D). This analysis revealed that expression outliers were predominantly located within 500 bp upstream and downstream of the TSS. While these enrichments improved when the analysis was restricted to conserved promoters (fig. S1B), selecting the most frequently used TSS across tissues was critical for maximizing enrichments (fig. S1C and table S3). We benchmarked our approach against other existing methods to quantify aberrant expression, including an unsupervised autoencoder model from OUTRIDER (31), and two summary statistics available from GTEx (32): median z-scores across tissues and median p-values from an allelic imbalance test (ANEVA-DOT) (33). Compared to the other methods, when choosing thresholds that result in the same false discovery rate in shuffled controls, our approach detected greater enrichment of multi-tissue outliers (Fig. 1E). Using allele-specific expression to assess variants with phasing data available, we found that alleles carrying overexpression outliers had significantly higher allelic expression compared to alleles carrying under-expression outliers (p = 7.4e-29; fig. S1D). As a further orthogonal validation, we observed that the proportion of promoter variants with either over- or under-expression effects was highest for positions with high constraint across 470 mammals (phyloP-470 way; Fig. 1F), indicating that both types of outliers disproportionately fall within conserved sequence elements. Next, we used multi-tissue outliers, as well as matched control variants that came from the promoters of the same genes but showed no effects on expression, and trained PromoterAI on the variants from the odd chromosomes (see Methods for details), while variants from the even chromosomes were set aside for testing and evaluating the model. We benchmarked PromoterAI, Evo2, DeepSEA, Basenji, ExPecto, Enformer, PuffinD, ChromBPNet, and Borzoi on three variant classification tasks, namely, under- vs overexpression, under-expression vs control variants that did not affect gene expression (null), and overexpression vs null (Fig. 1G, left). PromoterAI performed the best out of these models across all three metrics by achieving an auROC of 0.89, 0.80, and 0.74, respectively. We also benchmarked the performance of each classifier on multiple massively parallel reporter assay (MPRA) saturation mutagenesis datasets, including the Critical Assessment of Genome Interpretation (CAGI5) dataset, where we focused on promoters of disease-associated genes (3), and another dataset comprising promoters with known eQTLs and hits from genome-wide association studies (GWAS) (34). PromoterAI achieved the best performance (Fig. 1G, right, and fig. S1E) and also demonstrated strong per-gene predictive capability, as evidenced by correlations between model scores and MPRA effect sizes calculated
Downloaded from https://www.science.org at Cold Spring Harbor Laboratory on August 06, 2025


First release: 29 May 2025 science.org (Page numbers not final at time of first release) 3
separately for each gene (fig. S1F).
Under- and overexpression variants primarily act by disrupting motifs
To understand the basis of PromoterAI’s predictions, we systematically performed in-silico mutagenesis in the vicinity of GTEx outliers, using only the variants from the even chromosomes that were set aside for testing and evaluating the model. Consistent with the model learning the underlying biological mechanisms, PromoterAI assigned high scores to variants that disrupted the motifs of TFs with widely known effects on gene expression, with several representative examples shown (Fig. 2A and fig. S2A). Under-expression variants were particularly enriched for disrupting ETS motifs (Fig. 2B), recognized to be involved in a wide range of regulatory processes (35), along with YY1, ATF1, and NRF1 motifs, which are known to have broad regulatory functions at promoters (36–39). In contrast, overexpression outliers were enriched for disrupting E2F motifs (Fig. 2B), which have roles in cell cycle control and tumor suppression (40), alongside other transcriptional families such as NFKB, INSM1, ERR-alpha, and TFAP2 (41–44). We inserted these motifs along the lengths of promoters to assess their effects using PromoterAI and found that the predictions were consistent with the evidence from expression outliers (fig. S2B), while also localizing the optimal region for maximal effect to 100 bp immediately upstream and downstream of the TSS (45). These in-silico experiments reinforce that PromoterAI successfully learns sequence determinants of promoter regulation. We further examined the relative contributions of variants that affect gene expression by strengthening or weakening motifs, using a curated list of motifs (table S4) exclusively associated with activating or repressive Gene Ontology (GO) terms (46). We found that under-expression outliers were enriched for weakening of motifs associated with only activating GO terms (Fig. 2C, left). In contrast, overexpression outliers were enriched for both weakening of motifs associated with only repressive GO terms and, to a lesser extent, strengthening of motifs associated with only activating GO terms (Fig. 2C, right). Similar trends were observed for motifs associated with both activating and repressive GO terms, subcategorized based on their relative proportions, although the effects were less pronounced due to the complexity of these motifs (fig. S2C). Our results show that it is generally easier for newly occurring variants to cause outlier gene expression by disrupting existing regulatory components rather than creating new ones, as well as highlighting the important role of transcriptional repressors in promoter regulation. Next, we investigated the contribution of fine-tuning to PromoterAI’s performance, aiming to discern the specific variant types which saw the largest gains. We first stratified expression outliers by the motifs they overlapped, and
calculated the correlation between the actual and predicted effects of variants before and after fine-tuning. We saw substantial improvement in performance across nearly all motifs (Fig. 2D) as well as across TSS distance (fig. S2D). We also repeated the earlier analysis of inserting motifs along the lengths of promoters, using the models before and after finetuning. We observed that the predicted direction of effect aligned more consistently with the evidence from expression outliers after fine-tuning (fig. S2E), suggesting that fine-tuning may overcome limitations in predicting the directionality of motifs which has been highlighted in recent papers (26, 27). Moreover, we found that the nucleotide positions which experienced the largest differences in PromoterAI scores before and after fine-tuning tended to be highly conserved (Fig. 2E). Given that conservation information was not used during training or fine-tuning, these results provide independent evidence suggesting that fine-tuning with expression outliers enables the model to more effectively learn the underlying causal biological mechanisms. We show representative examples where fine-tuning led to substantial changes in PromoterAI predictions, noting that the model before fine-tuning frequently recognized the underlying motif but predicted the direction of effect incorrectly (Fig. 2F and fig. S2F). In order to investigate PromoterAI’s internal representations of promoter biology, we extracted feature embeddings of all canonical promoters of protein-coding genes, finding that they clustered into three distinct classes (fig. S3A and table S5): a first class of 9,234 genes abundantly expressed across most tissues and enriched for activating histone marks and ubiquitously expressed TFs, a second class of 3,430 genes characterized by bivalent chromatin, and a third class of 6,632 genes with mainly tissue-specific expression and enriched for repressive histone marks (fig. S3, B to F) (47). The three classes also differed in expression variability, with the first class resembling low-variability promoters, while the second and third classes resembled highly variable promoters (48). Observing that the third promoter class bears many enhancer-like characteristics (49), we used PromoterAI to extract feature embeddings from distal enhancers and compared them with those from promoters. Indeed, the feature embeddings of distal enhancers closely resembled those of the third promoter class (fig. S3G). This enhancer-like promoter class was notably depleted for conserved elements, further emphasizing its distinct regulatory properties (fig. S3H).
Promoter variants that impact gene expression are under strong negative selection
To measure the extent of negative selection acting on promoter variants in human populations, we turned to the Genome Aggregation Database v3 (gnomAD) cohort (50, 51), which includes whole genome sequencing data from 71,702 individuals. We used PromoterAI to score each promoter
Downloaded from https://www.science.org at Cold Spring Harbor Laboratory on August 06, 2025


First release: 29 May 2025 science.org (Page numbers not final at time of first release) 4
variant, excluding those with protein-coding or splicing effects, and compared the number of predicted expression-altering variants at common allele frequencies (> 0.1%) to rare singletons. We based our analysis on the observation that natural selection purges deleterious variants over time, preventing them from becoming widespread in the population (50, 52).
Variants predicted by PromoterAI to alter gene expression were strongly depleted of common variants relative to rare variants, indicating that expression-altering variants are being actively removed by natural selection (Fig. 3A). This depletion applied to both predicted under- and overexpression variants and became more pronounced with increasing magnitude of PromoterAI scores. At PromoterAI scores near +0.5 or -0.5, the observed depletion approached that of missense variants, while scores closer to +1 or -1 resulted in a depletion percentage comparable to that of nonsense variants. We further stratified genes by evolutionary constraint, as measured by probability of loss-of-function intolerance (pLI) (52). We observed a greater depletion for expression-altering variants in the promoters of high pLI genes, consistent with these genes being dosage sensitive and associated with haploinsufficient disease mechanisms (Fig. 3B). We also stratified variants by TSS distance, dividing the promoter region into 10 bp non-overlapping windows and estimating the number of singleton variants expected to be depleted by natural selection in each window (Fig. 3C). Similar to expression outliers being enriched near the TSS, the depletion signal also peaked in a region within 100 bp of the TSS, with a notably stronger signal for under-expression than overexpression variants, suggesting that under-expression variants tend to be more deleterious on average.
PromoterAI accurately predicts the effects of finemapped promoter eQTLs
We assessed PromoterAI’s ability to predict the effects of finemapped GTEx eQTLs. For this, we identified 409 eQTLs that mapped to a single variant with posterior probability > 0.9 in multiple tissues using CAVIAR (53), after excluding splice variants, protein-coding variants, and GTEx 