Curie Brief
Turn on cookies to sign in
Signing in saves your progress to your Curie account. We can only do that with cookies on — turn them on to continue.

A new machine learning model from Stanford can map how DNA "enhancers" control genes across thousands of cell types using single-cell data. Trained on over 10,000 CRISPR-tested enhancer–gene pairs, the scE2G models can pinpoint disease-linked genetic variants with unprecedented precision. The tool has already connected two genes to lymphocyte counts in the blood—a link nearly impossible to make from noncoding DNA alone.
Scientists at Stanford University have developed a scalable machine learning model called scE2G that maps how DNA enhancers—regulatory stretches of DNA that control when and how strongly genes are turned on—interact with specific genes across diverse cell types. Published in Nature Genetics, the study addresses a long-standing challenge: enhancer activity is highly cell-type-specific, making it hard to predict which enhancers regulate which genes in any given tissue.
The scE2G models were trained on data from CRISPR experiments that directly tested more than 10,000 candidate enhancer–gene pairs. They work with single-cell genomic data (scATAC-seq and multiomic scATAC + scRNA-seq), allowing researchers to study rare or hard-to-isolate cell types that bulk methods can't easily reach. The models perform reliably across datasets of varying sizes and sequencing depths, making them broadly applicable to existing single-cell datasets.
Key Takeaways:
Why it matters: Most disease-associated genetic variants fall in noncoding regions of the genome, making them hard to interpret. Tools like scE2G could transform how researchers connect these variants to specific genes and cell types—accelerating the path from genetic discovery to therapeutic targets.