Large Language Models (LLMs) have revolutionized Artificial intelligence (AI). Despite being originally designed for text, these models turn out to be incredibly effective for biological sequences like DNA and proteins. We were among the first to develop protein language models (pLMs) as global predictors of protein structure and protein function. We also showed that protein and DNA language models can accurately distinguish between disease-causing and benign mutations. The Brandes Lab continues to work on challenges that would unlock the full potential of genomic AI to prediagnose and treat disease and understand our genomes.

Join Us!

If you are excited about this research agenda, come work with us.

1. Predict diverse mutation effects in coding & noncoding regions

We are looking to further improve mutation effect predictions with protein and DNA language models. Existing variant effect prediction algorithms try to predict whether a given variant is damaging or neutral, but variant effect is not a one-dimensional phenomenon. For example, different mutations in the same gene may lead to loss-of-function, gain-of-function or dominant-negative effects. Deep-learning embeddings are inherently high-dimensional, allowing us to tease apart these different effects. We also use data from high-throughput experiments such as deep mutational scans and single-cell RNA-sequencing to improve our predictions.

2. Mechanistic interpretability of genomic foundation models

AI models are notorious black boxes, but the rapid advances in mechanistic interpretability provide an opportunity to study what is really going on inside them. This is critical for promoting their adoption (especially in the clinic), know when to trust them, learn about the mechanisms of mutations, understand the failure modes of existing models so that we can improve them, and even discover entirely new biology. This is also a fascinating case study for studying AI that has been trained on natural rather than man-made artefacts: what concepts does an AI learn entirely on its own?

3. Genetic engineering

We leverage our improved models of genetic effects to search for mutation combinations that optimize the genetic background of cells, for example to make immune cells more potent against tumors in cancer immunotherapy.

4. Improve genomic foundation models

A key bottleneck in genomic AI is insufficiently challenging benchmarks that mislead model developers. Human genetics provides uniquely challenging problems that only models with true knowledge about the underlying biology can tackle. With better benchmarks we can then train better genomic foundation models that address these important but neglected tasks. For example we are training phylogenetic-aware models that scale better.

5. Identify rare variant effects and causal genes

Genome-wide association studies (GWAS) and polygenic risk scores (PRS) are purely statistical: they search for genetic variants correlated with disease status without knowing anything about the molecular effects of these variants. Variant effect predictions, especially those made by frontier AI models, can change this and help guide GWAS and PRS towards variants more likely to have an effect. These functional priors are especially important when statistical evidence is limited, as in the case of rare mutations, and when attempting to distinguish between causal and non-causal associations. We test our methods on large-scale genetic cohorts such as the UK Biobank.

6. Clinical implementation of genomic AI in genetic testing

We are looking to improve the clinical guidelines to make better use of advanced genomic AI, for example when diagnosing rare genetic diseases. To do that, we evaluate the capacity of these models to provide better clinical evidence compared to existing protocols.

7. Prevent dangers from advanced AI

We are evaluating biosecurity risks from genomic AI models to inform researchers and policymakers.

* * * * *

To learn more about the philosophy behind this research agenda, you can read this blog post.