AI for predicting and understanding antibiotic resistance in Mycobacteria: learning from genomes and protein structures.
In 2024, tuberculosis (TB) caused 1.5 million deaths worldwide. Treatment is challenging, and the increasing prevalence of drug-resistant strains makes rapid diagnosis essential. Conventional drug susceptibility testing requires culturing Mycobacterium tuberculosis complex (MTBC) with antibiotics, a process that takes ~2 months because of the bacterium’s slow growth. Whole-genome sequencing (WGS) has reduced this to around one week and is now used routinely in several high-income countries. The UK Health Security Agency was the first public health agency to adopt WGS for routine TB diagnostics in 2018, building on research from the University of Oxford. However, current WGS approaches rely on detecting known resistance mutations and cannot reliably predict resistance to newer drugs [1], such as bedaquiline, which are increasingly used to treat multidrug-resistant TB (MDR-TB).
AI offers the potential to transform WGS-based antimicrobial resistance prediction. Early ML models showed promise but often relied on shortcut learning, for example inferring bedaquiline resistance from mutations associated with rifampicin resistance [2]. Graph convolutional networks (GCNs) avoid this problem by focussing on proteins known to mediate resistance whilst also incorporating structural and chemical information, but they ignore the rest of the genome and cannot predict the effects of non-coding variants [3]. Meanwhile, biological foundation models, including those developed at EIT Oxford, are trained on large multimodal datasets that capture broader genetic information but have not yet been able to outperform simpler approaches for AMR prediction [4].
This DPhil will develop AI models that combine these complementary approaches while remaining computationally tractable. Stretch objectives include evaluating whether models generalise to new antibiotics and applying them to early-stage antibiotic development to anticipate resistance mechanisms.
The project is underpinned by the world’s largest curated mycobacterial genomics and drug susceptibility dataset [5], maintained at the University of Oxford. It currently comprises 53,897 MTBC isolates with matched WGS and phenotypic drug susceptibility data, with a further 20,000 being processed, alongside 56,829 additional mycobacterial genomes, including 7,168 with susceptibility data.
- Adlard D, Malone KM, Westhead J, Hunt M, Thai H, Colpus M, Turner RD, Omar SV, Eyre DW, Ismail N, Walker TM, Peto TEA, Crook DW, Iqbal Z, Fowler PW
Rapidly and reproducibly building a comprehensive catalogue of resistance-associated variants for M. tuberculosis
bioRxiv preprint doi:10.1101/2025.10.02.679941 - The CRyPTIC Consortium
Quantitative drug susceptibility testing for Mycobacterium tuberculosis using unassembled sequencing data and machine learning
PLoS Comp Biol doi: 10.1371/journal.pcbi.1012260 - Dissanayake D, Brunner VM, Adlard D, Morrone J, Fowler PW
Predicting pyrazinamide resistance in M. tuberculosis using a graph convolutional network
BMC Microbiology doi:10.1186/s12866-026-04876-1 - Tasmin M, Mohanty S, Kulkarni S, Farhat MR, Green AG
BIG-TB: A benchmark for prediction and interpretability of sequence-based machine learning using Mycobacterium tuberculosis genomes.
bioRxiv doi:10.64898/2026.01.30.702134 - The CRyPTIC dataset v3.4.0, https://crypticproject.info/datasets/

