The dataset, released through the AlphaFold Database on September 24, focuses on viral families known to infect humans and is intended to give scientists a faster starting point for studying unfamiliar pathogens, identifying protein interactions and investigating possible vaccine, diagnostic and treatment targets before or during outbreaks.
The collaboration brings together NVIDIA, Google DeepMind, the European Molecular Biology Laboratory’s European Bioinformatics Institute, the Coalition for Epidemic Preparedness Innovations, Seoul National University, Sungkyunkwan University, the Swiss Institute of Bioinformatics and the University of Glasgow.
Researchers analysed sequences for 41,774 proteins from about 2,800 viruses, including viruses associated with mpox, measles and hepatitis B. AlphaFold2 was used to predict interacting protein structures, while NVIDIA’s BioNeMo Inference Runtime helped scale the computational work across thousands of viral proteomes.
The project generated predictions for 40,746 homodimers, complexes containing two identical protein molecules, and almost 1.7 million heterodimers, which contain two different molecules. Quality thresholds narrowed the entries added to the AlphaFold Database to 2,749 homodimers and 5,279 heterodimers. The wider set of predictions has also been made publicly accessible.
NVIDIA is separately opening the BioNeMo Structure Prediction Pipeline used for the work. The GPU-accelerated workflow takes researchers from a protein sequence to a predicted three-dimensional structure, allowing laboratories to apply the same approach to targets beyond those covered by the viral dataset.
The structures were inferred with AlphaFold2 rather than determined experimentally. Predictions carry confidence information, enabling researchers to assess which computational models warrant closer investigation. High-confidence predictions can guide experiments, but laboratory methods remain necessary to establish whether predicted interactions and structures accurately reflect biological behaviour.
Jo McEntyre, interim director of EMBL-EBI, said open access was important for understanding viral diagnostics and developing treatments and vaccines, while also lowering barriers for scientists in resource-constrained settings dealing directly with outbreaks.
Risha Patel, life sciences partnerships manager at Google DeepMind, said the collaboration would provide researchers worldwide with structural insights intended to strengthen preparation for future outbreaks. The AlphaFold Database now contains more than 260 million protein and protein-complex predictions spanning much of the protein sequence information catalogued by science.
The viral work also addresses a persistent gap in structural biology. Experimental techniques such as X-ray crystallography can require substantial time and resources, whereas AI systems can produce candidate structures far faster and at large scale. That speed is particularly useful when researchers must decide which proteins or interactions merit scarce laboratory capacity.
One reason for building a library before an emergency is that structural knowledge can shorten the phase of pathogen research. The three-dimensional shape of a protein influences how it works and how it interacts with other molecules, information that can help scientists formulate testable hypotheses about viral replication, immune evasion and intervention points. The approach does not identify a vaccine or medicine by itself, but can help researchers decide which questions to test.
The Swiss Institute of Bioinformatics contributed curated viral information, including data on polyproteins, precursor molecules that are later cut into functional proteins. Such processing can make it difficult to identify precise protein boundaries directly from viral genetic sequences and can undermine structure predictions if the underlying sequence definitions are incomplete.
The institute said it supplied unique protein data for 699 viruses and helped check the biological plausibility of predicted interactions. Its work covered polyproteins found in viruses including Zika, dengue and poliovirus.
The dataset nevertheless has important limits. Some viral proteins carry sugar molecules that are not represented in the predictions, while many function in assemblies larger than the two-protein complexes modelled in this project. Those omissions mean the structures should be treated as computational hypotheses rather than substitutes for experimental evidence.
Follow Arabian Post
Select Arabian Post as your preferred source on Google and MSN News for trusted business news and Arab politics and updates.