OmicsPred Compiles Multi-Omics Genetic Prediction Models, Expanding the Landscape for Drug Target Exploration

Background
Genetic prediction models that estimate the impact of genetic variants on RNA expression levels, protein concentrations, and metabolite levels have become essential tools for elucidating disease mechanisms and discovering new drug targets. Transcriptome-wide association studies (TWAS) and proteome-wide association studies (PWAS) link diseases to molecular traits by indirectly estimating them from genome-wide association study (GWAS) data. These approaches allow researchers to explore the pathway from genetic variants to molecular traits and diseases without directly sampling patient tissues.
The challenge has been that these models are scattered across paper appendices, lab servers, and individual repositories like PredictDB. Variant coordinates, effect alleles, and weight formats vary, and information on training population ancestry, measured tissues, or platforms is often missing. This inconsistency made it difficult to apply the same model to independent cohorts or compare across omics layers. In particular, models developed in European-ancestry populations may yield different association power and effect estimates when applied to other ancestry groups.
A research team from the University of Cambridge addressed this fragmentation by building OmicsPred, an open platform for registering, searching, and distributing genetic prediction models for multi-omics traits. The study was published in Nature Genetics on September 1, 2026.
Key Findings
OmicsPred assigns each predictive score an OPGS identifier and provides standardized formats for chromosomal positions, rsIDs, effect alleles, and variant-specific weights. It also links to the original genome build, molecular traits, measured tissues and platforms, model development methods, training and validation sample sizes and ancestry compositions, performance metrics, and usage conditions. Gene, transcriptome, proteome, and metabolome data are integrated with Ensembl, UniProt, and ChEBI, while tissues and phenotypes are normalized using established ontologies.
The file format is designed based on the multi-polygenic score catalog (PGS Catalog) standard. Users can download models in pgsc_calc-compatible score files or PredictDB formats executable in MetaXcan. A REST application programming interface (API) is also provided for querying metadata within the platform. The implementation code for the platform and database is publicly available under the Apache 2.0 license.
The research team validated the platform's utility using disease GWAS data from the U.S. Million Veteran Program (MVP). They analyzed 9 datasets of blood- and plasma-based transcriptome and proteome predictive scores, totaling 38,450 scores, using S-PrediXcan. In European-ancestry data, 1,233 PheCode disease traits were evaluated, 986 in African-ancestry data, and 719 in mixed-ancestry data. Only GWAS with ancestry and training cohorts that were similar were matched. Scores with less than 75% overlap with GWAS variants were excluded, and multiple testing was corrected per dataset.
The analysis identified 194,138 false discovery rate-corrected significant associations. The association between CFH expression and age-related macular degeneration reached a P-value of 1.29×10⁻⁶¹, and that between PCSK9 expression and coronary artery disease reached 7.61×10⁻²⁵. Thirty-one gene-disease combinations were consistently replicated across all six European-ancestry datasets, and five combinations were commonly detected across ancestry groups in the SomaLogic proteome models.
Implications and Outlook
The value of OmicsPred lies in transforming existing models into a reusable research infrastructure, rather than introducing new prediction algorithms. Comparing predictions for the same gene or protein across tissues, measurement platforms, omics layers, and ancestry groups can help prioritize candidates repeatedly linked to specific diseases. The strong replication of known associations such as CFH and PCSK9 demonstrates the platform's ability to recover disease-related signals.
The open API and standardized metadata are also useful for automating large-scale target discovery pipelines. Pharmaceutical companies can evaluate associations between genetically predicted molecular traits and hundreds of diseases in bulk, assessing not only potential indications but also possible adverse signals. If a candidate is consistently identified across multiple omics layers, it can strengthen the rationale for prioritizing follow-up functional experiments.
However, associations between genetic predictions and diseases alone do not prove causality or drug efficacy. Linkage disequilibrium, horizontal pleiotropy, tissue specificity, and measurement platform differences can distort results. The continued dominance of European-ancestry models and data remains a limitation. To identify candidates with high potential for therapeutic targets, more models developed and validated in diverse ancestry groups and cell types are needed, combined with colocalization analysis, Mendelian randomization, and experimental validation.
Nature Genetics, Published online: 01 September 2026; doi:10.1038/s41588-026-02726-4We present OmicsPred, an open platform for the deposition and dissemination of genetic prediction models of multi-omic traits, with key metadata required for reproducibility and independent applications. We illustrate the utility of the resource for target discovery with a multi-omic multi-ancestry phenome-wide association analysis using data from the Million Veterans Program.
In drug development, researchers can compare disease GWAS with OmicsPred scores to prioritize RNA or protein candidates regulated by genetics. For example, in a cardiovascular disease program, candidates consistently identified across multiple transcriptome and proteome datasets can be passed to colocalization analysis and cell experiments, and their associations with other diseases can be checked to review indications and safety risks early.
Hospital and biobank researchers can apply standardized score files to their own genomic data to estimate molecular traits in large patient populations where direct tissue measurements are difficult. However, to use these as clinical decision-making tools, separate validation in the target population and quantification of prediction accuracy and ancestry-specific biases are necessary.