Enhancing Disease Prediction Accuracy with a Bayesian Algorithm for Estimating the Effects of Rare Genetic Variants

Background
Existing genomic studies have primarily focused on analyzing the impact of common variants found in populations. Tracking common variants is relatively straightforward, enabling the identification of large-scale correlations and serving as a useful basis for Polygenic Risk Score (PRS) models. However, predicting the disease-causing risk of unique and rare genetic variants in individual patients is a challenging task. Extremely rare variants pose difficulties in achieving statistical significance due to insufficient sample sizes.
To overcome this limitation, researchers have relied on gene-burden tests, which aggregate the effects of multiple variants. This approach ignores the detailed differences between individual variants and sums the number of variants at the gene level. This can lead to a critical error where the effects of harmful variants and beneficial variants are offset. Furthermore, even non-disease-related, benign variants can be included in the analysis, reducing predictive performance. Therefore, there is a need for an analytical model that can precisely extract the specific impact of individual rare variants to improve the performance of PRS models.
Key Findings
The research team at Seoul National University's School of Public Health, led by Professor Seung-Geun Lee, developed 'RareEffect,' a statistical analysis method that calculates the impact of rare variants individually. The team designed a method that first measures the heritability at the gene or genomic region level and then applies a Bayesian framework to calculate the specific effect size of each variant. This variance component model precisely captures the pure impact of individual variants in regions where harmful and beneficial variants are mixed.
Technical improvements were also made to efficiently handle large-scale biobank data. The team incorporated the 'FaST-LMM' (Factored Spectrally Transformed Linear Mixed Models) technique, which utilizes the mathematical structure of the genetic kinship matrix. This technique significantly reduces the computational cost and time required for analysis, bringing the analysis efficiency to a realistic level. In addition, to stabilize the analysis of disease data with extremely unbalanced ratios of patients and controls, a fast algorithm for 'Firth bias correction' was integrated into the calculation process. To prevent the loss of signals from extremely rare variants, a variant reduction strategy was used in conjunction to combine them appropriately, further enhancing the analytical power.
The research team demonstrated the practical performance of RareEffect using data from the UK Biobank (UKB), a cohort of 450,000 individuals with Whole Exome Sequencing (WES) data. The analysis was performed on 100 quantitative traits and binary disease traits, and the results showed a more precise predictive power compared to existing burden test models. In particular, the results were notable in identifying genetic risk factors for major chronic diseases such as type 2 diabetes and cancer, as well as continuous indicators such as cholesterol levels.
Significance and Prospects
The emergence of RareEffect is expected to mark a turning point in genomic research, shifting the focus from common variants to rare variants. This analytical technique enables the identification of the risk of rare variants in individual patients at the single-variant level, paving the way for true personalized precision medicine. In particular, as the analysis of rare variants in drug-metabolizing genes becomes easier, the efficiency of clinical trial design for pharmaceutical companies developing new drugs is likely to be significantly improved.
However, there are still challenges to be addressed before it can be fully implemented in clinical practice. One of the challenges is to reduce the cost of the computing infrastructure required to decode large-scale WES data and run the RareEffect algorithm. In addition, since the UK Biobank data is mainly from European populations, it is necessary to verify whether the same predictive performance is observed in other ethnic groups, including Asians, through multi-ethnic cohort validation. It has also been pointed out that the statistical assumptions are somewhat simplistic to explain the functional interactions of numerous variants in the non-coding region of the genome.
Nature Genetics, Published online: 17 August 2026; doi:10.1038/s41588-026-02705-9RareEffect is an accurate and computationally efficient Bayesian method for estimating rare variant effect sizes. These estimates can be leveraged to improve polygenic prediction models, thereby outperforming other methods for rare variant polygenic scoring.
In clinical practice, RareEffect can be used to generate rare variant-based PRSs, enabling the early identification of high-risk individuals for serious diseases such as cardiovascular disease and cancer. Patients with borderline conditions, who were previously difficult to diagnose with existing tests, can now identify potential risk factors through precise analysis of individual genetic variants and receive personalized preventive treatment. A screening process that identifies individuals with rare variants that cause serious adverse reactions or reduced efficacy to specific drugs before the start of clinical trials. This is a useful strategy to improve clinical trial success rates and significantly reduce development costs. In the development of Companion Diagnostics (CDx) devices, the results of this analysis can also be reflected to increase the likelihood of approval for precision medicines that target specific genetic variants.