Genetic Risk Scores Lose Accuracy the Further a Person Is From the Study Group Behind Them

Genetic risk scores for common diseases predict less accurately for a person the further that person's ancestry sits from the group the score was built in, and the decline follows a straight line. The measurement comes from 245,388 whole-genome sequences in the U.S. All of Us Research Program, reported on September 14, 2026 in Nature Genetics.
These scores add up many small genetic effects to estimate someone's chance of developing a disease. Kristin Tsuo, Alicia R. Martin and colleagues, at the Broad Institute of MIT and Harvard, and Massachusetts General Hospital, report that scores trained on people of several ancestries can improve prediction for groups under-represented in genetic databases, but that datasets linking genetic and health records across broad human diversity remain limited.
Using the All of Us sequences together with UK Biobank data, the team built scores for 32 traits and diseases and tested them on All of Us participants of varied ancestry. Individual accuracy declined linearly as a participant's ancestry diverged from the group in the study where the genetic associations were first found. Training on several ancestries at once did not remove that decline; the authors report that it was reduced.
More data was not automatically better. Pooling the two collections to maximize sample size was not the best option for every trait: for traits driven by fewer genetic variants, training on All of Us alone performed best in participants of African ancestry, which the authors read as evidence of genetic effects concentrated within particular ancestries.
The authors argue that the results underscore the value of more representative biobanks. The score weights and the analysis code from the study have been deposited publicly on Zenodo.
Sources
- Peer-reviewedNature Genetics
