bioRxiv · 10.64898/2026.09.15.751721
Discriminating rare disease cases from their controls based on observed and excluded phenotypes
Abstract
Rare diseases are individually uncommon but collectively prevalent. Their primary clinical challenge lies not in treatment but in diagnosis. In the early stages of clinical management, it is frequently unclear whether the observed phenotypes are associated with a rare disease. Leveraging machine learning methods to mine latent associations between these phenotypes and rare diseases for early diagnosis offers a viable strategy to alleviate this diagnostic dilemma. In the present study, we demonstrated that machine learning can effectively discriminate rare disease cases from their controls using observed and excluded phenotypes as features. Among them, the Random Forest model achieved the best classification performance with a certain degree of generalizability. Based on further analysis of the feature selection results, we conclude that the two factors, specificity and occurrence count, are important for phenotype selection in rare disease discrimination, and comparable importance should be attached to both observed and excluded phenotypes, during feature construction.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Guo, Y., Yao, Y., Zhang, Y., Duan, G., Zhang, N., shan, g.. 2026-09-21. Discriminating rare disease cases from their controls based on observed and excluded phenotypes. https://doi.org/10.64898/2026.09.15.751721
Cite the original work for its findings. Save a collection to share your selection of sources.