bioRxiv · 10.1101/2024.08.07.606674
LitGene: a transformer-based model that uses contrastive learning to integrate textual information into gene representations
Abstract
Representation learning approaches leverage sequence, expression, and network data, but utilize only a fraction of the rich textual knowledge accumulated in the scientific literature. We present LitGene, an interpretable transformer-based model that refines gene representations by integrating textual information. The model is enhanced through a Contrastive Learning (CL) approach that identifies semantically similar genes sharing a Gene Ontology (GO) term. LitGene demonstrates accuracy across eight benchmark predictions of protein properties and robust zero-shot learning capabilities, enabling the prediction of new potential disease risk genes in obesity, asthma, hypertension, and schizophrenia. LitGenes SHAP-based interpretability tool illuminates the basis for identified disease-gene associations. An automated statistical framework gauges literature support for AI biomedical predictions, providing validation and improving reliability. LitGenes integration of textual and genetic information mitigates data biases, enhances biomedical predictions, and promotes ethical AI practices by ensuring transparent, equitable, open, and evidence-based insights. LitGene code is open source and also available for use via a public web interface at litgene.avisahuai.com.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jararweh, A., Macaulay, O., Arredondo, D., Oyebamiji, O., Hu, Y., Tafoya, L., Zhang, Y., Virupakshappa, K., Sahu, A.. 2024-08-07. LitGene: a transformer-based model that uses contrastive learning to integrate textual information into gene representations. https://doi.org/10.1101/2024.08.07.606674
Cite the original work for its findings. Save a collection to share your selection of sources.