Search bioRxiv⌕ Search

Biology subjects

Tran, V. D.

Publications and source records attributed to Tran, V. D..

3 recordsLinked to original sources

First-Trimester Non-Invasive Prediction of Preterm Birth Using Cell-Free DNA Fragmentomics

ObjectiveTo develop and validate a cell-free DNA (cfDNA) fragmentomic classifier for the early prediction of spontaneous preterm birth (PTB) using routine first-trimester non-invasive prenatal testing (NIPT) data. MethodsA nested case-control study was conducted within a prospective multicenter Vietnamese cohort comprising 286 pregnancies, including 82 spontaneous PTB cases and 204 term controls. Maternal plasma cfDNA collected during routine first-trimester NIPT (median gestational age, 12 weeks) was sequenced to a depth of approximately 20 million reads per sample. Five fragmentomic feature categories including copy number alterations, end-motif composition, nucleosome distance, fragment length, and joint fragment-lengthxend-motif were evaluated for PTB prediction. Machine learning classifiers were developed in a training cohort (n = 228, 65 PTB vs 163TB) and tested in a validation cohort (n = 58, 17 PTB vs 41 TB). ResultsAmong the five fragmentomic feature classes evaluated, 4-mer end-motif (EM) profiles exhibited the most pronounced differences between PTB and term control samples. Consistent with these findings, the EM-based classifier demonstrated the highest discriminative performance in the validation cohort, achieving an AUC of 0.970 (95% CI, 0.912-1.000). At a specificity >90%, the model achieved a sensitivity of 94% (95% CI, 78-100%). ConclusionThese findings demonstrate that cfDNA EM signatures derived from routine first-trimester NIPT can accurately identify pregnancies at risk of spontaneous preterm birth, without additional blood collection or sequencing, thereby extending the clinical utility of existing prenatal screening infrastructure. KEY POINTSO_ST_ABSWhat is already known about this topic?C_ST_ABSO_LICurrent first-trimester prediction strategies based on maternal characteristics, cervical length, and biochemical markers have limited predictive accuracy, particularly in nulliparous women. C_LIO_LIExisting cfDNA-based approaches have shown only modest performance or require additional assays, limiting clinical applicability. C_LI What does this study add?O_LIExisting NIPT sequencing data can be repurposed (without additional blood sampling or sequencing) for accurate prediction of spontaneous preterm birth (AUC=0.970). C_LIO_LIA classifier employing 4-mer end-motif (EM) profiles achieved an AUC of 0.970. At a specificity >90%, the model achieved a sensitivity of 94%. C_LI

genomics↗

DisGeneFormer: Precise Disease Gene Prioritization by Integrating Local and Global Graph Attention

Identifying genes associated with human diseases is essential for effective diagnosis and treatment. Experimentally identifying disease-causing genes is time-consuming and expensive. Computational prioritization methods aim to streamline this process by ranking genes based on their likelihood of association with a given disease. However, existing methods often report long ranked lists consisting of thousands of potential disease genes, often containing a high number of false positives. This fails to meet the practical needs of clinicians who require shorter, more precise candidate lists. To address this problem, we introduce DisGeneFormer (DGF), an end-to-end disease-gene prioritization pipeline. Our approach is based on two distinct graph representations, modeling gene and disease relationships, respectively. Each graph is first processed separately by graph attention and then jointly by a transformer module to combine within-graph and cross-graph knowledge through local and global attention. We propose an evaluation pipeline based on the precision of a top K ranked gene list, with K set to clinically feasible values between 5 and 50, relying solely on experimentally verified associations as ground truth. Our evaluation demonstrates that DGF substantially outperforms existing methods. We additionally assessed the influence of the negative data sampling strategy as well as analyses of the effect of graph topology and features on the performance of our model.

bioinformatics↗

GraphProt2: A novel deep learning-based method for predicting binding sites of RNA-binding proteins

CLIP-seq is the state-of-the-art technique to experimentally determine transcriptome-wide binding sites of RNA-binding proteins (RBPs). However, it relies on gene expression which can be highly variable between conditions, and thus cannot provide a complete picture of the RBP binding landscape. This creates a demand for computational methods to predict missing binding sites. Here we present GraphProt2, a computational RBP binding site prediction framework based on graph convolutional neural networks (GCNs). In contrast to current CNN methods, GraphProt2 offers native support for the encoding of base pair information as well as variable length input, providing increased flexibility and the prediction of nucleotide-wise RBP binding profiles. We demonstrate its superior performance compared to GraphProt and two CNN-based methods on single as well as combined CLIP-seq datasets. Conceived as an end-to-end method, GraphProt2 includes all necessary functionalities, from dataset generation over model training to the evaluation of binding preferences and binding site prediction. Various input types and features are supported, accompanied by comprehensive statistics and visualizations to inform the user about datatset characteristics and learned model properties. All this makes GraphProt2 the most versatile and complete RBP binding site prediction method available so far.

bioinformatics↗