Search bioRxiv⌕ Search

Biology subjects

Murtaza, G.

Publications and source records attributed to Murtaza, G..

2 recordsLinked to original sources

Investigating the performance of deep learning methods for Hi-C resolution improvement

MotivationHi-C is a widely used technique to study the 3D organization of the genome. Due to its high sequencing cost, most of the generated datasets are of coarse resolution, which makes it impractical to study finer chromatin features such as Topologically Associating Domains (TADs) and chromatin loops. Multiple deep-learning-based methods have recently been proposed to increase the resolution of these data sets by imputing Hi-C reads (typically called upscaling). However, the existing works evaluate these methods on either synthetically downsampled or a small subset of experimentally generated sparse Hi-C datasets, making it hard to establish their generalizability in the real-world use case. We present our framework - Hi-CY - that compares existing Hi-C resolution upscaling methods on seven experimentally generated low-resolution Hi-C datasets belonging to various levels of read sparsities originating from three cell lines on a comprehensive set of evaluation metrics. Hi-CY also includes four downstream analysis tasks, such as TAD and chromatin loops recall, to provide a thorough report on the generalizability of these methods. ResultsWe observe that existing deep-learning methods fail to generalize to experimentally generated sparse Hi-C datasets showing a performance reduction of up to 57 %. As a potential solution, we find that retraining deep-learning based methods with experimentally generated Hi-C datasets improves performance by up to 31%. More importantly, Hi-CY shows that even with retraining, the existing deep-learning based methods struggle to recover biological features such as chromatin loops and TADs when provided with sparse Hi-C datasets. Our study, through Hi-CY framework, highlights the need for rigorous evaluation in future. We identify specific avenues for improvements in the current deep learning-based Hi-C upscaling methods, including but not limited to using experimentally generated datasets for training. Availabilityhttps://github.com/rsinghlab/Hi-CY Author SummaryWe evaluate deep learning-based Hi-C upscaling methods with our framework Hi-CY using seven datasets originating from three cell lines evaluated using three correlation metrics, four Hi-C similarity metrics, and four downstream analysis tasks, including TAD and chromatin loop recovery. We identify a distributional shift between Hi-C contact matrices generated from downsampled and experimentally generated sparse Hi-C datasets. We use Hi-CY to establish that the existing methods trained with downsampled Hi-C datasets tend to perform significantly worse on experimentally generated Hi-C datasets. We explore potential strategies to alleviate the drop in performance such as retraining models with experimentally generated datasets. Our results suggest that retraining improves performance up to 31 % on five sparse GM12878 datsets but provides marginal improvement in cross cell-type setting. Moreover, we observe that regardless of the training scheme, all deep-learning based methods struggle to recover biological features such as TADs and chromatin loops when provided with very sparse experimentally generated datasets as inputs.

bioinformatics↗

Population structure and association mapping for hundred seed weight in mungbean minicore

Mungbean is an important legume rich in protein and carbohydrates. It is cultivated mostly in south to south east Asia and east Africa. One of the important yield contributing traits in mungbean is 100 seed weight. Variation for this trait is available in the mungbean germplasm, and to facilitate improvement of this trait through breeding, genome-wide association mapping for 100 seed weight was carried out in the World Vegetable Center mungbean minicore collection. A total of 24,870 single nucleotide polymorphic markers were tested for association with 100 see weight in plants grown in Pakistan, Bangladesh, Myanmar and Taiwan and candidate loci associated with seed weight on chromosomes 1, 4, 6, 8 and 9 were identified. None of the QTLs was stable across all environments, but loci on chromosomes 4 and 8 were significant in Pakistan and Bangladesh and other loci on chromosome 8 were significantly associated in plants grown in Myanmar and Taiwan.

bioinformatics↗