Search bioRxiv⌕ Search

Biology subjects

Arangasamy, Y.

Publications and source records attributed to Arangasamy, Y..

3 recordsLinked to original sources

Enhancing genome recovery across metagenomic samples using MAGmax

SummaryThe number of metagenome-assembled genomes (MAGs) is rapidly increasing with the growing scale of metagenomic studies, driving fast progress in microbiome research. Sample-wise assembly has become the standard due to its computational efficiency and strain-level resolution. It requires dereplication, the removal of near-identical genomes assembled in different metagenomic samples. We present MAGmax, an efficient dereplication tool that enhances both the quantity and quality of MAGs through a strategy of bin merging and re-assembly. Unlike dRep, which selects a single representative bin per genome cluster, MAGmax merges multiple bins within a cluster and reassembles them to increase coverage. MAGmax produces more dereplicated, higher-quality MAGs than dRep at 1.6x its speed and using three times less memory. Availability and implementationThe MAGmax open source software, implemented in Rust, is available under the GPLv3 license at https://github.com/soedinglab/MAGmax.

bioinformatics↗

Evaluation of Metagenome Binning: Advances and Challenges

BackgroundSeveral recent deep learning methods for metagenome binning claim improvements in the recovery of high quality metagenome-assembled genomes. These methods differ in their approaches to learn the contig embeddings and to cluster them. Rapid advances in binning require rigorous benchmarking to evaluate the effectiveness of new methods. We have benchmarked newly developed state-of-the-art deep learning binners on CAMI2 datasets, including our own, McDevol. ResultsThe results show that COMEBin and GenomeFace give the best binning accuracy, although not always the best embedding accuracy. Interestingly, post-binning reassembly consistently improves the quality of low coverage bins. We find that binning coassembled contigs with multi-sample coverage is effective for low coverage dataset while binning multi-sample contigs with multi-sample coverage ( multi-sample) is effective for high-coverage samples. In multi-sample binning, splitting the embedding space by sample before clustering showed enhanced performance compared to the standard approach of splitting final clusters by sample. ConclusionsCOMEBin and GenomeFace emerged as the top-performing tools overall, with MetaBAT2 and GenomeFace demonstrating superior speed. To facilitate future development, we provide workflows for standardized benchmarking of metagenome binners.

bioinformatics↗

Understanding the roles of secondary shell hotspots in protein-protein complexes

Hotspots are interfacial residues in protein-protein complexes that contribute significantly to complex stability. Methods for identifying interfacial residues in protein-protein complexes are based on two approaches, namely, (a) distance-based methods, which identify residues that form direct interactions with the partner protein and (b) Accessibility Surface Area (ASA)-based methods, which identify those residues which are solvent-exposed in the isolated form of the protein and become buried upon complex formation. In this study, we introduce the concept of secondary shell hotspots, which are hotspots uniquely identified by the distance-based approach, staying buried in both the bound and isolated forms of the protein and yet forming direct interactions with the partner protein. From the analysis of the dataset curated from Docking Benchmark 5.5, comprising of 94 protein-protein complexes, we find that secondary shell hotspots are more evolutionarily conserved and have distinct Chou-Fasman propensities and interaction patterns compared to other hotspots. Finally, we present detailed case studies to show that the interaction network formed by the secondary shell hotspots is crucial for complex stability and activity. Further, they act as potentially allosteric propagators and bridge interfacial and non-interfacial sites in the protein. Their mutations to any other amino acid types cause significant destabilization. Overall, this study sheds light on the uniqueness and importance of secondary shell hotspots in protein-protein complexes.

bioinformatics↗