Search bioRxivSearch

Biology subjects

Turgut, D.

Publications and source records attributed to Turgut, D..

3 recordsLinked to original sources

Population-specific genome graphs improve high-throughput sequencing data analysis: A case study on the Pan-African genome

Graph-based genome reference representations have seen significant development, motivated by the inadequacy of the current human genome reference to represent the diverse genetic information from different human populations and its inability to maintain the same level of accuracy for non-European ancestries. While there have been many efforts to develop computationally efficient graph-based toolkits for NGS read alignment and variant calling, methods to curate genomic variants and subsequently construct genome graphs remains an understudied problem that inevitably determines the effectiveness of the overall bioinformatics pipeline. In this study, we discuss obstacles encountered during graph construction and propose methods for sample selection based on population diversity, graph augmentation with structural variants and resolution of graph reference ambiguity caused by information overload. Moreover, we present the case for iteratively augmenting tailored genome graphs for targeted populations and demonstrate this approach on the whole-genome samples of African ancestry. Our results show that population-specific graphs, as more representative alternatives to linear or generic graph references, can achieve significantly lower read mapping errors and enhanced variant calling sensitivity, in addition to providing the improvements of joint variant calling without the need of computationally intensive post-processing steps.

bioinformatics

precisionFDA Truth Challenge V2: Calling variants from short- and long-reads in difficult-to-map regions

The precisionFDA Truth Challenge V2 aimed to assess the state-of-the-art of variant calling in difficult-to-map regions and the Major Histocompatibility Complex (MHC). Starting with FASTQ files, 20 challenge participants applied their variant calling pipelines and submitted 64 variant callsets for one or more sequencing technologies (~35X Illumina, ~35X PacBio HiFi, and ~50X Oxford Nanopore Technologies). Submissions were evaluated following best practices for benchmarking small variants with the new GIAB benchmark sets and genome stratifications. Challenge submissions included a number of innovative methods for all three technologies, with graph-based and machine-learning methods scoring best for short-read and long-read datasets, respectively. New methods out-performed the 2016 Truth Challenge winners, and new machine-learning approaches combining multiple sequencing technologies performed particularly well. Recent developments in sequencing and variant calling have enabled benchmarking variants in challenging genomic regions, paving the way for the identification of previously unknown clinically relevant variants.

bioinformatics

Could network structures generated with simple rules imposed on a cubic lattice reproduce the structural descriptors of globular proteins?

A direct way to spot structural features that are universally shared among proteins is to find proper analogues from simpler condensed matter systems. In most cases, sphere-packing arguments provide a straightforward route for structural comparison, as they successfully characterize a wide array of materials such as close packed crystals, dense liquids, and structural glasses. In the current study, the feasibility of creating ensembles of artificial structures that can automatically reproduce a large number of geometrical and topological descriptors of globular proteins is investigated. Towards this aim, a simple cubic (SC) arrangement is shown to provide the best background lattice after a careful analysis of the residue packing trends from 210 proteins. It is shown that a minimalistic set of ground rules imposed on this lattice is sufficient to generate structures that can mimic real proteins. In the proposed method, 210 such structures are generated by randomly removing residues (beads) from clusters that have a SC lattice arrangement until a predetermined residue concentration is achieved. All generated structures are checked for residue connectivity such that a path exists between any two residues. Two additional sets are prepared from the initial structures via random relaxation and a reverse Monte Carlo simulated annealing (RMC-SA) algorithm, which targets the average radial distribution function (RDF) of 210 globular proteins. The initial and relaxed structures are compared to real proteins via RDF, bond orientational order parameters, and several descriptors of network topology. Based on these features, results indicate that the structures generated with 40% occupancy via the proposed method closely resemble real residue networks. The broad correspondence established this way indicates a non-superficial link between the residue networks and the defect laden cubic crystalline order. The presented approach of identifying a minimalistic set of operations performed on a target lattice such that each resulting cluster possess structural characteristics largely indistinguishable from that of a coarse-grained globular protein opens up new venues in structural characterization, native state recognition, and rational design of proteins.

biophysics