Search bioRxivSearch

Biology subjects

Quinlan, A.

Publications and source records attributed to Quinlan, A..

4 recordsLinked to original sources

mosdepth: quick coverage calculation for genomes and exomes

Mosdepth is a new command-line tool for rapidly calculating genome-wide sequencing coverage. It measures depth from BAM (Li et al. 2009) or CRAM files at either each nucleotide position in a genome or for sets of genomic regions. Genomic regions may be specified as either a BED file to evaluate coverage across capture regions, or as a fixed-size window as required for copy-number calling. Mosdepth uses a simple algorithm that is computationally efficient and enables it to quickly produce optional coverage summaries. We demonstrate that mosdepth is faster than existing tools and provides flexibility in the types of coverage profiles produced.\n\nAvailabilitymosdepth is available from https://github.com/brentp/mosdepth under the MIT license..\n\nContactbpederse@gmail.com, aaronquinlan@gmail.com\n\nSupplementary informationDetailed documentation is available at https://github.com/brentp/mosdepth

bioinformatics

Fine-mapping identifies causal variants for RA and T1D in DNASE1L3, SIRPG, MEG3, TNFAIP3 and CD28/CTLA4 loci

We fine-mapped 76 rheumatoid arthritis (RA) and type 1 diabetes (T1D) loci outside of the MHC. After sequencing 799 1kb regulatory (H3K4me3) regions within these loci in 568 individuals, we observed accurate imputation for 89% of common variants. We fine-mapped1,2 these loci in RA (11,475 cases, 15,870 controls)3, T1D (9,334 cases and 11,111 controls) 4 and combined datasets. We reduced the number of potential causal variants to [≤]5 in 8 RA and 11 T1D loci. We identified causal missense variants in five loci (DNASE1L3, SIRPG, PTPN22, SH2B3 and TYK2) and likely causal non-coding variants in six loci (MEG3, TNFAIP3, CD28/CTLA4, ANKRD55, IL2RA, REL/PUS10). Functional analysis confirmed allele specific binding and differential enhancer activity for three variants: the CD28/CTLA4 rs117701653 SNP, the TNFAIP3 rs35926684 indel, and the MEG3 rs34552516 indel. This study demonstrates the potential for dense genotyping and imputation to pinpoint missense and non-coding causal alleles.

genetics

Nanopore sequencing and assembly of a human genome with ultra-long reads

Nanopore sequencing is a promising technique for genome sequencing due to its portability, ability to sequence long reads from single molecules, and to simultaneously assay DNA methylation. However until recently nanopore sequencing has been mainly applied to small genomes, due to the limited output attainable. We present nanopore sequencing and assembly of the GM12878 Utah/Ceph human reference genome generated using the Oxford Nanopore MinION and R9.4 version chemistry. We generated 91.2 Gb of sequence data ([~]30x theoretical coverage) from 39 flowcells. De novo assembly yielded a highly complete and contiguous assembly (NG50 [~]3Mb). We observed considerable variability in homopolymeric tract resolution between different basecallers. The data permitted sensitive detection of both large structural variants and epigenetic modifications. Further we developed a new approach exploiting the long-read capability of this system and found that adding an additional 5x-coverage of ultra-long reads (read N50 of 99.7kb) more than doubled the assembly contiguity. Modelling the repeat structure of the human genome predicts extraordinarily contiguous assemblies may be possible using nanopore reads alone. Portable de novo sequencing of human genomes may be important for rapid point-of-care diagnosis of rare genetic diseases and cancer, and monitoring of cancer progression. The complete dataset including raw signal is available as an Amazon Web Services Open Dataset at: https://github.com/nanopore-wgs-consortium/NA12878.

genomics

Limited contribution of rare, noncoding variation to autism spectrum disorder from sequencing of 2,076 genomes in quartet families

Genomic studies to date in autism spectrum disorder (ASD) have largely focused on newly arising mutations that disrupt protein coding sequence and strongly influence risk. We evaluate the contribution of noncoding regulatory variation across the size and frequency spectrum through whole genome sequencing of 519 ASD cases, their unaffected sibling controls, and parents. Cases carry a small excess of de novo (1.02-fold) noncoding variants, which is not significant after correcting for paternal age. Assessing 51,801 regulatory classes, no category is significantly associated with ASD after correction for multiple testing. The strongest signals are observed in coding regions, including structural variation not detected by previous technologies and missense variation. While rare noncoding variation likely contributes to risk in neurodevelopmental disorders, no category of variation has impact equivalent to loss-of-function mutations. Average effect sizes are likely to be smaller than that for coding variation, requiring substantially larger samples to quantify this risk.

genomics