Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,351 records · Page 75Linked to original sources

Ultra-sensitive mutation detection and genome-wide DNA copy number reconstruction by error corrected circulating tumour DNA sequencing

Minimally invasive circulating free DNA (cfDNA) analysis can portray cancer genome landscapes but highly sensitive and specific genetic approaches are necessary to accurately detect mutations with often low variant frequencies. We developed a targeted cfDNA sequencing technology using novel off-the-shelf molecular barcodes for error correction, in combination with custom solution hybrid capture enrichment. Modelling based on cfDNA yields from 58 patients shows that our assay, which requires 25ng of cfDNA input, should be applicable to >95% of patients with metastatic colorectal cancer. Sequencing of a 163.3 kb target region including 32 genes detected 100% of single nucleotide variants with 0.15% variant frequency in cfDNA spike-in experiments. Molecular barcode error correction reduced false positive mutation calls by 98.6%. In a series of 28 patients with metastatic colorectal cancers, 80 out of 91 (88%) mutations previously detected by tumour tissue sequencing were called in the cfDNA. Call rates were similar for single nucleotide variants and small insertions/deletions. Mutations only called in cfDNA but not detectable in matched tumour tissue included, among others, a subclonal resistance driver mutation to anti-EGFR antibodies in the KRAS gene, multiple activating PIK3CA mutations in each of two patients (indicative of parallel evolution), and TP53 mutations originating from clonal haematopoiesis. Furthermore, we demonstrate that cfDNA off-target read analysis allows the reconstruction of genome wide copy number aberration profiles from 71% of these 28 cases. This error-corrected ultra-deep cfDNA sequencing assay with a target region that can be readily customized enables broad insights into cancer genomes and evolution.

genomics

Genome-Wide Identification of Early-Firing Human Replication Origins by Optical Replication Mapping

The timing of DNA replication is largely regulated by the location and timing of replication origin firing. Therefore, much effort has been invested in identifying and analyzing human replication origins. However, the heterogeneous nature of eukaryotic replication kinetics and the low efficiency of individual origins in metazoans has made mapping the location and timing of replication initiation in human cells difficult. We have mapped early-firing origins in HeLa cells using Optical Replication Mapping, a high-throughput single-molecule approach based on Bionano Genomics genomic mapping technology. The single-molecule nature and 290-fold coverage of our dataset allowed us to identify origins that fire with as little as 1% efficiency. We find sites of human replication initiation in early S phase are not confined to well-defined efficient replication origins, but are instead distributed across broad initiation zones consisting of many inefficient origins. These early-firing initiation zones co-localize with initiation zones inferred from Okazaki-fragment-mapping analysis and are enriched in ORC1 binding sites. Although most early-firing origins fire in early-replication regions of the genome, a significant number fire in late-replicating regions, suggesting that the major difference between origins in early and late replicating regions is their probability of firing in early S-phase, as opposed to qualitative differences in their firing-time distributions. This observation is consistent with stochastic models of origin timing regulation, which explain the regulation of replication timing in yeast.

genomics

Genome-wide polygenic score to identify a monogenic risk-equivalent for coronary disease

Identification of individuals at increased genetic risk for a complex disorder such as coronary disease can facilitate treatments or enhanced screening strategies. A rare monogenic mutation associated with increased cholesterol is present in ~1:250 carriers and confers an up to 4-fold increase in coronary risk when compared with non-carriers. Although individual common polymorphisms have modest predictive capacity, their cumulative impact can be aggregated into a polygenic score. Here, we develop a new, genome-wide polygenic score that aggregates information from 6.6 million common polymorphisms and show that this score can similarly identify individuals with a 4-fold increased risk for coronary disease. In >400,000 participants from UK Biobank, the score conforms to a normal distribution and those in the top 2.5% of the distribution are at 4-fold increased risk compared to the remaining 97.5%. Similar patterns are observed with genome-wide polygenic scores for two additional diseases - breast cancer and severe obesity.\n\nOne Sentence SummaryA genome-wide polygenic score identifies 2.5% of the population born with a 4-fold increased risk for coronary artery disease.

genomics

The Evolutionary Genomic Dynamics of Peruvians Before, During, and After the Inca Empire

Native Americans from the Amazon, Andes, and coast regions of South America have a rich cultural heritage, but have been genetically understudied leading to gaps in our knowledge of their genomic architecture and demographic history. Here, we sequenced 150 high-coverage and genotyped 130 genomes from Native American and mestizo populations in Peru. A majority of our samples possess greater than 90% Native American ancestry and demographic modeling reveals, consistent with a rapid peopling model of the Americas, that most of Peru was peopled approximately 12,000 years ago. While the Native American populations possessed distinct ancestral divisions, the mestizo groups were admixtures of multiple Native American communities which occurred before and during the Inca Empire. The mestizo communities also show Spanish introgression only after Peruvian Independence. Thus, we present a detailed model of the evolutionary dynamics which impacted the genomes of modern day Peruvians.

genomics

Genome variants altering gene splicing in bovine are extensively shared between tissues

BackgroundMammalian phenotypes are shaped by numerous genome variants, many of which may regulate gene transcription or RNA splicing. To identify variants with regulatory functions in cattle, an important economic and model species, we used sequence variants to map a type of expression quantitative trait loci (expression QTLs) that are associated with variations in the RNA splicing, i.e., sQTLs. To further the understanding of regulatory variants, sQTLs were compare with other two types of expression QTLs, 1) variants associated with variations in gene expression, i.e., geQTLs and 2) variants associated with variations in exon expression, i.e., eeQTLs, in different tissues.\n\nResultsUsing whole genome and RNA sequence data from four tissues of over 200 cattle, sQTLs identified using exon inclusion ratios were verified by matching their effects on adjacent intron excision ratios. sQTLs contained the highest percentage of variants that are within the intronic region of genes and contained the lowest percentage of variants that are within intergenic regions, compared to eeQTLs and geQTLs. Many geQTLs and sQTLs are also detected as eeQTLs. Many expression QTLs, including sQTLs, were significant in all four tissues and had a similar effect in each tissue. To verify such expression QTL sharing between tissues, variants surrounding ({+/-}1Mb) the exon or gene were used to build local genomic relationship matrices (LGRM) and estimated genetic correlations between tissues. For many exons, the splicing and expression level was determined by the same cis additive genetic variance in different tissues. Thus, an effective but simple-to-implement meta-analysis combining information from three tissues is introduced to increase power to detect and validate sQTLs. sQTLs and eeQTLs together were more enriched for variants associated with cattle complex traits, compared to geQTLs. Several putative causal mutations were identified, including an sQTL at Chr6:87392580 within the 5th exon of kappa casein (CSN3) associated with milk production traits.\n\nConclusionsUsing novel analytical approaches, we report the first identification of numerous bovine sQTLs which are extensively shared between multiple tissue types. The significant overlaps between bovine sQTLs and complex traits QTL highlight the contribution of regulatory mutations to phenotypic variations.

genomics

Genome-wide detection of genes under positive selection in worldwide populations of the barley scald pathogen

The coevolution between hosts and pathogens generates strong selection pressures to maintain resistance and infectivity, respectively. Genomes of plant pathogens often encode major effect loci for the ability to successfully infect a specific host. Hence, heterogeneity in the host genotypes and abiotic factors leads to locally adapted pathogen populations. However, the genetic basis of local adaptation is poorly understood. We analyzed global field populations of Rhynchosporium commune, the pathogen causing barley scald disease, to identify candidate genes for local adaptation. Whole genome sequencing data generated for 125 isolates showed that the pathogen is subdivided into three genetic clusters associated with distinct geographic and climatic regions. Using haplotype-based selection scans applied independently to each genetic cluster, we found strong evidence for selective sweeps throughout the genome. Comparisons of loci under selection among clusters revealed little overlap, suggesting that ecological differences associated with each cluster led to variable selection regimes. The strongest signals of selection were found predominantly in the two clusters composed of isolates from Central Europe and Ethiopia. The strongest selective sweep regions encoded proteins with functions related to both biotic and abiotic stresses. We found that selective sweep regions were enriched in genes encoding functions in cellular localization, protein transport activity, and DNA damage responses. In contrast to the prevailing view that a small number of gene-for-gene interactions govern plant pathogen evolution, our analyses suggest that the evolutionary trajectory is largely determined by spatially heterogeneous biotic and abiotic selection pressures.

genomics

Deep coverage whole genome sequences and plasma lipoprotein(a) in individuals of European and African ancestries

Lipoprotein(a), Lp(a), is a modified low-density lipoprotein particle where apolipoprotein(a) (protein product of the LPA gene) is covalently attached to apolipoprotein B. Lp(a) is a highly heritable, causal risk factor for cardiovascular diseases and varies in concentrations across ancestries. To comprehensively delineate the inherited basis for plasma Lp(a), we performed deep-coverage whole genome sequencing in 8,392 individuals of European and African American ancestries. Through whole genome variant discovery and direct genotyping of all structural variants overlapping LPA, we quantified the 5.5kb kringle IV-2 copy number (KIV2-CN), a known LPA structural polymorphism, and developed a model for its imputation. Through common variant analysis, we discovered a novel locus (SORT1) associated with Lp(a)-cholesterol, and also genetic modifiers of KIV2-CN. Furthermore, in contrast to previous GWAS studies, we explain most of the heritability of Lp(a), observing Lp(a) to be 85% heritable among African Americans and 75% among Europeans, yet with notable inter-ethnic heterogeneity. Through analyses of aggregates of rare coding and non-coding variants with Lp(a)-cholesterol, we found the only genome-wide significant signal to be at a non-coding SLC22A3 intronic window also previously described to be associated with Lp(a); however, this association was mitigated by adjustment with KIV2-CN. Finally, using an additional imputation dataset (N=27,344), we performed Mendelian randomization of LPA variant classes, finding that genetically regulated Lp(a) is more strongly associated with incident cardiovascular diseases than directly measured Lp(a), and is significantly associated with measures of subclinical atherosclerosis in African Americans.

genomics

Population genomics of hypervirulent Klebsiella pneumoniae clonal group 23 reveals early emergence and rapid global dissemination

Since the mid-1980s there have been increasing reports of severe community-acquired pyogenic liver abscess, meningitis and bloodstream infections caused by hypervirulent Klebsiella pneumoniae, predominantly encompassing clonal group (CG) 23 serotype K1 strains. Common features of CG23 include a virulence plasmid associated with iron scavenging and hypermucoidy, and a chromosomal integrative and conjugative element (ICE) encoding the siderophore yersiniabactin and the genotoxin colibactin. Here we investigate the evolutionary history and genomic diversity of CG23 based on comparative analysis of 98 genomes. Contrary to previous reports with more limited samples, we show that CG23 comprises several deep branching sublineages dating back to the 1870s, many of which are associated with distinct chromosomal insertions of ICEs encoding yersiniabactin. We find that most liver abscess isolates (>80%) belong to a dominant sublineage, CG23-I, which emerged in the 1920s following acquisition of ICEKp10 (encoding colibactin in addition to yersiniabactin) and has undergone clonal expansion and global dissemination within the human population. The unique genomic feature of CG23-I is the production of colibactin, which has been reported previously as a promoter of gut colonisation and dissemination to the liver and brain in a mouse model of CG23 K. pneumoniae infection, and has been linked to colorectal cancer. We also identify an antibiotic-resistant subclade of CG23-I associated with sexually-transmitted infections in horses dating back to the 1980s. These data show that hypervirulent CG23 K. pneumoniae was circulating in humans for decades before the liver abscess epidemic was first recognised, and has the capacity to acquire and maintain AMR plasmids. These data provide a framework for future epidemiological and experimental studies of hypervirulent K. pneumoniae. To further support such studies we present an open access and completely sequenced human liver abscess isolate, SGH10, which is typical of the globally disseminated CG23-I sublineage.

genomics

BLINK: A Package for Next Level of Genome Wide Association Studies with Both Individuals and Markers in Millions

Big data, accumulated from biomedical and agronomic studies, provides the potential to identify genes controlling complex human diseases and agriculturally important traits through genome-wide association studies (GWAS). However, big data also leads to extreme computational challenges, especially when sophisticated statistical models are employed to simultaneously reduce false positives and false negatives. The newly developed Fixed and random model Circulating Probability Unification (FarmCPU) method uses a bin method under the assumption that Quantitative Trait Nucleotides (QTNs) are evenly distributed throughout the genome. The estimated QTNs are used to separate a mixed linear model into a computationally efficient fixed effect model (FEM) and a computationally expensive random effect model (REM), which are then used iteratively. To completely eliminate the computationally expensive REM, we replaced REM with FEM by using Bayesian information criteria. To eliminate the requirement that QTNs be evenly distributed throughout the genome, we replaced the bin method with linkage disequilibrium information. The new method is called Bayesian-information and Linkage-disequilibrium Iteratively Nested Keyway (BLINK). Both real and simulated data analyses demonstrated that BLINK improves statistical power compared to FarmCPU, in addition to a remarkable improvement in computing time. Now, a dataset with half million markers and one million individuals can be analyzed within five hours, compared with one week using FarmCPU.

genomics

Multiplex Confounding Factor Correction for Genomic Association Mapping with Squared Sparse Linear Mixed Model

Genome-wide Association Study has presented a promising way to understand the association between human genomes and complex traits. Many simple polymorphic loci have been shown to explain a significant fraction of phenotypic variability. However, challenges remain in the non-triviality of explaining complex traits associated with multifactorial genetic loci, especially considering the confounding factors caused by population structure, family structure, and cryptic relatedness. In this paper, we propose a Squared-LMM (LMM2) model, aiming to jointly correct population and genetic confounding factors. We offer two strategies of utilizing LMM2 for association mapping: 1) It serves as an extension of univariate LMM, which could effectively correct population structure, but consider each SNP in isolation. 2) It is integrated with the multivariate regression model to discover association relationship between complex traits and multifactorial genetic loci. We refer to this second model as sparse Squared-LMM (sLMM2). Further, we extend LMM2/sLMM2 by raising the power of our squared model to the LMMn/sLMMn model. We demonstrate the practical use of our model with synthetic phenotypic variants generated from genetic loci of Arabidopsis Thaliana. The experiment shows that our method achieves a more accurate and significant prediction on the association relationship between traits and loci. We also evaluate our models on collected phenotypes and genotypes with the number of candidate genes that the models could discover. The results suggest the potential and promising usage of our method in genome-wide association studies.

genomics

Correcting for population stratification reduces false positive and false negative results in joint analyses of host and pathogen genomes

Studies of host genetic determinants of pathogen sequence variation can identify sites of genomic conflicts, by highlighting variants that are implicated in immune response on the host side and adaptive escape on the pathogen side. However, systematic genetic differences in host and pathogen populations can lead to inflated type I (false positive) and type II (false negative) error rates in genome-wide association analyses. Here, we demonstrate through simulation that correcting for both host and pathogen stratification reduces spurious signals and increases power to detect real associations in a variety of tested scenarios. We confirm the validity of the simulations by showing comparable results in an analysis of paired human and HIV genomes.

genomics

Structural disruption of genomic regions containing ultraconserved elements is associated with neurodevelopmental phenotypes

The development of the human brain and nervous system can be affected by genetic or environmental factors. Here we focus on characterizing the genetic perturbations that accompany and may contribute to neurodevelopmental phenotypes. Specifically, we examine two types of structural variants, namely, copy number variation and balanced chromosome rearrangements, discovered in subjects with neurodevelopmental disorders and related phenotypes. We find that a feature uniting these types of genetic aberrations is a proximity to ultraconserved elements (UCEs), which are sequences that are perfectly conserved between the reference genomes of distantly related species. In particular, while UCEs are generally depleted from copy number variant regions in healthy individuals, they are, on the whole, enriched in genomic regions disrupted by copy number variants or breakpoints of balanced rearrangements in affected individuals. Additionally, while genes associated with neurodevelopmental disorders are enriched in UCEs, this does not account for the excess of UCEs either in copy number variants or close to the breakpoints of balanced rearrangements in affected individuals. Indeed, our data are consistent with some manifestations of neurodevelopmental disorders resulting from a disruption of genome integrity in the vicinity of UCEs.

genomics

The genomic ancestry, landscape genetics, and invasion history of introduced mice in New Zealand

1. SummaryThe house mouse (Mus musculus) provides a fascinating system for studying both the genomic basis of reproductive isolation, and the patterns of human-mediated dispersal. New Zealand has a complex history of mouse invasions, and the living descendants of these invaders have genetic ancestry from all three subspecies, although most are primarily descended from M. m. domesticus. We used the GigaMUGA genotyping array (~135,000 loci) to describe the genomic ancestry of 161 mice, sampled from 34 locations from across New Zealand (and one Australian city - Sydney). Of these, two populations, one in the south of the South Island, and one on Chatham Island, showed complete mitochondrial lineage capture, featuring two different lineages of M. m. castaneus mitochondrial DNA but with only M. m. domesticus nuclear ancestry detectable. Mice in the northern and southern parts of the North Island had small traces (~2-3%) of M. m. castaneus nuclear ancestry, and mice in the upper South Island had ~7-8% M. m. musculus nuclear ancestry including some Y-chromosomal ancestry - though no detectable M. m. musculus mitochondrial ancestry. This is the most thorough genomic study of introduced populations of house mice yet conducted, and will have relevance to studies of the isolation mechanisms separating subspecies of mice.

genomics

Improved Aedes aegypti mosquito reference genome assembly enables biological discovery and vector control

Female Aedes aegypti mosquitoes infect hundreds of millions of people each year with dangerous viral pathogens including dengue, yellow fever, Zika, and chikungunya. Progress in understanding the biology of this insect, and developing tools to fight it, has been slowed by the lack of a high-quality genome assembly. Here we combine diverse genome technologies to produce AaegL5, a dramatically improved and annotated assembly, and demonstrate how it accelerates mosquito science and control. We anchored the physical and cytogenetic maps, resolved the size and composition of the elusive sex-determining \"M locus\", significantly increased the known members of the glutathione-S-transferase genes important for insecticide resistance, and doubled the number of chemosensory ionotropic receptors that guide mosquitoes to human hosts and egg-laying sites. Using high-resolution QTL and population genomic analyses, we mapped new candidates for dengue vector competence and insecticide resistance. We predict that AaegL5 will catalyse new biological insights and intervention strategies to fight this deadly arboviral vector.

genomics

Easy Hi-C: A simple efficient protocol for 3D genome mapping in small cell populations

Despite the growing interest in studying the mammalian genome organization, it is still challenging to map the DNA contacts genome-wide. Here we present easy Hi-C (eHi-C), a highly efficient method for unbiased mapping of 3D genome architecture. The eHi-C protocol only involves a series of enzymatic reactions and maximizes the recovery of DNA products from proximity ligation. We show that eHi-C can be performed with 0.1 million cells and yields high quality libraries comparable to Hi-C.

genomics

An atlas of silencer elements for the human and mouse genomes

The study of gene regulation is dominated by a focus on the control of gene activation or controlling an increase in the level of expression. Just as critical is the process of gene repression or silencing. Chromatin signatures have allowed for the global mapping of enhancer cis-regulatory elements, however, the identification of silencer elements by computational or experimental approaches in a genome-wide manner are lacking. We present a simple but powerful computational approach to identify putative silencers genome-wide. We used a series of consortia data to predict silencers in over 100 human and mouse cell or tissue types. We performed several analyses to determine if these elements exhibited characteristics expected of a silencers. Motif enrichment analyses on putative silencers determined that motifs belonging to known transcriptional repressors are enriched, as well as overlapping known transcription repressor binding sites. Leveraging promoter capture HiC data from several human and mouse cell types, we found that over 50% of putative silencer elements are interacting with gene promoters having very low to no expression. Next, to validate our silencer predictions, we quantified silencer activity using massively parallel reporter assays (MPRAs) on 7500 selected elements in K562 cells. We trained a support vector machine model classifier on MPRA data and used it to refine potential silencers in other cell types. We also show that similar to enhancer elements, silencer elements are enriched in disease-associated variants. Our results suggest a general strategy for genome-wide identification and characterization of silencer elements.

genomics

Whole genome sequencing reveals the emergence of a Pseudomonas aeruginosa shared strain sub-lineage among patients treated within a single cystic fibrosis centre

BackgroundChronic lung infections by Pseudomonas aeruginosa are a significant cause of morbidity and mortality in people with cystic fibrosis (CF). Shared P. aeruginosa strains, that can be transmitted between patients, are of concern and in Australia the AUST-02 shared strain is predominant in individuals attending CF centres in Queensland and Western Australia. M3L7 is a multidrug resistant sub-type of AUST-02 that was recently identified in a Queensland CF centre and was shown to be associated with poorer clinical outcomes. The main aim of this study was to resolve the relationship of the emergent M3L7 sub-type within the AUST-02 group of strains using whole genome sequencing.\n\nResultsA whole-genome core phylogeny of 63 isolates indicated that M3L7 is a monophyletic sub-lineage within the context of the broader AUST-02 group. Relatively short branch lengths connected all of the M3L7 isolates. A phylogeny based on nucleotide polymorphisms present across the genome showed that the chronological estimation of the most recent common ancestor was around 2001 ({+/-} 3 years). SNP differences between sequential M3L7 isolates collected 3-4 years apart from five patients suggested both continuous infection of the same strain and cross-infection of some M3L7 variants between patients. The majority of polymorphisms that were characteristic of M3L7 (i.e. acquired after divergence from all other AUST-02 isolates sequenced) were found to produce non-synonymous mutations in virulence and antibiotic resistance genes.\n\nConclusionsM3L7 has recently diverged from a common ancestor indicating descent from a single carrier at a CF treatment centre in Australia. Both adaptation to the lung and transmission of M3L7 between adults attending this centre may have contributed to its rapid dissemination. The study emphasises the importance of clinical management in controlling the emergence of shared strains in CF.

genomics

Generating genomic platforms to study Candida albicans pathogenesis

The advent of the genomic era has made elucidating gene function at large scale a pressing challenge. ORFeome collections, whereby almost all ORFs of a given species are cloned and can be subsequently leveraged in multiple functional genomic approaches, represent valuable resources towards this endeavor. Here we provide novel, genome-scale tools for the study of Candida albicans, a commensal yeast that is also responsible for frequent superficial and disseminated infections in humans. We have generated an ORFeome collection composed of 5,102 ORFs cloned in a Gateway donor vector, representing 83% of the currently annotated coding sequences of C. albicans. Sequencing data of the cloned ORFs are available in the CandidaOrfDB database at http://candidaorfeome.eu. We also engineered 49 expression vectors with a choice of promoters, tags, and selection markers and demonstrated their applicability to the study of target ORFs transferred from the C. albicans ORFeome. In addition, the use of the ORFeome in the detection of protein-protein interaction was demonstrated. Mating-compatible strains as well as Gateway-compatible two-hybrid vectors were engineered, validated and used in a proof of concept experiment. These unique and valuable resources should greatly facilitate future functional studies in C. albicans and the elucidation of mechanisms that underlie its pathogenicity.

genomics