Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,675 records · Page 93Linked to original sources

Reduced representation sequencing for symbiotic anthozoans: are reference genomes necessary to eliminate endosymbiont contamination and make robust phylogeographic inference?

Anthozoan cnidarians form the backbone of coral reefs. Their success relies on endosymbiosis with photosynthetic dinoflagellates in the family Symbiodiniaceae. Photosymbionts represent a hurdle for researchers using population genomic techniques to study these highly imperiled and ecologically critical species because sequencing datasets harbor unknown mixtures of anthozoan and photosymbiont loci. Here we use range-wide sampling and a double-digest restriction-site associated DNA sequencing (ddRADseq) of the sea anemone Bartholomea annulata to explore how symbiont loci impact the interpretation of phylogeographic patterns and population genetic parameters. We use the genome of the closely related Exaiptasia diaphana (previously Aiptasia pallida) to create an anthozoan-only dataset from a genomic dataset containing both B. annulata and its symbiodiniacean symbionts and then compare this to the raw, holobiont dataset. For each, we investigate spatial patterns of genetic diversity and use coalescent model-based approaches to estimate demographic history and population parameters. The Florida Straits are the only phylogeographic break we recover for B. annulata, with divergence estimated during the last glacial maximum. Because B. annulata hosts multiple members of Symbiodiniaceae, we hypothesize that, under moderate missing data thresholds, de novo clustering algorithms that identify orthologs across datasets will have difficulty identifying shared non-coding loci from the photosymbionts. We infer that, for anthozoans hosting diverse members of Symbiodinaceae, clustering algorithms act as de facto filters of symbiont loci. Thus, while at least some photosymbiont loci remain, these are swamped by orders of magnitude greater numbers of anthozoan loci and thus represent genetic \"noise,\" rather than contributing genetic signal.

genomics

Capturing the dynamics of genome replication on individual ultra-long nanopore sequence reads

The replication of eukaryotic genomes is highly stochastic, making it difficult to determine the replication dynamics of individual molecules with existing methods. We now report a sequencing method for the measurement of replication fork movement on single molecules by Detecting Nucleotide Analogue signal currents on extremely long nanopore traces (D-NAscent). Using this method, we detect BrdU incorporated by Saccharomyces cerevisiae to reveal, at a genomic scale and on single molecules, the DNA sequences replicated during a pulse labelling period. Under conditions of limiting BrdU concentration, D-NAscent detects the differences in BrdU incorporation frequency across individual molecules to reveal the location of active replication origins, fork direction, termination sites, and fork pausing/stalling events. We used sequencing reads of 20-160 kb, to generate the first whole genome single-molecule map of DNA replication dynamics and discover a new class of low frequency stochastic origins in budding yeast.

genomics

The jellyfish genome sheds light on the early evolution of active predation

BackgroundUnique among cnidarians, jellyfish have remarkable morphological and biochemical innovations that allow them to actively hunt in the water column. One of the first animals to become free-swimming, jellyfish employ pulsed jet propulsion and venomous tentacles to capture prey.\n\nResultsTo understand these key innovations, we sequenced the genome of the giant Nomuras jellyfish (Nemopilema nomurai), the transcriptomes of its bell and tentacles, and transcriptomes across tissues and developmental stages of the Sanderia malayensis jellyfish. Analyses of Nemopilema and other cnidarian genomes revealed adaptations associated with swimming, marked by codon bias in muscle contraction and expansion of neurotransmitter genes, along with expanded Myosin type II family and venom domains; possibly contributing to jellyfish mobility and active predation. We also identified gene family expansions of Wnt and posterior Hox genes, and discovered the important role of retinoic acid signaling in this ancient lineage of metazoans, which together may be related to the unique jellyfish body plan (medusa formation).\n\nConclusionsTaken together, the jellyfish genome and transcriptomes genetically confirm their unique morphological and physiological traits that have combined to make these animals one of the worlds earliest and most successful multi-cellular predators.

genomics

Physiological genomics of dietary adaptation in a marine herbivorous fish

Adopting a new diet is a significant evolutionary change and can profoundly affect an animals physiology, biochemistry, ecology, and its genome. To study this evolutionary transition, we investigated the physiology and genomics of digestion of a derived herbivorous fish, the monkeyface prickleback (Cebidichthys violaceus). We sequenced and assembled its genome and digestive transcriptome and revealed the molecular changes related to important dietary enzymes, finding abundant evidence for adaptation at the molecular level. In this species, two gene families experienced expansion in copy number and adaptive amino acid substitutions. These families, amylase, and bile salt activated lipase, are involved digestion of carbohydrates and lipids, respectively. Both show elevated levels of gene expression and increased enzyme activity. Because carbohydrates are abundant in the pricklebacks diet and lipids are rare, these findings suggest that such dietary specialization involves both exploiting abundant resources and scavenging rare ones, especially essential nutrients, like essential fatty acids.

genomics

Integrating Culture-based Antibiotic Resistance Profiles with Whole-genome Sequencing Data for 11,087 Clinical Isolates

Emerging antibiotic resistance is a major global health threat. The analysis of nucleic acid sequences linked to susceptibility phenotypes facilitates the study of genetic antibiotic resistance determinants to inform molecular diagnostics and drug development. We collected genetic data (11,087 newly sequenced whole genomes) and culture-based resistance profiles (10,991 of 11,087 isolates were comprehensively tested against 22 antibiotics in total) of clinical isolates including 18 main species spanning a time period of 30 years. Species and drug specific resistance patterns could be observed including increasing resistance rates for Acinetobacter baumannii to carbapenems and for Escherichia coli to fluoroquinolones. Species-level pan-genomes were constructed to reflect the genetic repertoire of the respective species such as conserved essential genes and known resistance factors. Integrating phenotypes and genotypes through species-level pan-genomes allowed to infer gene-drug resistance associations using statistical testing. The isolate collection and the analysis results have been integrated into a resource, GEAR-base, available for academic research use free of charge at https://gear-base.com.

genomics

A genome-wide algal mutant library reveals a global view of genes required for eukaryotic photosynthesis

Photosynthetic organisms provide food and energy for nearly all life on Earth, yet half of their protein-coding genes remain uncharacterized1,2. Characterization of these genes could be greatly accelerated by new genetic resources for unicellular organisms that complement the use of multicellular plants by enabling higher-throughput studies. Here, we generated a genome-wide, indexed library of mapped insertion mutants for the flagship unicellular alga Chlamydomonas reinhardtii (Chlamydomonas hereafter). The 62,389 mutants in the library, covering 83% of nuclear, protein-coding genes, are available to the community. Each mutant contains unique DNA barcodes, allowing the collection to be screened as a pool. We leveraged this feature to perform a genome-wide survey of genes required for photosynthesis, which identified 303 candidate genes. Characterization of one of these genes, the conserved predicted phosphatase CPL3, showed it is important for accumulation of multiple photosynthetic protein complexes. Strikingly, 21 of the 43 highest-confidence genes are novel, opening new opportunities for advances in our understanding of this biogeochemically fundamental process. This library is the first genome-wide mapped mutant resource in any unicellular photosynthetic organism, and will accelerate the characterization of thousands of genes in algae, plants and animals.

genomics

Population genomic analyses reveal a highly differentiated genetic cluster of northern goshawks (Accipiter gentilis laingi) in Haida Gwaii

Accurate knowledge of geographic ranges and genetic relationships among populations is important when managing a species or population of conservation concern. In the western Canadian province of British Columbia, a subspecies of the northern goshawk (Accipiter gentilis laingi) is designated as Threatened under the Canadian Species at Risk Act. Historically, the range of this bird of prey has been ambiguous and its genetic distinctness from the other North American subspecies (Accipiter gentilis atricapillus) has not been well established. Given the uncertainty in using morphological traits to assign individual goshawks to these two subspecies, we analyzed genomic relationships in tens of thousands of single nucleotide polymorphisms identified using genotyping-bysequencing of high-quality genetic samples. This genome-wide analysis revealed a genetically distinct population of northern goshawks on the archipelago of Haida Gwaii and subtle genetic structuring among the remainder of our sampling sites within North America. Following from this analysis, we developed targeted genotyping assays for ten loci that are highly differentiated between the two main genetic clusters, allowing the addition of hundreds of low-quality samples to our analysis. This additional information confirmed that the distinct genetic cluster on Haida Gwaii is restricted to that archipelago. As the laingi form was originally described as being based in Haida Gwaii, where the type specimen of that form is from, further study (especially of morphological traits) may indicate a need to restrict this name to the Haida Gwaii genetic cluster. Regardless of taxonomic treatment, our finding of a distinct Haida Gwaii genetic cluster along with the small and historically declining population size of the Haida Gwaii population suggests a high risk of extinction of an ecologically and genetically distinct form of northern goshawk. Outside of Haida Gwaii, sampling regions along the coast of BC and southeast Alaska (often considered regions inhabited by laingi) show some subtle differentiation from other North American regions. We anticipate that these results will increase the effectiveness of conservation management of northern goshawks in northwestern North America. More broadly, other conservation-related studies of genetic variation may benefit from the two-step approach we employed that first surveys genomic variation using high-quality samples and then genotypes low-quality samples at particularly informative loci.

genomics

Transcription-factor centric genome mining strategy for discovery of diverse unprecedented RiPP gene clusters

Ribosomally synthesized and post-translationally modified peptides (RiPPs) are a rapidly emerging group of natural products with diverse biological activity. Most of their biosynthetic mechanisms are well studied and the \"genome mining\" strategy based on homology has led to the unearthing of many new ribosomal natural products, including lantipeptides, lasso peptides, cyanobactins. These precursor-centric or biosynthetic protein-centric genome mining strategies have encouraged the discovery of RiPPs natural products. However, a limitation of these strategies is that the newly identified natural products are similar to the known products and novel families of RiPP pathways were overlooked by these strategies. In this work, we applied a transcription-factor centric genome mining strategy and diverse unique crosslinked RiPP gene clusters were predicted in several sequenced microorganisms. Our research could significantly expand the category of biosynthetic pathways of RiPP natural products and predict new resources for novel RiPPs.

genomics

First Draft Genome of a Brazilian Atlantic Rainforest Burseraceae reveals commercially-promising genes involved in terpenic oleoresins synthesis

BackgroundProtium species produce abundant aromatic oleoresins composed mainly of different types of terpenes, which are highly sought after by the flavor and fragrance industry.\n\nResultsHere we present (i) the first draft genome of an endemic tree of the Brazils Atlantic Rainforest (Mata Atlantica), Protium kleinii Cuatrec., (ii) a first characterization of its genes involved in the terpene pathways, and (iii) the composition of the resins volatile fraction. The de novo draft genome was assembled using Illumina paired-end-only data, totalizing 407 Mb in size present in 229,912 scaffolds. The N50 is 2.60 Kb and the longest scaffold is 52.26 Kb. Despite its fragmentation, we were able to infer 53,538 gene models of which 5,434 were complete. The draft genome of P. kleinii presents 76.67 % (62.01 % complete and 14.66 % partial) of plant-core BUSCO genes. InterProScan was able to assign at least one Gene Ontology annotation and one Pfam domain for 13,629 and 26,469 sequences, respectively. We were able to identify 116 enzymes involved in terpene biosynthesis, such as monoterpenes -terpineol, 1,8-cineole, geraniol, (+)-neomenthol and (+)-(R)-limonene. Through the phylogenetic analysis of the Terpene Synthases gene family, three candidates of limonene synthase were identified. Chemical analysis of the resins volatile fraction identified four monoterpenes: terpinolene, limonene, -pinene and -phellandrene.\n\nConclusionThese results provide resources for further studies to identify the molecular bases of the main aroma compounds and new biotechnological approaches to their production.

genomics

Discovery of Drosophila melanogaster from Wild African Environments and Genomic Insights into Species History

A long-standing enigma concerns the geographic and ecological origins of the intensively studied vinegar fly, Drosophila melanogaster, a globally widespread species [1] which \"has invariably appeared to be a strict human commensal\" [2]. In spite of its sub-Saharan origins, this species has never been reported from undisturbed wilderness environments that might reflect its pre-commensal niche [3]. Here, we document the collection of 288 D. melanogaster individuals from African wilderness areas in Zambia, Zimbabwe, and Namibia. After sequencing the genomes of 17 flies collected from Kafue National Park, Zambia, we found reduced genetic diversity relative to town populations, elevated chromosomal inversion frequencies, and strong differences at specific genes including known insecticide targets. Combining these new genomes with prior data enabled us to gain novel insights into the history of this species geographic expansion. Our demographic estimates indicated that an expansion from southern Africa began approximately 10,000 years ago, with a Saharan crossing soon after, but expansion from the Middle East into Europe did not begin until roughly 1,400 years ago. This improved model of demographic history will provide a critical resource for future evolutionary and genomic studies of this key model organism. Our results add historical context to the species human association, and the opportunity to study wilderness populations opens the door for future studies on the biological basis of its adaptation to human environments.

genomics

Hi-C yields chromosome-length scaffolds for a legume genome, Trifolium subterraneum

We present a chromosome-length assembly of the genome of subterranean clover, Trifolium subterraneum, a key Australian pasture legume. Specifically, in situ Hi-C data (48X) was used to correct misjoins and anchor, order, and orient scaffolds in a previously published genome assembly (TSUd_r1.1; scaffold N50: 287kb). This resulted in an improved genome assembly (TrSub3; scaffold N50: 56Mb) containing eight chromosome-length scaffolds that span 95% of the sequenced bases in the input assembly.

genomics

Genome-wide by environment interaction studies (GWEIS) of depressive symptoms and psychosocial stress in UK Biobank and Generation Scotland.

Stress is associated with poorer physical and mental health. To improve our understanding of this link, we performed genome-wide association studies (GWAS) of depressive symptoms and genome-wide by environment interaction studies (GWEIS) of depressive symptoms and stressful life events (SLE) in two UK population cohorts (Generation Scotland and UK Biobank). No SNP was individually significant in either GWAS, but gene-based tests identified six genes associated with depressive symptoms in UK Biobank (DCC, ACSS3, DRD2, STAG1, FOXP2 and KYNU; p < 2.77x10-6). Two SNPs with genome-wide significant GxE effects were identified by GWEIS in Generation Scotland: rs12789145 (53kb downstream PIWIL4; p = 4.95x10-9; total SLE) and rs17070072 (intronic to ZCCHC2; p = 1.46x10-8; dependent SLE). A third locus upstream CYLC2 (rs12000047 and rs12005200, p < 2.00x10-8; dependent SLE) when the joint effect of the SNP main and GxE effects was considered. GWEIS gene-based tests identified: MTNR1B with GxE effect with dependent SLE in Generation Scotland; and PHF2 with the joint effect in UK Biobank (p < 2.77x10-6). Polygenic risk scores (PRS) analyses incorporating GxE effects improved the prediction of depressive symptom scores, when using weights derived from either the UK Biobank GWAS of depressive symptoms (p = 0.01) or the PGC GWAS of major depressive disorder (p = 5.91x10-3). Using an independent sample, PRS derived using GWEIS GxE effects provided evidence of shared aetiologies between depressive symptoms and schizotypal personality, heart disease and COPD. Further such studies are required and may result in improved treatments for depression and other stress-related conditions.

genetics

CompStor Novos: a low cost yet fast assembly-based variant calling for personal genomes

Application of assembly methods for personal genome analysis from next generation sequencing data has been limited by the requirement for an expensive supercomputer hardware or long computation times when using ordinary resources. We describe CompStor Novos, achieving supercomputer-class performance in de novo assembly computation time on standard server hardware, based on a tiered-memory implementation. Run on commercial off-the-shelf servers, Novos assembly is more precise and 10-20 times faster than that of existing assembly algorithms. Furthermore, we integrated Novos into a variant calling pipeline and demonstrate that both compute times and precision of calling point variants and indels compare well with standard alignment-based pipelines. Additionally, assembly eliminates bias in the estimation of allele frequency for indels and naturally enables discovery of breakpoints for structural variants with base pair resolution. Thus, Novos bridges the gap between alignment-based and assembly-based genome analyses. Extension and adaption of its underlying algorithm will help quickly and fully harvest information in sequencing reads for personal genome reconstruction.

genomics

Evolution of pathogenic and nonpathogenic yeasts mitochondrial genomes inferred by supertrees and supermatrices with divergence estimates based on relaxed molecular clocks

The evolution of mitochondrial genomes is essential for the adaptation of yeasts to the variation of environmental levels of oxygen. Although Saccharomyces cerevisiae mitochondrial DNA lacks all complex I genes, respiration is possible because alternative NADH dehydrogenases are encoded by NDE1 and NDI1 nuclear genes. The proposed whole genome duplication (WGD) in the yeast ancestor at 150-100 million years ago caused nuclear gene duplications and secondary losses, although its relation to the loss of complex I mitocondrial is unknown. Here we present phylogenomic supertrees and supermatrix tree of 46 mitochondrial genomes showing that the loss of complex I predates WGD and occurred independently in the S. cerevisiae group and the fission yeast Schizosaccharomyces pombe. We also show that the branching patterns do not differ dramatically in supertrees and supermatrix phylogenies. Our inferences indicated consistent relations between conserved mitochondrial chromosomal gene order (synteny) in closely related yeasts. Correlation of mitochondrial molecular clock estimates and atmospheric oxygen variation in the Phanerozoic suggests that the Saccharomyces lineage might have lost the complex I during hypoxic periods near Perminian-Triassic or Triassic-Jurassic mass extinction events, while the Schizosaccharomyces lineage possibly lost the complex I during hypoxic environment periods during Middle Cambrian until Lower Devonian. The loss of mitochondrial complex I during low oxygen might not affect yeast metabolism due to fermentative switch. The return to increased oxygen periods might have favored adaptations to aerobic metabolism. Additionally, we also showed that NDE1 and NDI1 phylogenies indicate evolutionary convergence in yeasts where mitochondrial complex I is absent.

genomics

Functional and evolutionary impact of polymorphic inversions in the human genome

Inversions are one type of structural variants linked to phenotypic differences and adaptation in multiple organisms. However, there is still very little information about inversions in the human genome due to the difficulty of their detection. Here, thanks to the development of a new high-throughput genotyping method, we have performed a complete study of a representative set of 45 common human polymorphic inversions. Most inversions promoted by homologous recombination occur recurrently both in humans and great apes and, since they are not tagged by SNPs, they are missed by genome-wide association studies. Furthermore, there is an enrichment of inversions showing signatures of positive or balancing selection, diverse functional effects, such as gene disruption and gene-expression changes, or association with phenotypic traits. Therefore, our results indicate that the genome is more dynamic than previously thought and that human inversions have important functional and evolutionary consequences, making possible to determine for the first time their contribution to complex traits.

genomics

Testing Structural Models of Psychopathology at the Genomic Level

Genome-wide association studies (GWAS) have revealed hundreds of genetic loci associated with the vulnerability to major psychiatric disorders, and post-GWAS analyses have shown substantial genetic correlations among these disorders. This evidence supports the existence of a higher-order structure of psychopathology at both the genetic and phenotypic levels. Despite recent efforts by collaborative consortia such as the Hierarchical Taxonomy of Psychopathology (HiTOP), this structure remains unclear. In this study, we tested multiple alternative structural models of psychopathology at the genomic level, using the genetic correlations among fourteen psychiatric disorders and related psychological traits estimated from GWAS summary statistics. The best-fitting model included four correlated higher-order factors - externalizing, internalizing, thought problems, and neurodevelopmental disorders - which showed distinct patterns of genetic correlations with external validity variables and accounted for substantial genetic variance in their constituent disorders. A bifactor model including a general factor of psychopathology as well as the four specific factors fit worse than the above model. Several model modifications were tested to explore the placement of some disorders - such as bipolar disorder, obsessive-compulsive disorder, and eating disorders - within the broader psychopathology structure. The best-fitting model indicated that eating disorders and obsessive-compulsive disorder, on the one hand, and bipolar disorder and schizophrenia, on the other, load together on the same thought problems factor. These findings provide support for several of the HiTOP higher-order dimensions and suggest a similar structure of psychopathology at the genomic and phenotypic levels.

genomics

Genome-wide association study identifies common genetic risk factors for alcohol, heroin and methamphetamine dependence

BackgroundCommon molecular and cellular foundations underlie different types of substance dependence (SD). However direct evidence for common genetic factors of SD is lacking. Here we aimed to identify specific genetic variants that are shared between alcoholism, heroin and methamphetamine dependence.\n\nMethodsWe first conducted a combined case-control genome-wide association analysis (GWAS) of 521 alcoholic, 1,026 heroin and 1,749 methamphetamine patients and 2,859 healthy controls. We then replicated the significant loci using an independent cohort (146 alcoholic, 1,045 heroin, 763 methamphetamine and 1,904 controls). Second, we examined the genetic effects of these identified SNPs on gene expression, addiction characteristics and brain images (gray and white matter). Furthermore, we investigated the effects of these genetic variants on addiction behaviors using self-administration rat models.\n\nResultsWe identified and validated four genome-wide significant loci in the combined cohorts in the discovery stage: ADH1B rs1229984 (P=6.45x10-10), ANKS1B rs2133896 (P=4.09x10-8), AGBL4 rs147247472 (P=4.30x10-8) and CTNNA2 rs10196867 (P=4.67x10-8). Association results for each dependence group showed that ADH1B rs1229984 was only associated with alcoholism, while the other three loci were associated with heroin, methamphetamine addiction and alcoholism respectively. Variants that were strongly linked to rs2133896 affected ANKS1B gene expression, heroin use frequency and interacted with heroin dependence to affect gray matter of the left calcarine and white matter of the right superior longitudinal fasciculus. In addition, the reduced anks1b expression in the ventral tegmental area increased addiction vulnerability for heroin and methamphetamine in self-administration rat models.\n\nConclusionOur findings revealed several novel genome-wide significant SNPs and genes that synchronously affected the vulnerability and phenotypes for alcoholism, heroin and MA dependence. These findings could shed light on the root cause and the generalized vulnerability for SD.

genomics

Whole-genome sequencing of rare disease patients in a national healthcare system

Most patients with rare diseases do not receive a molecular diagnosis and the aetiological variants and mediating genes for more than half such disorders remain to be discovered. We implemented whole-genome sequencing (WGS) in a national healthcare system to streamline diagnosis and to discover unknown aetiological variants, in the coding and non-coding regions of the genome. In a pilot study for the 100,000 Genomes Project, we generated WGS data for 13,037 participants, of whom 9,802 had a rare disease, and provided a genetic diagnosis to 1,138 of the 7,065 patients with detailed phenotypic data. We identified 95 Mendelian associations between genes and rare diseases, of which 11 have been discovered since 2015 and at least 79 are confirmed aetiological. Using WGS of UK Biobank1, we showed that rare alleles can explain the presence of some individuals in the tails of a quantitative red blood cell (RBC) trait. Finally, we reported 4 novel non-coding variants which cause disease through the disruption of transcription of ARPC1B, GATA1, LRBA and MPL. Our study demonstrates a synergy by using WGS for diagnosis and aetiological discovery in routine healthcare.

genomics