Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Genome-wide patterns of copy number variation in the diversified chicken genomes using next-generation sequencing

Copy number variation (CNV) is important and widespread in the genome, and is a major cause of disease and phenotypic diversity. Herein, we perform genome-wide CNV analysis in 12 diversified chicken genomes based on whole genome sequencing. A total of 9,025 CNV regions (CNVRs) covering 100.1 Mb and representing 9.6% of the chicken genome are identified, ranging in size from 1.1 to 268.8 kb with an average of 11.1 kb. Sequencing-based predictions are confirmed at high validation rate by two independent approaches, including array comparative genomic hybridization (aCGH) and quantitative PCR (qPCR). The Pearson?s correlation values between sequencing and aCGH results range from 0.395 to 0.740, and qPCR experiments reveal a positive validation rate of 91.71% and a false negative rate of 22.43%. In total, 2,188 predicted CNVRs (24.2%) span 2,182 RefSeq genes (36.8%) associated with specific biological functions. Besides two previously accepted copy number variable genes EDN3 and PRLR, we also find some promising genes with potential in phenotypic variants. FZD6 and LIMS1, two genes related to diseases susceptibility and resistance are covered by CNVRs. Highly duplicated SOCS2 may lead to higher bone mineral density. Entire or partial duplication of some genes like POPDC3 and LBFABP may have great economic importance in poultry breeding. Our results based on extensive genetic diversity provide the first individualized chicken CNV map and genome-wide gene copy number estimates and warrant future CNV association studies for important traits of chickens.

Genomics

Whole genome sequencing of Plasmodium falciparum from dried blood spots using selective whole genome amplification

Translating genomic technologies into healthcare applications for the malaria parasite Plasmodium falciparum has been limited by the technical and logistical difficulties of obtaining high quality clinical samples from the field. Sampling by dried blood spot (DBS) finger-pricks can be performed safely and efficiently with minimal resource and storage requirements compared with venous blood (VB). Here, we evaluate the use of selective whole genome amplification (sWGA) to sequence the P. falciparum genome from clinical DBS samples, and compare the results to current methods using leucodepleted VB. Parasite DNA with high (> 95%) human DNA contamination was selectively amplified by Phi29 polymerase using short oligonucleotide probes of 8-12 mers as primers. These primers were selected on the basis of their differential frequency of binding the desired (P. falciparum DNA) and contaminating (human) genomes. Using sWGA method, we sequenced clinical samples from 156 malaria patients, including 120 paired samples for head-to-head comparison of DBS and leucodepleted VB. Greater than 18-fold enrichment of P. falciparum DNA was achieved from DBS extracts. The parasitaemia threshold to achieve >5x coverage for 50% of the genome was 0.03% (40 parasites per 200 white blood cells). Over 99% SNP concordance between VB and DBS samples was achieved after excluding missing calls. The sWGA methods described here provide a reliable and scalable way of generating P. falciparum genome sequence data from DBS samples. Our data indicate that it will be possible to get good quality sequence data on most if not all drug resistance loci from the majority of symptomatic malaria patients. This technique overcomes a major limiting factor in P. falciparum genome sequencing from field samples, and paves the way for large-scale epidemiological applications.

Genomics

Genome-Wide Association Study And Genomic Predictions For Resistance Against Piscirickettsia salmonis In Coho Salmon (Oncorhynchus kisutch) Using ddRAD Sequencing

Piscirickettsia salmonis is one of the main infectious diseases affecting coho salmon (Oncorhynchus kisutch) farming. Current treatments have been ineffective for the control of the disease. Genetic improvement for P. salmonis resistance has been proposed as a feasible alternative for the control of this infectious disease in farmed fish. Genotyping by sequencing (GBS) strategies allow genotyping hundreds of individuals with thousands of single nucleotide polymorphisms (SNPs), which can be used to perform genome wide association studies (GWAS) and predict genetic values using genome-wide information. We used double-digest restriction-site associated DNA (ddRAD) sequencing to dissect the genetic architecture of resistance against P. salmonis in a farmed coho salmon population and identify molecular markers associated with the trait. We also evaluated genomic selection (GS) models in order to determine the potential to accelerate the genetic improvement of this trait by means of using genome-wide molecular information. 764 individuals from 33 full-sib families (17 highly resistant and 16 highly susceptible) which were experimentally challenged against P. salmonis were sequenced using ddRAD sequencing. A total of 4,174 SNP markers were identified in the population. These markers were used to perform a GWAS and testing genomic selection models. One SNP related with iron availability was genome-wide significantly associated with resistance to P. salmonis defined as day of death. Genomic selection models showed similar accuracies and predictive abilities than traditional pedigree-based best linear unbiased prediction (PBLUP) method.

genomics

Improved de novo Genome Assembly: Synthetic long read sequencing combined with optical mapping produce a high quality mammalian genome at relatively low cost

Current short-read methods have come to dominate genome sequencing because they are cost-effective, rapid, and accurate. However, short reads are most applicable when data can be aligned to a known reference. Two new methods for de novo assembly are linked-reads and restriction-site labeled optical maps. We combined commercial applications of these technologies for genome assembly of an endangered mammal, the Hawaiian Monk seal.\n\nWe show that the linked-reads produced with 10X Genomics Chromium chemistry and assembled with Supernova v1.1 software produced scaffolds with an N50 of 22.23 Mbp with the longest individual scaffold of 84.06 Mbp. When combined with Bionano Genomics optical maps using Bionano RefAligner, the scaffold N50 increased to 29.65 Mbp for a total of 170 hybrid scaffolds, the longest of which was 84.78 Mbp. These results were 161X and 215X, respectively, improved over DISCOVAR de novo assemblies. The quality of the scaffolds was assessed using conserved synteny analysis of both the DNA sequence and predicted seal proteins relative to the genomes of humans and other species. We found large blocks of conserved synteny suggesting that the hybrid scaffolds were high quality. An inversion in one scaffold complementary to human chromosome 6 was found and confirmed by optical maps.\n\nThe complementarity of linked-reads and optical maps is likely to make the production of high quality genomes more routine and economical and, by doing so, significantly improve our understanding of comparative genome biology.

genomics

De-novo assembly of zucchini genome reveals a whole genome duplication associated with the origin of the Cucurbita genus

The Cucurbita genus (squashes, pumpkins, gourds) includes important domesticated species such as C. pepo, C. maxima and C. moschata. In this study, we present a high-quality draft of the zucchini (C. pepo) genome. The assembly has a size of 263 Mb, a scaffold N50 of 1.8 Mb, 34,240 gene models, includes 92% of the conserved BUSCO core gene set, and it is estimated to cover 93.0% of the genome. The genome is organized in 20 pseudomolecules, that represent 81.4% of the assembly, and it is integrated with a genetic map of 7,718 SNPs. Despite its small genome size three independent evidences support that the C. pepo genome is the result of a Whole Genome Duplication: the topology of the gene family phylogenies, the karyotype organization, and the distribution of 4DTv distances. Additionally, 40 transcriptomes of 12 species of the genus were assembled and analyzed together with all the other published genomes of the Cucurbitaceae family. The duplication was detected in all the Cucurbita species analyzed, including C. maxima and C. moschata, but not in the more distant cucurbits belonging to the Cucumis and Citrullus genera, and it is likely to have happened 30 {+/-} 4 Mya in the ancestral species that gave rise to the genus.

genomics

Extensive genomic diversity among Mycobacterium marinum strains revealed by whole genome sequencing

Mycobacterium marinum is the causative agent for the tuberculosis-like disease mycobacteriosis in fish and skin lesions in humans. Ubiquitous in its geographical distribution, M. marinum is known to occupy diverse fish as hosts. However, information about its genomic diversity is limited. Here, we provide the genome sequences for 15 M. marinum strains isolated from infected humans and fish. Comparative genomic analysis of these and four available genomes of the M. marinum strains M, E11, MB2 and Europe reveal high genomic diversity among the strains, leading to the conclusion that M. marinum should be divided into two different clusters, the \"M\"- and the \"Aronson\"-type. We suggest that these two clusters should be considered, if not two separate species, at least two M. marinum subspecies. Our data also show that the M. marinum pan-genome for both groups is open and expanding and we provide data showing high number of mutational hotspots in M. marinum relative to other mycobacteria such as Mycobacterium tuberculosis. This high genomic diversity might be related to that M. marinum occupy different ecological niches.

genomics

Emergence and features of the multipartite genome structure of the family Burkholderiaceae revealed through comparative and evolutionary genomics

The multipartite genome structure is found in a diverse group of important symbiotic and pathogenic bacteria; however, the advantage of this genome structure remains incompletely understood. Here, we perform comparative genomics of hundreds of finished {beta}-proteobacterial genomes to study the role and emergence of multipartite genomes. Nearly all essential secondary replicons (chromids) of the {beta}-proteobacteria are found in the family Burkholderiaceae. These replicons arose from just two plasmid acquisition events, and they were likely stabilized early in their evolution by the presence of core genes, at least some of which were likely acquired through an inter-replicon translocation event. On average, Burkholderiaceae genera with multipartite genomes had a larger total genome size, but smaller chromosome, than genera without secondary replicons. Pangenome-level functional enrichment analyses suggested that inter-replicon functional biases are partially driven by the enrichment of secondary replicons in the accessory pangenome fraction. Nevertheless, the small overlap in orthologous groups present in each replicons pangenome indicates a clear functional separation of the replicons. Chromids appeared biased to environmental adaptation, as the functional categories enriched on chromids were also over-represented on the chromosomes of the environmental genera (Paraburkholderia, Cupriavidus) compared to the pathogenic genera (Burkholderia, Ralstonia). Using ancestral state reconstruction, it was predicted that the rate of accumulation of modern-day genes by chromids was more rapid than the rate of gene accumulation by the chromosomes. Overall, the data are consistent with a model where the primary advantage of secondary replicons is in facilitating increased rates of gene acquisition through horizontal gene transfer, consequently resulting in a replicon enriched in genes associated with adaptation to novel environments.

genomics

Population genomics of the Anthropocene: urbanization is negatively associated with genome-wide variation in white-footed mouse populations

Urbanization results in pervasive habitat fragmentation and reduces standing genetic variation through bottlenecks and drift. Loss of genome-wide variation may ultimately reduce the evolutionary potential of animal populations experiencing rapidly changing conditions. In this study, we examined genome-wide variation among 23 white-footed mouse (Peromyscus leucopus) populations sampled along an urbanization gradient in the New York City metropolitan area. Genome-wide variation was estimated as a proxy for evolutionary potential using more than 10,000 SNP markers generated by ddRAD-Seq. We found that genome-wide variation is inversely related to urbanization as measured by percent impervious surface cover, and to a lesser extent, human population density. We also report that urbanization results in enhanced genome-wide differentiation between populations in cities. There was no pattern of isolation by distance among these populations, but an isolation by resistance model based on impervious surface significantly explained patterns of genetic differentiation. Isolation by environment modeling also indicated that urban populations deviate much more strongly from global allele frequencies than suburban or rural populations. This study is the first to examine loss of genome-wide SNP variation along an urban-to-rural gradient and quantify urbanization as a driver of population genomic patterns.

Evolutionary Biology

Integrated Genome Browser: visual analytics platform for genomics

MotivationGenome browsers that support fast navigation and interactive visual analytics can help scientists achieve deeper insight into large-scale genomic data sets more quickly, thus accelerating the discovery process. Toward this end, we developed Integrated Genome Browser (IGB), a highly configurable, interactive and fast open source desktop genome browser.\n\nResultsHere we describe multiple updates to IGB, including all-new capability to display and interact with data from high-throughput sequencing experiments. To demonstrate, we describe example visualizations and analyses of data sets from RNA-Seq, ChIP-Seq, and bisulfite sequencing experiments. Understanding results from genome-scale experiments requires viewing the data in the context of reference genome annotations and other related data sets. To facilitate this, we enhanced IGBs ability to consume data from diverse sources, including Galaxy, Distributed Annotation, and IGB-specific Quickload servers. To support future visualization needs as new genome-scale assays enter wide use, we transformed the IGB codebase into a modular, extensible platform for developers to create and deploy all-new visualizations of genomic data.\n\nAvailabilityIGB is open source and is freely available from http://bioviz.org/igb.\n\nContactaloraine@uncc.edu

Bioinformatics

Prophage genomics reveals patterns in phage genome organization and replication

Temperate phage genomes are highly variable mosaic collections of genes that infect a bacterial host, integrate into the hosts genome or replicate as low copy number plasmids, and are regulated to switch from the lysogenic to lytic cycles to generate new virions and escape their host. Genomes from most Bacterial phyla contain at least one or more prophages. We updated our PhiSpy algorithm to improve detection of prophages and to provide a web-based framework for PhiSpy. We have used this algorithm to identify 36,488 prophage regions from 11,941 bacterial genomes, including almost 600 prophages with no known homology to any proteins. Transfer RNA genes were abundant in the prophages, many of which alleviate the limits of translation efficiency due to host codon bias and presumably enable phages to surpass the normal capacity of the hosts translation machinery. We identified integrase genes in 15,765 prophages (43% of the prophages). The integrase was routinely located at either end of the integrated phage genome, and was used to orient and align prophage genomes to reveal their underlying organization. The conserved genome alignments of phages recapitulate early, middle, and late gene order in transcriptional control of phage genes, and demonstrate that gene order, presumably selected by transcription timing and/or coordination among functional modules has been stably conserved throughout phage evolution.

biochemistry

Comparative genomics of bdelloid rotifers: evaluating the effects of asexuality and desiccation tolerance on genome evolution

Bdelloid rotifers are microscopic invertebrates that have existed for millions of years apparently without sex or meiosis. They inhabit a variety of temporary and permanent freshwater habitats globally, and many species are remarkably tolerant of desiccation. Bdelloids offer an opportunity to better understand the evolution of sex and recombination, but previous work has emphasized desiccation as the cause of several unusual genomic features in this group. Here, we evaluate the relative effects of asexuality and desiccation tolerance on genome evolution by comparing whole genome sequences for three bdelloid species: Adineta ricciae (desiccation tolerant), Rotaria macrura and Rotaria magnacalcarata (both desiccation intolerant) to the only published bdelloid genome to date, that of Adineta vaga (also desiccation tolerant). We find that tetraploidy is conserved among all four bdelloid species, but homologous divergence in obligately aquatic Rotaria genomes is low, well within the range observed between alleles in obligately sexual, diploid animals. In addition, we find that homologous regions in A. ricciae are largely collinear and do not form palindromic repeats as observed in the published A. vaga assembly. These findings are contrary to current understanding of the role of desiccation in shaping the bdelloid genome, and indicate that various features interpreted as genomic evidence for long-term ameiotic evolution are not general to all bdelloid species, even within the same genus. Finally, we substantiate previous findings of high levels of horizontally transferred non-metazoan genes encoded in both desiccating and non-desiccating bdelloid species, and show that this is a unique feature of bdelloids among related animal phyla. Comparisons within bdelloids and to other desiccation-tolerant animals, however, again call into question the purported role of desiccation in horizontal transfer.

evolutionary biology

Genome sequence of Jaltomata addresses rapid reproductive trait evolution and enhances comparative genomics in the hyper-diverse Solanaceae

Within the economically important plant family Solanaceae, Jaltomata is a rapidly evolving genus that has extensive diversity in flower size and shape, as well as fruit and nectar color, among its [~]80 species. Here we report the whole-genome sequencing, assembly, and annotation, of one representative species (Jaltomata sinuosa) from this genus. Combining PacBio long-reads (25X) and Illumina short-reads (148X) achieved an assembly of approximately 1.45 Gb, spanning [~]96% of the estimated genome. 96% of curated single-copy orthologs in plants were detected in the assembly, supporting a high level of completeness of the genome. Similar to other Solanaceous species, repetitive elements made up a large fraction ([~]80%) of the genome, with the most recently active element, Gypsy, expanding across the genome in the last 1-2 million years.\n\nComputational gene prediction, in conjunction with a merged transcriptome dataset from 11 tissues, identified 34725 protein-coding genes. Comparative phylogenetic analyses with six other sequenced Solanaceae species determined that Jaltomata is most likely sister to Solanum, although a large fraction of gene trees supported a conflicting bipartition consistent with substantial introgression between Jaltomata and Capsicum after these species split. We also identified gene family dynamics specific to Jaltomata, including expansion of gene families potentially involved in novel reproductive trait development, and loss of gene families that accompanied the loss of self-incompatibility. This high-quality genome will facilitate studies of phenotypic diversification in this rapidly radiating group, and provide a new point of comparison for broader analyses of genomic evolution across the Solanaceae.

evolutionary biology

Signal sequences in the genome of Mononegavirales regulate the generation of copy-back defective viral genomes

Defective viral genomes of the copy-back type (cbDVGs) are the primary initiators of the antiviral immune response during infection with respiratory syncytial virus (RSV) both in vitro and in vivo. However, the mechanism governing cbDVG generation remains unknown, thereby limiting our ability to manipulate cbDVG content in order to modulate the host response to infection. Here we report a specific genomic signal that mediates the generation of RSV cbDVGs. Using a customized bioinformatics tool, we identified regions in the RSV genome frequently used to generate cbDVGs during infection. We then created a minigenome system to validate the function of one of these sequences and to determine if specific nucleotides were essential for cbDVG generation at that position. Further, we created a recombinant virus that selectively produced a specific cbDVG based on variations introduced in this sequence. The identified sequence was also found as a common site for cbDVG generation during natural RSV infections, and common cbDVGs generated at this sequence were found among samples from various infected patients. These data demonstrate that sequences encoded in the viral genome are critical determinants of the location of cbDVG generation and, therefore, this is not a stochastic process. Most importantly, these findings open the possibility of genetically manipulating cbDVG formation to modulate infection outcome. Author summaryCopy-back defective viral genomes (cbDVGs) regulate infection and pathogenesis of Mononegavirales. cbDVG are believed to arise from random errors that occur during virus replication and the predominant hypothesis is that the viral polymerase is the main driver of cbDVG generation. Here we describe a specific genomic sequence in the RSV genome that is necessary for the generation of a large proportion of the cbDVG population present during infection. We identified specific nucleotides that when modified altered cbDVG generation at this position, and we created a recombinant virus that selectively produced cbDVGs based on mutations in this sequence. These data demonstrate that the generation of RSV cbDVGs is regulated by specific viral sequences and that these sequences can be manipulated to alter the content and quality of cbDVG generated during infection.

microbiology

The causal meaning of genomic predictors and how it affects the construction and comparison of genome-enabled selection models

The additive genetic effect is arguably the most important quantity inferred in animal and plant breeding analyses. The term effect indicates that it represents causal information, which is different from standard statistical concepts as regression coefficient and association. The process of inferring causal information is also different from standard statistical learning, as the former requires causal (i.e. non-statistical) assumptions and involves extra complexities. Remarkably, the task of inferring genetic effects is largely seen as a standard regression/prediction problem, contradicting its label. This widely accepted analysis approach is by itself insufficient for causal learning, suggesting that causality is not the point for selection. Given this incongruence, it is important to verify if genomic predictors need to represent causal effects to be relevant for selection decisions, especially because applying regression studies to answer causal questions may lead to wrong conclusions. The answer to this question defines if genomic selection models should be constructed aiming maximum genomic predictive ability or aiming identifiability of genetic causal effects. Here, we demonstrate that selection relies on a causal effect from genotype to phenotype, and that genomic predictors are only useful for selection if they distinguish such effect from other sources of association. Conversely, genomic predictors capturing non-causal signals provide information that is less relevant for selection regardless of the resulting predictive ability. Focusing on covariate choice decision, simulated examples are used to show that predictive ability, which is the criterion normally used to compare models, may not indicate the quality of genomic predictors for selection. Additionally, we propose using alternative criteria to construct models aiming for the identification of the genetic causal effects.

Genomics

Analysis of clinical Bordetella pertussis isolates using whole genome sequences reveals novel genomic regions associated with recent outbreaks in the United States of America

BackgroundDespite high-levels of vaccination, whooping cough, primarily caused by Bordetella pertussis (BP), has persisted and resurged. It remains a major cause of infant death worldwide and is the most prevalent vaccine-preventable disease in developed countries. To date, most genomic studies have focused on a small subset of the BP genome, biasing our clinical understanding and public health awareness.\n\nMethodsWe performed a Genome-Wide Association Study (GWAS) on 76 U.S. BP whole genomes, including strains from recent outbreaks.\n\nResultsA GWAS of the 76 BP isolates revealed a sharp increase in genetic variation associated with the Minnesota 2012 outbreak and identified 52 variants unique to the Minnesota outbreak and 19 unique to the California and Washington outbreaks. None of the identified variants were shared between the outbreaks and the vast majority were previously uncharacterized. We further identified variation associated with pertactin negative strains and acellular vaccination.\n\nConclusionsWe identified novel genomic regions associated with recent BP outbreaks. Our results underscore the need for increased whole genome sequencing of BP isolates, which can reduce costly misdiagnosis and improve surveillance. The genes containing these variants warrant further investigation into their possible roles in BP pathogenicity and the ongoing resurgence in the U.S.

Genomics

Using MinION nanopore sequencing to generate a de novo eukaryotic draft genome: preliminary physiological and genomic description of the extremophilic red alga Galdieria sulphuraria strain SAG 107.79

We report here the de novo assembly of a eukaryotic genome using only MinION nanopore DNA sequence data by examining a novel Galdieria sulphuraria genome: strain SAG 107.79. This extremophilic red alga was targeted for full genome sequencing as we found that it could grow on a wide variety of carbon sources and could uptake several precious and rare-earth metals, which places it as an interesting biological target for disparate industrial biotechnological uses. Phylogenetic analysis clearly places this as a species of G. sulphuraria. Here we additionally show that the genome assembly generated via nanopore long read data was of a high quality with regards to low total number of contiguous DNA sequences and long length of assemblies. Collectively, the MinION platform looks to rival other competing approaches for de novo genome acquisition with available informatics tools for assembly. The genome assembly is publically released as NCBI BioProject PRJNA330791. Further work is needed to reduce small insertion-deletion errors, relative to short-read assemblies.

Genomics

Revealing the causative variant in Mendelian patient genomes without revealing patient genomes

Given the rapidly growing utility of critical health information revealed in the human genome, secure genomic computation is essential to moving forward, especially as genome sequencing becomes commonplace. We devise and implement proof-of-principle computational operations for precisely identifying causal variants in Mendelian patients using secure multiparty computation methods based on Yaos protocol. We show multiple real scenarios (small patient cohorts, trio analysis, two hospital collaboration) where the causal variant is discovered jointly, while keeping up to 99.7% of all participants most sensitive genomic information private. All similar operations performed today to diagnose such cases are done openly, keeping 0% of participants genomic information private. Our work will help usher in an era where genomes can be both utilized and truly protected.

genomics

Mammalian genomic regulatory regions predicted by utilizing human genomics, transcriptomics and epigenetics data

Genome sequences for hundreds of mammalian species are available, but an understanding of their genomic regulatory regions, which control gene expression, is only beginning. A comprehensive prediction of potential active regulatory regions is necessary to functionally study the roles of the majority of genomic variants in evolution, domestication, and animal production. We developed a computational method to predict regulatory DNA sequences (promoters, enhancers and transcription factor binding sites) in production animals (cows and pigs) and extended its broad applicability to other mammals. The method utilizes human regulatory features identified from thousands of tissues, cell lines, and experimental assays to find homologous regions that are conserved in sequences and genome organization and are enriched for regulatory elements in the genome sequences of other mammalian species. Importantly, we developed a filtering strategy, including a machine learning classification method, to utilize a very small number of species-specific experimental datasets available to select for the likely active regulatory regions. The method finds the optimal combination of sensitivity and accuracy to unbiasedly predict regulatory regions in mammalian species. Furthermore, we demonstrated the utility of the predicted regulatory datasets in cattle for prioritizing variants associated with multiple production and climate change adaptation traits, and identifying potential genome editing targets.

genomics