Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,297 records · Page 72Linked to original sources

Declaring a tuberculosis outbreak over with genomic epidemiology

We report an updated method for inferring the time at which an infectious disease was transmitted between persons from a time-labelled pathogen genome phylogeny. We applied the method to 48 Mycobacterium tuberculosis genomes as part of a real-time public health outbreak investigation, demonstrating that although active tuberculosis (TB) cases were diagnosed through 2013, no transmission events took place beyond mid-2012. Subsequent cases were the result of progression from latent TB infection to active disease and not recent transmission. This evolutionary genomic approach was used to declare the outbreak over in January 2015.

Genomics

Phenotypic and genomic analysis of P elements in natural populations of Drosophila melanogaster

The Drosophila melanogaster P transposable element provides one of the best cases of horizontal transfer of a mobile DNA sequence in eukaryotes. Invasion of natural populations by the P element has led to a syndrome of phenotypes known as \"P-M hybrid dysgenesis\" that emerges when strains differing in their P element composition mate and produce offspring. Despite extensive research on many aspects of P element biology, questions remain about the stability and genomic basis of variation in P-M dysgenesis phenotypes. Here we report the P-M status for a number of populations sampled recently from Ukraine that appear to be undergoing a shift in their P element composition. Gondal dysgenesis assays reveal that Ukrainian populations of D. melanogaster are currently dominated by the P cytotype, a cytotype that was previously thought to be rare in nature, suggesting that a new active form of the P element has recently spread in this region. We also compared gondal dysgenesis phenotypes and genomic P element predictions for isofemale strains obtained from three worldwide populations of D. melanogaster in order to guide further work on the molecular basis of differences in cytotype status across populations. We find that the number of euchromatic P elements per strain can vary significantly across populations but that total P element numbers are not strongly correlated with the degree of gondal dysgenesis. Our work shows that rapid changes in cytotype status can occur in natural populations of D. melanogaster, and informs future efforts to decode the genomic basis of geographic and temporal differences in P element induced phenotypes.

Genomics

A scored human protein-protein interaction network to catalyze genomic interpretation

Human protein-protein interaction networks are critical to understanding cell biology and interpreting genetic and genomic data, but are challenging to produce in individual large-scale experiments. We describe a general computational framework that through data integration and quality control provides a scored human protein-protein interaction network (InWeb_IM). Juxtaposed with five comparable resources, InWeb_IM has 2.8 times more interactions (~585K) and a superior functional signal showing that the added interactions reflect real cellular biology. InWeb_IM is a versatile resource for accurate and cost-efficient functional interpretation of massive genomic datasets illustrated by annotating candidate genes from >4,700 cancer genomes and genes involved in neuropsychiatric diseases.

Genomics

Gene mapping of nine agronomic traits and genome assembly by resequencing a foxtail millet RIL population

Foxtail millet (Setaria italica) provides food and fodder in semi-arid regions and infertile land. Resequencing of 184 foxtail millet recombinant inbred lines (RILs) was carried out to aid essential research on foxtail millet improvement. Bin map were constructed based on the RILs recombination data. By anchoring some unseated scaffolds and filling gaps, we update two original millet reference genomes Zhanggu and Yugu to produce second editions. Gene mapping of nine agronomic traits were done based on this RIL population. The genome resequencing and QTL mapping provided important tools for foxtail millet research and breeding. Resequencing of the RILs could also provide an effective way for high quantity genome assembly and gene identification.

Genomics

Local PCA shows how the effect of population structure differs along the genome

Population structure leads to systematic patterns in measures of mean relatedness between individuals in large genomic datasets, which are often discovered and visualized using dimension reduction techniques such as principal component analysis (PCA). Mean relatedness is an average of the relationships across locus-specific genealogical trees, which can be strongly affected on intermediate genomic scales by linked selection and other factors. We show how to use local principal components analysis to describe this meso-scale heterogeneity in patterns of relatedness, and apply the method to genomic data from three species, finding in each that the effect of population structure can vary substantially across only a few megabases. In a global human dataset, localized heterogeneity is likely explained by polymorphic chromosomal inversions. In a range-wide dataset of Medicago truncatula, factors that produce heterogeneity are shared between chromosomes, correlate with local gene density, and may be caused by linked selection, such as background selection or local adaptation. In a dataset of primarily African Drosophila melanogaster, large-scale heterogeneity across each chromosome arm is explained by known chromosomal inversions thought to be under recent selection, and after removing samples carrying inversions, remaining heterogeneity is correlated with recombination rate and gene density, again suggesting a role for linked selection. The visualization method provides a flexible new way to discover biological drivers of genetic variation, and its application to data highlights the strong effects that linked selection and chromosomal inversions can have on observed patterns of genetic variation.

Genomics

Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly

The human reference genome assembly plays a central role in nearly all aspects of todays basic and clinical research. GRCh38 is the first coordinate-changing assembly update since 2009 and reflects the resolution of roughly 1000 issues and encompasses modifications ranging from thousands of single base changes to megabase-scale path reorganizations, gap closures and localization of previously orphaned sequences. We developed a new approach to sequence generation for targeted base updates and used data from new genome mapping technologies and single haplotype resources to identify and resolve larger assembly issues. For the first time, the reference assembly contains sequence-based representations for the centromeres. We also expanded the number of alternate loci to create a reference that provides a more robust representation of human population variation. We demonstrate that the updates render the reference an improved annotation substrate, alter read alignments in unchanged regions and impact variant interpretation at clinically relevant loci. We additionally evaluated a collection of new de novo long-read haploid assemblies and conclude that while the new assemblies compare favorably to the reference with respect to continuity, error rate, and gene completeness, the reference still provides the best representation for complex genomic regions and coding sequences. We assert that the collected updates in GRCh38 make the newer assembly a more robust substrate for comprehensive analyses that will promote our understanding of human biology and advance our efforts to improve health.

Genomics

HiChIP: Efficient and sensitive analysis of protein-directed genome architecture

Genome conformation is central to gene control but challenging to interrogate. Here we present HiChIP, a protein-centric chromatin conformation method. HiChIP improves the yield of conformation-informative reads by over 10-fold and lowers input requirement over 100-fold relative to ChIA-PET. HiChIP of cohesin reveals multi-scale genome architecture with greater signal to background than in situ Hi-C. Thus, HiChIP adds to the toolbox of 3D genome structure and regulation for diverse biomedical applications.

Genomics

Whole genome sequence analysis of Salmonella Typhi isolated in Thailand before and after the introduction of a national immunization program

Vaccines against Salmonella Typhi, the causative agent of typhoid fever, are commonly used by travellers, however, there are few examples of national immunization programs in endemic areas. There is therefore a paucity of data on the impact of typhoid immunization programs on localised populations of S. Typhi. Here we have used whole genome sequencing (WGS) to characterise 44 historical bacterial isolates collected before and after a national typhoid immunization program that was implemented in Thailand in 1977 in response to a large outbreak; the program was highly effective in reducing typhoid case numbers. Thai isolates were highly diverse, including 10 distinct phylogenetic lineages or genotypes. Novel prophage and plasmids were also detected, including examples that were previously only reported in Shigella sonnei and Escherichia coli. The majority of S. Typhi genotypes observed prior to the immunization program were not observed following it. Post-vaccine era isolates were more closely related to S. Typhi isolated from neighbouring countries than to earlier Thai isolates, providing no evidence for the local persistence of endemic S. Typhi following the national immunization program. Rather, later cases of typhoid appeared to be caused by the occasional importation of common genotypes from neighbouring Vietnam, Laos, and Cambodia. These data show the value of WGS in understanding the impacts of vaccination on pathogen populations and provide support for the proposal that large-scale typhoid immunization programs in endemic areas could result in lasting local disease elimination, although larger prospective studies are needed to test this directly.\n\nAuthor SummaryTyphoid fever is a systemic infection caused by the bacterium Salmonella Typhi. Typhoid fever is associated with inadequate hygiene in low-income settings and a lack of sanitation infrastructure. A sustained outbreak of typhoid fever occurred in Thailand in the 1970s, which peaked in 1975-1976. In response to this typhoid fever outbreak the government of Thailand initiated an immunization program, which resulted in a dramatic reduction in the number of typhoid cases in Thailand. To better understand the population of S. Typhi circulating in Thailand at this time, as well as the impact of the immunization program on the pathogen population, we sequenced the genomes of 44 S. Typhi obtained from hospitals in Thailand before and after the immunization program. The genome sequences showed that isolates of S. Typhi bacteria isolated from post-immunization era typhoid cases were likely imported from neighbouring countries, rather than strains that have persisted in Thailand throughout the immunization period. Our work provides the first historical insights into S. Typhi in Thailand during the 1970s, and provides a model for the impact of immunization on S. Typhi populations.

Genomics

Contrasting genome dynamics between domesticated and wild yeasts

Structural rearrangements have long been recognized as an important source of genetic variation with implications in phenotypic diversity and disease, yet their evolutionary dynamics are difficult to characterize with short-read sequencing. Here, we report long-read sequencing for 12 strains representing major subpopulations of the partially domesticated yeast Saccharomyces cerevisiae and its wild relative Saccharomyces paradoxus. Complete genome assemblies and annotations generate population-level reference genomes and allow for the first explicit definition of chromosome partitioning into cores, subtelomeres and chromosome-ends. High-resolution view of structural dynamics uncovers that, in chromosomal cores, S. paradoxus exhibits higher accumulation rate of balanced structural rearrangements (inversions, translocations and transpositions) whereas S. cerevisiae accumulates unbalanced rearrangements (large insertions, deletions and duplications) more rapidly. In subtelomeres, recurrent interchromosomal reshuffling was found in both species, with higher rate in S. cerevisiae. Such striking contrasts between wild and domesticated yeasts reveal the influence of human activities on structural genome evolution.

Genomics

Choice of reference genome can introduce massive bias in bisulfite sequencing data

In the study of DNA methylation, genetic variation between species, strains, or individuals can result in CpG sites that are exclusive to a subset of samples, and insertions and deletions can rearrange the spatial distribution of CpGs. How to account for this variation in an analysis of the interplay between sequence variation and DNA methylation is not well understood, especially when the number of CpG differences between samples is large. Here we use whole-genome bisulfite sequencint data on two highly divergent inbred mouse strains to study this problem. We find that while the large number of strain-specific CpGs necessitates considerations regarding the reference genomes used during alignment, properties such as CpG density are surprisingly conserved across the genome. We introduce a method for including strain-specific CpGs in differential analysis, and show that accounting for strain-specific CpGs increases the power to find differentially methylated regions between the strains. Our method uses smoothing to impute methylation levels at strain-specific sites, thereby allowing strain-specific CpGs to contribute to the analysis, and also allowing us to account for differences in the spatial occurrences of CpGs. Our results have implications for analysis of genetic variation and DNA methylation using bisulfite-converted DNA.

Genomics

Contamination as a major factor in poor Illumina assembly of microbial isolate genomes

The nonhybrid hierarchical assembly of PacBio long reads is becoming the most preferred method for obtaining genomes for microbial isolates. On the other hand, among massive numbers of Illumina sequencing reads produced, there is a slim chance of re-evaluating failed microbial genome assembly (high contig number, large total contig size, and/or the presence of low-depth contigs). We generated Illumina-type test datasets with various levels of sequencing error, pretreatment (trimming and error correction), repetitive sequences, contamination, and ploidy from both simulated and real sequencing data and applied k-mer abundance analysis to quickly detect possible diagnostic signatures of poor assemblies. Contamination was the only factor leading to poor assemblies for the test dataset derived from haploid microbial genomes, resulting in an extraordinary peak within low-frequency k-mer range. When thirteen Illumina sequencing reads of microbes belonging to genera Bacillus or Paenibacillus from a single multiplexed run were subjected to a k-mer abundance analysis, all three samples leading to poor assemblies showed peculiar patterns of contamination. Read depth distribution along the contig length indicated that all problematic assemblies suffered from too many contigs with low average read coverage, where 1% to 15% of total reads were mapped to low-coverage contigs. We found that subsampling or filtering out reads having rare k-mers could efficiently remove low-level contaminants and greatly improve the de novo assemblies. An analysis of 16S rRNA genes recruited from reads or contigs and the application of read classification tools originally designed for metagenome analyses can help identify the source of a contamination. The unexpected presence of proteobacterial reads across multiple samples, which had no relevance to our lab environment, implies that such prevalent contamination might have occurred after the DNA preparation step, probably at the place where sequencing service was provided.

genomics

Detecting hierarchical 3-D genome domain reconfiguration with network modularity

Mammalian genomes are folded in a hierarchy of topologically associating domains (TADs), subTADs and looping interactions. The nested nature of chromatin domains has rendered it challenging to identify a sensitive and specific metric for detecting subTADs and quantifying their dynamic reconfiguration across cellular states. Here, we apply graph theoretic principles to quantify hierarchical folding patterns in high-resolution chromatin topology maps. We discover that TADs can be accurately detected using a Louvain-like locally greedy algorithm to maximize network modularity. By varying a resolution parameter in the modularity quality function, we accurately partition the mouse genome across length scales into a hierarchical nested structure of network communities exhibiting a wide range of sizes. To distinguish high probability subTADs from the full detected set, we developed and applied a new hierarchical spatial variance minimization method. Moreover, we identified a large number of dynamically altered communities between pluripotent embryonic stem cells and multipotent neural progenitor cells. Cell type specific boundaries correlate with trends in dynamic occupancy of the architectural protein CTCF, thereby validating their biological relevance. Together, these data demonstrate the utility of metrics from network science in quantifying a nested hierarchy of dynamic 3D chromatin communities across length scales. Our findings are significant toward unraveling the link between higher-order genome folding and gene expression during healthy development and the deregulation of molecular pathways linked to disease.

genomics

Genomic characterization of serial-passaged Ebola virus in a boa constrictor cell line

Ebola virus disease (EVD) is a viral hemorrhagic fever with a high case-fatality rate in humans. EVD is caused by four members of the filoviral genus Ebolavirus, with Ebola virus (EBOV) being the most notorious one. Although bats are discussed as potential ebolavirus reservoirs, limited data actually support this hypothesis. Glycoprotein 2 (GP2) of reptarenaviruses, known to infect only boa constrictors and pythons, are similar in sequence and structure to ebolaviral glycoprotein 2 (GP2), suggesting that EBOV may be able to infect snake cells. We therefore serially passaged EBOV and a distantly related filovirus, Marburg virus (MARV), in the boa constrictor kidney cell line, JK, and characterized viral growth and mutational frequency by sequencing. We observed that EBOV efficiently infected and replicated in JK cells, but MARV did not. In contrast to most cell lines, EBOV infected JK cells did not result in obvious cytopathic effect (CPE). Genomic characterization of serial-passaged EBOV in JK cells revealed that genomic adaptation was not required for infection. Deep sequencing coverage (>10,000x) demonstrated the existence of only a single non-synonymous variant (EBOV glycoprotein precursor preGP T544I) of unknown significance within the viral population that exhibited a shift in frequency of at least 10% over six passages. Our data suggest that boid snake derived cells are competent for filovirus infection without appreciable genomic adaptation; that cellular filovirus infection without CPE may be more common than currently appreciated; and that there may be significant differences between the natural host spectra of ebolaviruses and marburgviruses.\n\nIMPORTANCEEbola virus (EBOV) causes a high case-fatality form of viral hemorrhagic fever. The natural reservoir of EBOV remains unknown. EBOV is distantly related to Marburg virus (MARV), which has been found in bats in the wild. The glycoprotein of a reptarenavirus known to infect boid snakes (pythons and boas) shows similarity in sequence and structure to these viruses, suggesting that EBOV and MARV may be able to infect and replicate in snake cells. We demonstrate that JK, a boa constrictor cell line, does not support MARV infection, but does support EBOV infection without causing overt cytopathic effect or the need for appreciable adaptation. These findings suggest different filoviruses may have a more diverse natural host spectra than previously thought.

genomics

The dynamic three-dimensional organization of the diploid yeast genome

The budding yeast Saccharomyces cerevisiae is a long-standing model for the three-dimensional organization of eukaryotic genomes. Even in this well-studied model, it is unclear how homolog pairing in diploids and environment-induced gene relocalization influence overall genome organization. Here, we performed high-throughput chromosome conformation capture on diverged Saccharomyces hybrid diploids to obtain the first global view of chromosome conformation in diploid yeasts. After controlling for the Rabl-like orientation, we observe significant homolog proximity that increased in saturated culture conditions. Surprisingly, we observe a localized increase in homologous interactions between the HAS1 alleles specifically under galactose induction and saturated growth, mediated by association with nuclear pore complexes at the nuclear periphery. Together, these results reveal that the diploid yeast genome has a dynamic and complex 3D organization.

genomics

Lineage-specific rediploidization is a mechanism to explain time-lags between genome duplication and evolutionary diversification

The functional divergence of duplicate genes (ohnologues) retained from whole genome duplication (WGD) is thought to promote evolutionary diversification. However, species radiation and phenotypic diversification is often highly temporally-detached from WGD. Salmonid fish, whose ancestor experienced WGD by autotetraploidization ~95 Ma (i.e. Ss4R), fit such a time-lag model of post-WGD radiation, which occurred alongside a major delay in the rediploidization process. Here we propose a model called Lineage-specific Ohnologue Resolution (LORe) to address the phylogenetic and functional consequences of delayed rediploidization. Under LORe, speciation precedes rediploidization, allowing independent ohnologue divergence in sister lineages sharing an ancestral WGD event. Using cross-species sequence capture, phylogenomics and genome-wide analyses of ohnologue expression divergence, we demonstrate the major impact of LORe on salmonid evolution. One quarter of each salmonid genome, harbouring at least 4,500 ohnologues, has evolved under LORe, with rediploidization and functional divergence occurring on multiple independent occasions > 50 Myr post-WGD. We demonstrate the existence and regulatory divergence of many LORe ohnologues with functions in lineage-specific physiological adaptations that promoted salmonid species radiation. We show that LORe ohnologues are enriched for different functions than older ohnologues that began diverging in the salmonid ancestor. LORe has unappreciated significance as a nested component of post-WGD divergence that impacts the functional properties of genes, whilst providing ohnologues available solely for lineage-specific adaptation. Under LORe, which is predicted following many WGD events, the functional outcomes of WGD need not appear explosively, but can arise gradually over tens of Myr, promoting lineage-specific diversification regimes under prevailing ecological pressures.

genomics

Genomic and Environmental Contributions to Chronic Diseases in Urban Populations

Uncovering the interaction between genomes and the environment is a principal challenge of modern genomics and preventive medicine. While theoretical models are well defined, little is known of the GxE interactions in humans. We used a system biology approach to comprehensively assess the interactions between 1.6 million environmental exposure data, health, and expression phenotypes, together with whole genome genetic variation, for [~]1000 individuals from a founder-population in Quebec. We reveal a substantial impact of the urbanization gradient on the transcriptome and clinical endophenotypes, overpowering that of genetic ancestry. In detail, air pollution impacts gene expression and pathways affecting cardio-metabolic and respiratory traits when controlling for genetic ancestry. Finally, we capture 34 clinically associated expression quantitative trait loci that interact with the environment (air pollution). Our findings demonstrate how the local environment directly affects chronic disease development, and that genetic variation, including rare variants, can modulate individuals response to environmental challenges.\n\nHighlightsO_LIFine scale environmental effects overpower those of ancestry on gene expression\nC_LIO_LIAir pollution (geographic and temporal) is associated with transcriptional response\nC_LIO_LIGene-by-environment interactions with air pollution include asthma associated loci\nC_LIO_LIInflammatory pathways and cardio-respiratory clinical traits are among those affected\nC_LI

genomics

European Flint reference sequences complement the maize pan-genome

The genomic diversity of maize is reflected by a large number of SNPs and substantial structural variation. Here, we report the de novo assembly of two European Flint maize lines to remedy the scarcity of sequence resources for the Flint pool. EP1 and F7 are important founder lines of European hybrid breeding programs. The lines were sequenced on an Illumina platform at 320X and 225X coverage. Using NRGene{acute}s DeNovoMAGIC 2.0 technology, pseudochromosomes were assembled encompassing a total of 2,463 Mb for EP1 and 2,405 Mb for F7. Structural and functional annotation of the two genomes is currently in progress. The two high-quality de novo assemblies complement the existing maize pan-genome and will pave the way for future functional and comparative studies.

genomics

Repli-seq: genome-wide analysis of replication timing by next-generation sequencing

Cycling cells duplicate their DNA content during S phase, following a defined program called replication timing (RT). Early and late replicating regions differ in terms of mutation rates, transcriptional activity, chromatin marks and sub-nuclear position. Moreover, RT is regulated during development and is altered in disease. Exploring mechanisms linking RT to other cellular processes in normal and diseased cells will be facilitated by rapid and robust methods with which to measure RT genome wide. Here, we describe a protocol to analyse genome-wide RT by next-generation sequencing (NGS). This protocol yields highly reproducible results across laboratories and platforms. We also provide the computational pipelines for analysis, parsing phased genomes using single nucleotide polymorphisms (SNP) for analyzing imprinted RT, and for direct comparison to Repli-chip data obtained by analyzing nascent DNA by microarrays.

genomics