Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Building a genome browser with GIVE

Growing popularity and diversity of genomic data demands portable and versatile genome browsers. Here, we present an open source programming library, called GIVE that facilitates creation of personalized genome browsers without requiring a system administrator. By inserting HTML tags, one can add to a personal webpage interactive visualization of multiple types of genomics data, including genome annotation, \"linear\" quantitative data (wiggle), and genome interaction data. GIVE includes a graphical interface called HUG (HTML Universal Generator) that automatically generates HTML code for displaying user chosen data, which can be copy-pasted into users personal website or saved and shared with collaborators. The simplicity of use was enabled by encapsulation of novel data communication and visualization technologies, including new data structures, a memory management method, and a double layer display method. GIVE is available at: http://www.givengine.org/.

bioinformatics

Limited role of differential fractionation in genome content variation and function in maize (Zea mays L.) inbred lines

Maize is a diverse paleotetraploid species with widespread presence/absence variation and copy number variation. One mechanism through which presence/absence variation can arise is differential fractionation. Fractionation refers to the loss of duplicate gene pairs from one of the maize subgenomes during diploidization and differential fractionation refers to non-shared gene loss events between individuals. We investigated the prevalence of presence/absence variation resulting from differential fractionation in the syntenic portion of the genome using two whole genome de novo assemblies of the inbred lines B73 and PH207. Between these two genomes, syntenic genes were highly conserved with less than 1% of syntenic genes being subject to differential fractionation. The few variable syntenic genes that were identified are unlikely to contribute to functional phenotypic variation, as there is a significant depletion of these genes in annotated gene sets. In further comparisons of 60 diverse inbred lines, non-syntenic genes were six times more likely to be variable compared to syntenic genes, suggesting that comparisons among additional genome assemblies are not likely to result in the discovery of large-scale presence/absence variation among syntenic genes.\n\nSIGNIFICANCE STATEMENTThere is a large amount of presence/absence variation for gene content in maize. One mechanism that has been hypothesized to contribute to this variation is differential fractionation between individuals following the maize whole genome duplication event. Using comparative genomics, with sorghum and rice representing the ancestral state, we observed little evidence of differential fractionation among elite inbred lines and the few differentially fractionated genes identified did not appear to confer functional significance.

plant biology

Identifying the genetic determinants of particular phenotypes in microbial genomes with very small training sets

Machine learning (ML) encompasses numerous algorithms that aim at discovering complex patterns between elements within large data using limited prior assumptions or modeling. However, some scientific disciplines still produce small data sets: in particular, empirical studies that try to find the mutations responsible for complex phenotypes are often limited to very small sample sizes (n), while scanning a large number of amino acid sites (p) in a proteome. To date, little is known on how ML performs in this type of so-called \"large p, small n\" problem. To address this question, we evaluated the performance of two general ML classifiers, adaptive boosting (AB) and random forest, on two data sets. To assess the impact of proteome size, we contrasted a small (viral) genome with a larger (bacterial) one. To analyze large proteomes, we further developed a chunking algorithm, and introduce a repeated random forest (RRF) algorithm that stabilizes model predictions. With the influenza data, we were able to rediscover amino acid sites experimentally implicated in three different complex phenotypes (infectivity, transmissibility, and pathogenicity). Results for the larger proteome, pertaining to three types of drug resistance (Ciprofloxacin, Ceftazidime, and Gentamicin), were more nuanced, with RRF making more sensible pre-dictions, with smaller errors rates, than AB. Furthermore, we show that chunking improved runtimes by an order of magnitude and may increase sensitivity of the predictions. Altogether, we demonstrate that ML algorithms can be used to identify genetic determinants in small proteomes (viruses), even with small numbers of individuals. We further show that even if the size of bacterial proteomes pushes AB to its limits in the context of small n, RRF may deserve more scrutiny, which should be facilitated by the plummeting costs of sequencing and, more critically, by phenotyping large cohorts of individuals.\n\nAuthor SummaryFinding the genetic determinants of a phenotype is typically performed by testing for an association between a particular allele and a trait, carrying out the testing over a large number of loci in a large cohort of individuals, itself divided into two subsets of individuals: those who have the trait (cases), and those who do not (controls). However, recruiting large cohorts can be problematic in some experimental fields, while using genotypic information rather than complete genomes can miss some mutations. To address these issues, we implemented two machine learning (ML) algorithms, tweaked for analyzing large genomes and providing stable results. The analysis of a small viral genome, for which genetic determinants of three phenotypes are already known, showed that our approach can rediscover known mutations, almost irrespective of the ML algorithm used. However, the analysis of a larger bacterial genome, for which genetic determinants of three phenotypes are unknown, suggested that the simpler of our modified algorithms performed better, returning more sensitive predictions with lower error rates. This work demonstrates the feasibility of finding genetic determinants of complex phenotypes based on a small number of complete genomes.

bioinformatics

Local and global chromatin interactions are altered by large genomic deletions associated with human brain development

BackgroundLarge copy number variants (CNVs) in the human genome are strongly associated with common neurodevelopmental, neuropsychiatric disorders such as schizophrenia and autism. Using Hi-C analysis of long-range chromosome interactions and ChIP-Seq analysis of regulatory histone marks we studied the epigenomic effects of the prominent large deletion CNV on chromosome 22q11.2 and also replicated a subset of the findings for the large deletion CNV on chromosome 1q21.1.\n\nResultsWe found that, in addition to local and global gene expression changes, there are pronounced and multilayered effects on chromatin states, chromosome folding and topological domains of the chromatin, that emanate from the large CNV locus. Regulatory histone marks are altered in the deletion proximal regions, and in opposing directions for activating and repressing marks. There are also significant changes of histone marks elsewhere along chromosome 22q and genome wide. Chromosome interaction patterns are weakened within the deletion boundaries and strengthened between the deletion proximal regions. We detected a change in the manner in which chromosome 22q folds onto itself, namely by increasing the long-range contacts between the telomeric end and the deletion proximal region. Further, the large CNV affects the topological domain that is spanning its genomic region. Finally, there is a widespread and complex effect on chromosome interactions genome-wide, i.e. involving all other autosomes, with some of the effect directly tied to the deletion region on 22q11.2.\n\nConclusionsThese findings suggest novel principles of how such large genomic deletions can alter nuclear organization and affect genomic molecular activity.

genetics

Machine learning leveraging genomes from metagenomes identifies influential antibiotic resistance genes in the infant gut microbiome

Antibiotic resistance in pathogens is extensively studied, yet little is known about how antibiotic resistance genes of typical gut bacteria influence microbiome dynamics. Here, we leverage genomes from metagenomes to investigate how genes of the premature infant gut resistome correspond to the ability of bacteria to survive under certain environmental and clinical conditions. We find that formula feeding impacts the resistome. Random forest models corroborated by statistical tests revealed that the gut resistome of formula-fed infants is enriched in class D beta-lactamase genes. Interestingly, Clostridium difficile strains harboring this gene are at higher abundance in formula-fed infants compared to C. difficile lacking this gene. Organisms with genes for major facilitator superfamily drug efflux pumps have faster replication rates under all conditions, even in the absence of antibiotic therapy. Using a machine learning approach, we identified genes that are predictive of an organisms direction of change in relative abundance after administration of vancomycin and cephalosporin antibiotics. The most accurate results were obtained by reducing annotated genomic data into five principal components classified by boosted decision trees. Among the genes involved in predicting if an organism increased in relative abundance after treatment are those that encode for subclass B2 beta-lactamases and transcriptional regulators of vancomycin resistance. This demonstrates that machine learning applied to genome-resolved metagenomics data can identify key genes for survival after antibiotics and predict how organisms in the gut microbiome will respond to antibiotic administration.\n\nImportanceThe process of reconstructing genomes from environmental sequence data (genome-resolved metagenomics) allows for unique insight into microbial systems. We apply this technique to investigate how the antibiotic resistance genes of bacteria affect their ability to flourish in the gut under various conditions. Our analysis reveals that strain-level selection in formula-fed infants drives enrichment of beta-lactamase genes in the gut resistome. Using genomes from metagenomes, we built a machine learning model to predict how organisms in the gut microbial community respond to perturbation by antibiotics. This may eventually have clinical and industrial applications.

microbiology

Functional Difference of Mitochondrial Genome and Its Association with Traits of Common Complex Diseases in Humans

Recent evidence suggests that mitochondrial genomes harboring common mitochondrial DNA polymorphisms might have functional difference and could be associated with common complex human diseases such as metabolic syndrome and cancer that are related to mitochondrial dysfunction. However, there has been no report examining the functional difference of mitochondrial genome in the pathogenesis of such diseases at the cellular or molecular level. In order to examine the effect of mitochondrial genome on metabolic syndrome or cancer without interference from nuclear genes, we analyzed trans-mitochondrial cytoplasmic hybrid cells (cybrids) with common Asian mtDNA haplogroups A, B, D, and F from healthy volunteers. The mitochondrial oxygen consumption rates of cybrids were associated with multiple components of metabolic syndrome such as body mass index, waist circumference, serum triglyceride levels and high-density lipoprotein cholesterol levels. In addition, the cybrids showed varying degree of tumorigenicity both in vitro and in vivo. Especially, the cybrids harboring mtDNA haplogroup D had a significantly slower growth rate. These findings suggest that the phenotypes of common complex diseases in humans can be determined by their mitochondrial genomes. Therefore, not only nuclear genome but also mitochondrial genome should be considered in explaining the genetic pathogenesis of common complex human diseases.

genetics

Desert Tortoises in the Genomic Age: Population Genetics and the Landscape

The California Department of Fish and Wildlife (CDFW) provided research funds to study the conservation genomics and landscape genomics of the Mojave desert tortoise, Gopherus agassizii, in response to the Desert Renewable Energy Conservation Plan (DRECP). To do this, we consolidated tissue samples of the desert tortoise from across the species range within California and southern Nevada, generated a DNA dataset consisting of full genomes of 270 tortoises, and analyzed the way in which the environment of the desert tortoise has determined modern patterns of relatedness and genetic diversity across the landscape. Here we present the implications of these results for the conservation and landscape genomics of the desert tortoise. Our work strongly indicates that several well-defined genetic groups exist within the species, including a primary north-south genetic discontinuity at the Ivanpah Valley and another separating western from eastern Mojave samples. We also use existing desert tortoise habitat modeling data with a novel extension of genetic \"resistance distance\" using geographic maps of continuous space to predict the relative impacts of five proposed development alternatives within the DRECP and rank them with respect to their likely impacts on desert tortoise gene flow and connectivity in the Mojave. Finally, we analyzed the impacts of each of the 214 distinct proposed development area \"chunks,\" derived from the proposed development polygons, and ranked each chunk in terms of its range-wide impacts on desert tortoise gene flow.\n\n\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=150 SRC=\"FIGDIR/small/195743_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (147K):\norg.highwire.dtl.DTLVardef@b60416org.highwire.dtl.DTLVardef@1c649f2org.highwire.dtl.DTLVardef@120e625org.highwire.dtl.DTLVardef@e59b34_HPS_FORMAT_FIGEXP M_FIG C_FIG PrefaceO_ST_ABSContextC_ST_ABSThe following document is a report that was submitted to the California Department of Fish and Wildlife, describing a series of analyses to help understand the impacts of several alternative spatial configurations of renewable energy development on gene flow of the federally threatened Mojave desert tortoise. These development alternatives were the centerpiece of the Desert Renewable Energy Conservation Plan (DRECP), a landscape-level land use planning initiative undertaken by the Bureau of Land Management (BLM), U.S. Fish and Wildlife Service (USFWS), California Energy Commission (CEC), and the California Department of Fish and Wildlife (CDFW). We were tasked by the California Department of Fish and Wildlife with providing a detailed analysis of these alternative plans on desert tortoise gene flow, and submitted the report for the public comment period for the initial implementation of the DRECP.\n\nFuture PlansO_ST_ABSCurrent state of landscape-level planning for the Mojave desert tortoiseC_ST_ABSThe five proposed land use configuration alternatives analyzed in the subsequent report include public and private lands spread across several counties in California. Shortly after the end of the DRECPs public comment period, the government agencies that developed the DRECP announced that they would be splitting its implementation into two phases: one that deals with land use decisions on BLM-controlled lands and one that deals with non-BLM areas (Sahagun 2015).\n\nPhase I of the DRECP was approved by the Bureau of Land Management on September 14, 2016 (U.S. Bureau of Land Management 2016). This phase includes land use planning decisions for BLM-administered lands. Specifically, 388,000 acres of public lands were designated as development focus areas (DFAs). In applications for leasing lands for renewable energy development, DFAs will not require the same degree of environmental evaluation prior to permitting, as theyve already been evaluated in the context of the DRECP. The application process for renewable energy development within DFAs will be streamlined to encourage development in these areas. Phase I also designated a total of 6,527,000 acres for natural resource conservation. This includes California Desert National Conservation Lands, Areas of Critical Environmental Concern, and Wildlife Allocations. A further 2,691,000 acres were designated for recreation under Phase I. Phase II of the DRECP is currently under development in conjunction with county-level governments to extend this landscape-level planning beyond BLM-administered lands.\n\nAuthor ContributionsThis was a collaborative report. Evan McCartney-Melstad performed the simulations of the low-coverage full genome approach (see Figure 2); conducted all of the laboratory work to generate the genome sequences; performed all of the bioinformatic analyses to bring the raw sequence data to the various stages required for different analyses; wrote the software to quickly estimate pairwise genetic relationships between individuals using read count data in low coverage sequence data (www.github.com/atcg/cPWP); performed some of the population genetic analyses; and wrote and edited several sections of the report.\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=122 SRC=\"FIGDIR/small/195743_fig2.gif\" ALT=\"Figure 2\">\nView larger version (16K):\norg.highwire.dtl.DTLVardef@30a49forg.highwire.dtl.DTLVardef@187c3d8org.highwire.dtl.DTLVardef@4abe98org.highwire.dtl.DTLVardef@126febe_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 2.C_FLOATNO Comparison of two different sequencing approaches in their ability to differentiate very slightly differentiated populations (Fst=0.001)\n\nC_FIG Peter Ralph (in collaboration with Gideon Bradburd and Erik Lundgren) invented and implemented the random walk-based gene flow model that we used to estimate reductions in gene flow due to development, and also developed the theory behind the read-based pairwise pi and genetic covariance estimation used here, in addition to writing and editing several sections of the report. Gideon Bradburd also performed some of the population genetic analyses and wrote and edited several sections of the report. Jannet Vu collected and curated the spatial environmental data and generated the maps that are included in the report (Figures 10, A10-A13), and also wrote Appendices I and IV. Bridgette Hagerty, Fran Sandmeier, Chava Weitzman, and C. Richard Tracy contributed approximately 1,000 desert tortoise blood samples that they collected (at great effort), in addition to knowledge of tortoise ecology and conservation, as well as the results of previous microsatellite-based genetic analyses and editing of the report. H. Bradley Shaffer wrote and edited several sections of the report, and is listed as the lead author for his role in conceiving of and obtaining funding support for the project.\n\nO_FIG O_LINKSMALLFIG WIDTH=154 HEIGHT=200 SRC=\"FIGDIR/small/195743_fig10.gif\" ALT=\"Figure 10\">\nView larger version (82K):\norg.highwire.dtl.DTLVardef@11e99acorg.highwire.dtl.DTLVardef@1faf2f5org.highwire.dtl.DTLVardef@64c440org.highwire.dtl.DTLVardef@1907837_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 10.C_FLOATNO Spatial configuration of the proposed development chunks (see Appendix 4).\n\nC_FIG

evolutionary biology

Whole Genomes Define Concordance of Matched Primary, Xenograft, and Organoid Models of Pancreas Cancer

Pancreatic ductal adenocarcinoma (PDAC) has the worst prognosis among solid malignancies and improved therapeutic strategies are needed to improve outcomes. Patient-derived xenografts (PDX) and patient-derived organoids (PDO) serve as promising tools to identify new drugs with therapeutic potential in PDAC. For these preclinical disease models to be effective, they should both recapitulate the molecular heterogeneity of PDAC and validate patient-specific therapeutic sensitivities. To date however, deep characterization of PDAC PDX and PDO models and comparison with matched human tumour remains largely unaddressed at the whole genome level. We conducted a comprehensive assessment of the genetic landscape of 16 whole-genome pairs of tumours and matched PDX, from primary PDAC and liver metastasis, including a unique cohort of 5 trios of matched primary tumour, PDX, and PDO. We developed a new pipeline to score concordance between PDAC models and their paired human tumours for genomic events, including mutations, structural variations, and copy number variations. Comparison of genomic events in the tumours and matched disease models displayed single-gene concordance across major PDAC driver genes, and genome-wide similarities of copy number changes. Genome-wide and chromosome-centric analysis of structural variation (SV) events revealed high variability across tumours and disease models, but also highlighted previously unrecognized concordance across chromosomes that demonstrate clustered SV events. Our approach and results demonstrate that PDX and PDO recapitulate PDAC tumourigenesis with respect to simple somatic mutations and copy number changes, and capture major SV events that are found in both resected and metastatic tumours.

bioinformatics

Genomic Locus Modulating Corneal Thickness in the Mouse Identifies POU6F2 as a Potential Risk of Developing Glaucoma

Purpose: Central corneal thickness (CCT) is one of the most heritable ocular traits and it is also a phenotypic risk factor for primary open angle glaucoma (POAG). The present study uses the BXD Recombinant Inbred (RI) strains to identify novel quantitative trait loci (QTLs) modulating CCT in the mouse with the potential of identifying a molecular link between CCT and risk of developing POAG.\n\nMethods: The BXD RI strain set was used to define mammalian genomic loci modulating CCT, with a total of 818 corneas measured from 61 BXD RI strains (between 60-100 days of age). The mice were anesthetized and the eyes were positioned in front of the lens of the Phoenix Micron IV Image-Guided OCT system or the Bioptigen OCT system. CCT data for each strain was averaged and used to identify quantitative trait loci (QTLs) modulating this phenotype using the bioinformatics tools on GeneNetwork (www.genenetwork.org). The candidate genes and genomic loci identified in the mouse were then directly compared with the summary data from a human primary open-angle glaucoma (POGA) genome wide association study (NEIGHBORHOOD) to determine if any genomic elements modulating mouse CCT are also risk factors for POAG.\n\nResults: This analysis revealed one significant QTL on Chr 13 and a suggestive QTL on Chr 7. The significant locus on Chr 13 (13 to 19 Mb) was examined further to define candidate genes modulating this eye phenotype. For the Chr 13 QTL in the mouse, only one gene in the region (Pou6f2) contained nonsynonymous SNPs. Of these five nonsynonymous SNPs in Pou6f2, two resulted in changes in the amino acid proline which could result in altered secondary structure affecting protein function. The 7 Mb region under the mouse Chr 13 peak distributes over 2 chromosomes in the human: Chr 1 and Chr 7. These genomic loci were examined in the NEIGHBORHOOD database to determine if they are potential risk factors for human glaucoma identified using meta-data from human GWAS. The top 50 hits all resided within one gene (POU6F2), with the highest significance level of p = 10-6 for SNP rs76319873. POU6F2 is found in retinal ganglion cells and in corneal limbal stem cells. To test the effect of POU6F2 on CCT we examined the corneas of a Pou6f2-null mice and the corneas were thinner than those of wild-type littermates. In addition, these POU6F2 RGCs die early in the DBA/2J model of glaucoma than most RGCs.\n\nConclusions: Using a mouse genetic reference panel, we identified a transcription factor, Pou6f2, that modulates CCT in the mouse. POU6F2 is also found in a subset of retinal ganglion cells and these RGCs are sensitive to injury.\n\nAuthors SummaryGlaucoma is a complex group of diseases with several known causal mutations and many known risk factors. One well-known risk factor for developing primary open angle glaucoma is the thickness of the central cornea. The present study leverages a unique blend of systems biology methods using BXD recombinant inbred mice and genome-wide association studies from humans to define a putative molecular link between a phenotypic risk factor (central corneal thickness) and glaucoma. We identified a transcription factor, POU6F2, that is found in the developing retinal ganglion cells and cornea. POU6F2 is also present in a subpopulation of retinal ganglion cells and in stem cells of the cornea. Functional studies reveal that POU6F2 is associated the central corneal thickness and with susceptibility of retinal ganglion cells to injury.

genetics

Promyelocytic Leukemia (PML) Nuclear Bodies (NBs) Induce Latent/Quiescent HSV-1 Genomes Chromatinization Through a PML-NB/Histone H3.3/H3.3 Chaperone Axis

Herpes simplex virus 1 (HSV-1) latency establishment is tightly controlled by promyelocytic leukemia (PML) nuclear bodies (NBs) (or ND10), although their exact implication is still elusive. A hallmark of HSV-1 latency is the interaction between latent viral genomes and PML-NBs, leading to the formation of viral DNA-containing PML-NBs (vDCP-NBs). Using a replication-defective HSV-1-infected human primary fibroblast model reproducing the formation of vDCP-NBs, combined with an immuno-FISH approach developed to detect latent/quiescent HSV-1, we show that vDCP-NBs contain both histone H3.3 and its chaperone complexes, i.e., DAXX/ATRX and HIRA complex (HIRA, UBN1, CABIN1, and ASF1a). HIRA also co-localizes with vDCP-NBs present in trigeminal ganglia (TG) neurons from HSV-1-infected wild type mice. ChIP-qPCR performed on fibroblasts stably expressing tagged H3.3 (e-H3.3) or H3.1 (e-H3.1) show that latent/quiescent viral genomes are chromatinized almost exclusively with e-H3.3, consistent with an interaction of the H3.3 chaperones with multiple viral loci. Depletion by shRNA of single proteins from the H3.3 chaperone complexes only mildly affects H3.3 deposition on the latent viral genome, suggesting a compensation mechanism. In contrast, depletion (by shRNA) or absence of PML (in mouse embryonic fibroblast (MEF) pml-/- cells) significantly impacts the chromatinization of the latent/quiescent viral genomes with H3.3 without any overall replacement with H3.1. Consequently, the study demonstrates a specific epigenetic regulation of latent/quiescent HSV-1 through an H3.3-dependent HSV-1 chromatinization involving the two H3.3 chaperones DAXX/ATRX and HIRA complexes. Additionally, the study reveals that PML-NBs are major actors in latent/quiescent HSV-1 H3.3 chromatinization through a PML-NB/histone H3.3/H3.3 chaperone axis.\n\nAuthor summaryAn understanding of the molecular mechanisms contributing to the persistence of a virus in its host is essential to be able to control viral reactivation and its associated diseases. Herpes simplex virus 1 (HSV-1) is a human pathogen that remains latent in the PNS and CNS of the infected host. However, the latency is unstable, and frequent reactivations of the virus are responsible for PNS and CNS pathologies. It is thus crucial to understand the physiological, immunological and molecular levels of interplay between latent HSV-1 and the host. Promyelocytic leukemia (PML) nuclear bodies (NBs) play a major role in controlling viral infections by preventing the onset of lytic infection. In previous studies, we showed a major role of PML-NBs in favoring the establishment of a latent state for HSV-1. A hallmark of HSV-1 latency establishment is the formation of PML-NBs containing the viral genome, which we called \"viral DNA-containing PML-NBs\" (vDCP-NBs). The genome entrapped in the vDCP-NBs is transcriptionally silenced. This naturally occurring latent/quiescent state could, however, be transcriptionally reactivated. Therefore, understanding the role of PML-NBs in controlling the establishment of HSV-1 latency and its reactivation is essential to design new therapeutic approaches based on the prevention of viral reactivation.

microbiology

Phytophthora methylomes modulated by expanded 6mA methyltransferases are associated with adaptive genome regions

Filamentous plant pathogen genomes often display a bipartite architecture with gene sparse, repeat-rich compartments serving as a cradle for adaptive evolution. However, the extent to which this \"two-speed\" genome architecture is associated with genome-wide epigenetic modifications is unknown. Here, we show that the oomycete plant pathogens Phytophthora infestans and Phytophthora sojae possess functional adenine N6- methylation (6mA) methyltransferases that modulate patterns of 6mA marks across the genome. In contrast, 5-methylcytosine (5mC) could not be detected in the two Phytophthora species. Methylated DNA IP Sequencing (MeDIP-seq) of each species revealed that 6mA is depleted around the transcriptional starting sites (TSS) and is associated with low expressed genes, particularly transposable elements. Remarkably, genes occupying the gene-sparse regions have higher levels of 6mA compared to the remainder of both genomes, possibly implicating the methylome in adaptive evolution of Phytophthora. Among three putative adenine methyltransferases, DAMT1 and DAMT3 displayed robust enzymatic activities. Surprisingly, single knockouts of each of the 6mA methyltransferases in P. sojae significantly reduced in vivo 6mA levels, indicating that the three enzymes are not fully redundant. MeDIP-seq of the damt3 mutant revealed uneven patterns of 6mA methylation across genes, suggesting that PsDAMT3 may have a preference for gene body methylation after the TSS. Our findings provide evidence that 6mA modification is an epigenetic mark of Phytophthora genomes and that complex patterns of 6mA methylation by the expanded 6mA methyltransferases may be associated with adaptive evolution in these important plant pathogens.

molecular biology

Genomic and geographic footprints of differential introgression between two highly divergent fish species

Investigating variation in gene flow across the genome between closely related species is important to understand how reproductive isolation builds up during the speciation process. An efficient way to characterize differential gene flow is to study how the genetic interactions that take place in hybrid zones selectively filter gene exchange between species, leading to heterogeneous genome divergence. In the present study, genome-wide divergence and introgression patterns were investigated between two sole species, Solea senegalensis and Solea aegyptiaca, using a restriction-associated DNA sequencing (RAD-Seq) approach to analyze samples taken from a transect spanning the hybrid zone. An integrative approach combining geographic and genomic clines methods with an analysis of individual locus introgression taking into account the demographic history of divergence inferred from the joint allele frequency spectrum was conducted. Our results showed that only a minor fraction of the genome can still substantially introgress between the two species due to genome-wide congealing. We found multiple evidence for a preferential direction of introgression in the S. aegyptiaca genetic background, indicating a possible recent or ongoing movement of the hybrid zone. Deviant introgression signals found in the opposite direction supported that the Mediterranean populations of S. senegalensis could have benefited from adaptive introgression. Our study thus illustrates the varied outcomes of genetic interactions between divergent gene pools that recently met after a long history of divergence.

evolutionary biology

Optimal cross selection for long-term genetic gain in two-part programs with rapid recurrent genomic selection

This study evaluates optimal cross selection for balancing selection and maintenance of genetic diversity in two-part plant breeding programs with rapid recurrent genomic selection. The two-part program reorganizes a conventional breeding program into population improvement component with recurrent genomic selection to increase the mean of germplasm and product development component with standard methods to develop new lines. Rapid recurrent genomic selection has a large potential, but is challenging due to genotyping costs or genetic drift. Here we simulate a wheat breeding program for 20 years and compare optimal cross selection against truncation selection in the population improvement with one to six cycles per year. With truncation selection we crossed a small or a large number of parents. With optimal cross selection we jointly optimised selection, maintenance of genetic diversity, and cross allocation with AlphaMate program. The results show that the two-part program with optimal cross selection delivered the largest genetic gain that increased with the increasing number of cycles. With four cycles per year optimal cross selection had 78% (15%) higher long-term genetic gain than truncation selection with a small (large) number of parents. Higher genetic gain was achieved through higher efficiency of converting genetic diversity into genetic gain; optimal cross selection quadrupled (doubled) efficiency of truncation selection with a small (large) number of parents. Optimal cross selection also reduced the drop of genomic selection accuracy due to the drift between training and prediction populations. In conclusion, optimal cross-selection enables optimal management and exploitation of population improvement germplasm in two-part programs.\n\nKey messageOptimal cross selection increases long-term genetic gain of two-part programs with rapid recurrent genomic selection. It achieves this by optimising efficiency of converting genetic diversity into genetic gain through reducing the loss of genetic diversity and reducing the drop of genomic prediction accuracy with rapid cycling.

genetics

Peripatric speciation associated with genome expansion and female-biased sex ratios in the moss genus Ceratodon

PREMISE OF THE STUDYA period of allopatry is widely believed to be essential for the evolution of reproductive isolation. However, strict allopatry may be difficult to achieve in some cosmopolitan, spore-dispersed groups, like mosses. Here we examine the genetic and genome size diversity in Mediterranean populations of the moss Ceratodon purpureus s.l. to evaluate the role of allopatry and ploidy change in population divergence.\n\nMETHODSWe sampled populations of the genus Ceratodon from mountainous areas and lowlands of the Mediterranean region, and from western and central Europe. We performed phylogenetic and coalescent analyses on sequences from five nuclear introns and a chloroplast locus to reconstruct their evolutionary history. We also estimated the genome size using flow cytometry, employing propidium iodide, and determined their sex using a sex-linked PCR marker.\n\nKEY RESULTSTwo well differentiated clades were resolved, discriminating two homogeneous groups: the widespread C. purpureus and a local group mostly restricted to the mountains in southern Spain. The latter also possessed a genome size 25% larger than the widespread C. purpureus, and the samples of this group consist entirely of females. We also found hybrids, and some of them had a genome size equivalent to the sum of the C. purpureus and Spanish genome, suggesting that they arose by allopolyploidy.\n\nCONCLUSIONSThese data suggest that a new species of Ceratodon arose via peripatric speciation, potentially involving a genome size change and a strong female-biased sex ratio. The new species has hybridized in the past with C. purpureus.

molecular biology

PERSONAL GENOMICS: new concepts for future community data banks

INTRODUCTION INTRODUCTION SPLIT GENOMIC DATA TRANSACTIONS BETWEEN COMMUNITY... RESULTS AND DISCUSSION REFERENCES The advent of fast, low-cost sequencing techniques is boosting the number of sequenced human genomes thus driving the irresistible push towards personalized genomics and personalized medicine.1 The examination of the genomes of large cohorts of patients will permit to identify the origins of many diseases and to propose adapted, fine-tuned medical treatments. The genomic boom is also expected to generate an economic activity worth billions of dollars and a fierce race has already started between Internet giants like Google and Amazon. Big data companies have already gathered tens of thousands of individual human genomes hosted in the cloud (that is on a server accessible via the Internet) and th ...

bioinformatics

Comparative Annotation Toolkit (CAT) - simultaneous clade and personal genome annotation

The recent introductions of low-cost, long-read, and read-cloud sequencing technologies coupled with intense efforts to develop efficient algorithms have made affordable, high-quality de novo sequence assembly a realistic proposition. The result is an explosion of new, ultra-contiguous genome assemblies. To compare these genomes we need robust methods for genome annotation. We describe the fully open source Comparative Annotation Toolkit (CAT), which provides a flexible way to simultaneously annotate entire clades and identify orthology relationships. We show that CAT can be used to improve annotations on the rat genome, annotate the great apes, annotate a diverse set of mammals, and annotate personal, diploid human genomes. We demonstrate the resulting discovery of novel genes, isoforms and structural variants, even in genomes as well studied as rat and the great apes, and how these annotations improve cross-species RNA expression experiments.

bioinformatics

Outbreak of invasive wound mucormycosis in a burn unit due to multiple strains of Mucor circinelloides f. circinelloides resolved by whole genome sequencing

Mucorales are ubiquitous environmental molds responsible for mucormycosis in diabetic, immunocompromised, and severely burned patients. Small outbreaks of invasive wound mucormycosis (IWM) have already been reported in burn units without extensive microbiological investigations. We faced an outbreak of IWM in our center and investigated the clinical isolates with whole genome sequencing (WGS) analysis.\n\nWe analyzed M. circinelloides isolates from patients in our burn unit (BU1) together with non-outbreak isolates from burn unit 2 (BU2, Paris area) and from France over a two-year period (2013-2015). For each isolate, WGS and a de novo genome assembly was performed from read data extracted from the aligned contig sequences of the reference genome (1006PhL).\n\nA total of 21 isolates were sequenced including 14 isolates from six BU1 patients. Phylogenetic classification showed that the clinical isolates clustered in four highly divergent clades. Clade1 contained at least one of the strains from the six epidemiologically-linked BU1 patients. The clinical isolates seemed specific to each patient. Two patients were infected with more than two strains from different clades suggesting that an environmental reservoir of clonally unrelated isolates was the source of contamination. Only two patients shared one strain in BU1, suggesting direct transmission or contamination with the same environmental source.\n\nWGS coupled with precise epidemiological data and analysis of several isolates per patients revealed in our study a complex situation with both potential cross-transmission and multiple contaminations with a heterogeneous pool of strains from a cryptic environmental reservoir.\n\nImportanceInvasive wound mucormycosis (IWM) is a severe infection due to the environmental molds belonging to the order Mucorales. Severely burned patients are particularly at risk for IWM. Here, we used Whole Genome Sequencing (WGS) analysis to resolve an outbreak of IWM due to Mucor circinelloides that occurred in our hospital (BU1). We sequenced 21 clinical isolates, including 14 from BU1 and 7 unrelated isolates, and compared them to the reference genome (1006PhL). This analysis revealed that the outbreak was mainly due to multiple strains that seemed patient-specific, suggesting that the patients were more likely infected from a pool of diverse strains from the environment rather than from direct transmission between the patients. This study revealed the complexity of a Mucorales outbreak in the settings of IWM in burn patients, which has been highlighted based on whole genome sequencing and careful sampling.

microbiology

A Regression-based Framework for Scalable Pathway-guided Search in Genome-wide Association Studies.

Traditional unbiased genome-wide association studies (GWAS) have successfully identified thousands of loci associated with various complex diseases but there is evidence to suggest that many variants were missed at stringent genome-wide thresholds. Fortunately, there is a rapidly increasing amount of prior knowledge in publicly available genomic datasets and biological databases that can be harnessed to enhance the power of discovering SNPs/Genes from existing or new GWAS datasets. For most diseases, many of the identified loci tend to cluster into a few specific biological pathways/networks. From the point of view of disease etiology, such clustering is generally to be expected. This phenomenon can be exploited to conduct a more powerful genome-wide scan that is tailored to identify loci that are interconnected in pathways. We propose a scalable regression-based analytical framework to enable such a pathway-guided GWAS and demonstrate that it provides significant gains in power to detect disease associated SNPs. Our method requires two inputs, namely a) genome-wide summary level data (e.g., SNP p-values) and b) a grouping of genes into biologically meaningful categories (e.g., a database of pathways). It automatically adjusts the input p-values by incorporating the knowledge derived adaptively from the data and the pathways specified. The method involves a regularized logistic regression analysis to derive priors of each SNP and then re-weights the p-values of SNPs so as to maximize overall power of making discoveries. It increases the power to discover SNPs co-clustering into some of these pathways, while maintaining the global type-1 error (FWER) at the desired level. We used whole-genome simulations and summary data from real GWA studies of psoriasis, SLE, coronary artery disease and type-2 diabetes to illustrate the power improvement achieved by pathway-guided search. Our pipeline implemented as an R package can flexibly handle large number of prior annotations possibly derived from multiple databases.

genetics