Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Rhopalocnemis phalloides has one of the most reduced and mutated plastid genomes known

Although most plant species are photosynthetic, several hundred species have lost the ability to photosynthesize and instead obtain nutrients via various types of heterotrophic feeding. Their genomes, especially plastid genomes, markedly differ from the genomes of photosynthetic plants. In this work, we describe the sequenced plastid genome of the heterotrophic plant Rhopalocnemis phalloides, which belongs to the family Balanophoraceae and feeds by parasitizing on other plants. The genome is highly reduced (18 622 base pairs versus approximately 150 kilobase pairs in autotrophic plants) and possesses an outstanding AT content, 86.8%, the highest of all sequenced plant plastid genomes. The gene content of this genome is quite typical of heterotrophic plants, with all of the genes related to photosynthesis having been lost. The remaining genes are notably distorted by a high mutation rate and the aforementioned AT content. The high AT content has led to sequence convergence between some of the remaining genes and their homologues from AT-rich plastid genomes of protists. Overall, the plastid genome of R. phalloides is one of the most unusual plastid genomes known.

genomics

Comparison of the 3D organization of sperm and fibroblast genomes using the Hi-C approach

The 3D organization of the genome is tightly connected to its biological function. The Hi-C approach was recently introduced as a method that can be used to identify higher-order chromatin interactions genome-wide. The aim of this study was to determine genome-wide chromatin interaction frequencies using the Hi-C approach in mouse sperm cells and embryonic fibroblasts. The obtained results demonstrated that the 3D genome organizations of sperm and fibroblast cells show a high degree of similarity both with each other and with the previously described mouse embryonic stem (ES) cells. Both A- and B-compartments and topologically associated domains (TADs) are present in spermatozoa and fibroblasts. Nevertheless, sperm cells and fibroblasts exhibited statistically significant differences between each other in the contact probabilities of defined loci. Tight packaging of the sperm genome resulted in an enrichment of long-range contacts compared with the fibroblasts. However, only 30% of the differences in the number of contacts are based on differences in the densities of their genome packages; the main source of the differences is the gain or loss of contacts that are specific for defined genome regions. An analysis of interchromosomal contacts in both cell types demonstrated that the large chromosomes showed a tendency to interact with each other more than with the small chromosomes and vice versa. We found that the dependence of the contact probability P(s) on genomic distance for sperm is in a good agreement with the fractal globular folding of chromatin. The similarity of the spatial DNA organization in sperm and somatic cell genomes suggests the stability of the 3D structure of genomes through generations.

Developmental Biology

A statistical approach to genome size evolution: Observations and explanations

Genome size evolution is a fundamental problem in molecular evolution. Statistical analysis of genome sizes brings new insight into the evolution of genome size. Although the variation of genome sizes is complicated, it is indicated that the genome size evolution can be explained more clearly at taxon level than at species level. I find that the genome size distribution for species in a taxon fits log-normal distribution. And I find a relationship between the phylogeny of life and the statistical features of genome size distributions among taxa. I observed different statistical features of genome size distributions between animal taxa and plant taxa. A log-normal stochastic process model is developed to simulate the genome size evolution. The simulation results on the log-normal distributions of genome sizes and their statistical features agree with the observations.

Evolutionary Biology

Identification of physical interactions between genomic regions by enChIP-Seq

Physical interactions between genomic regions play critical roles in the regulation of genome functions, including gene expression. However, the methods for confidently detecting physical interactions between genomic regions remain limited. Here, we demonstrate the feasibility of using engineered DNA-binding molecule-mediated chromatin immunoprecipitation (enChIP) in combination with next-generation sequencing (NGS) (enChIP-Seq) to detect such interactions. In enChIP-Seq, the target genomic region is captured by an engineered DNA-binding complex, such as a CRISPR system consisting of a catalytically inactive form of Cas9 (dCas9) and a single guide RNA (sgRNA). Subsequently, the genomic regions that physically interact with the target genomic region in the captured complex are sequenced by NGS. Using enChIP-Seq, we found that the 5HS5 locus, which regulates expression of the {beta}-globin genes, interacts with multiple genomic regions upon erythroid differentiation in the human erythroleukemia cell line K562. Genes near the genomic regions inducibly associated with the 5HS5 locus were transcriptionally up-regulated in the differentiated state, suggesting the existence of a coordinated transcription mechanism directly or indirectly mediated by physical interactions between these loci. Our data suggest that enChIP-Seq is a potentially useful tool for detecting physical interactions between genomic regions in a non-biased manner, which would facilitate elucidation of the molecular mechanisms underlying regulation of genome functions.

Molecular Biology

Genomic epidemiology and global diversity of the emerging bacterial pathogen Elizabethkingia anophelis

Elizabethkingia anophelis is an emerging pathogen. Genomic analysis of strains from clinical, environmental or mosquito sources is needed to understand the epidemiological emergence of E. anophelis and to uncover genetic elements implicated in antimicrobial resistance, pathogenesis, or niche adaptation. Here, the genomic sequences of two nosocomial isolates that caused neonatal meningitis in Bangui, Central African Republic, were determined and compared with Elizabethkingia isolates from other world regions and sources. Average nucleotide identity firmly confirmed that E. anophelis, E. meningoseptica and E. miricola represent distinct genomic species and led to re-identification of several strains. Phylogenetic analysis of E. anophelis strains revealed several sublineages and demonstrated a single evolutionary origin of African clinical isolates, which carry unique antimicrobial resistance genes acquired by horizontal transfer. The Elizabethkingia genus and the species E. anophelis had pan-genomes comprising respectively 7,801 and 6,880 gene families, underlining their genomic heterogeneity. African isolates were capsulated and carried a distinctive capsular polysaccharide synthesis cluster. A core-genome multilocus sequence typing scheme applicable to all Elizabethkingia isolates was developed, made publicly available (http://bigsdb.web.pasteur.fr/elizabethkingia), and shown to provide useful insights into E. anophelis epidemiology. Furthermore, a clustered regularly interspaced short palindromic repeats (CRISPR) locus was uncovered in E. meningoseptica, E. miricola and in a few E. anophelis strains. CRISPR spacer variation was observed between the African isolates, illustrating the value of CRISPR for strain subtyping. This work demonstrates the dynamic evolution of E. anophelis genomes and provides innovative tools for Elizabethkingia identification, population biology and epidemiology.\n\nIMPORTANCEElizabethkingia anophelis is a recently recognized bacterial species involved in human infections and outbreaks in distinct world regions. Using whole-genome sequencing, we showed that the species comprises several sublineages, which differ markedly in their genomic features associated with antibiotic resistance and host-pathogen interactions. Further, we have devised high-resolution strain subtyping strategies and provide an open genomic sequence analysis tool, facilitating the investigation of outbreaks and tracking of strains across time and space. We illustrate the power of these tools by showing that two African healthcare-associated meningitis cases observed 5 years apart were caused by the same strain, providing evidence that E. anophelis can persist in the hospital environment.

Microbiology

Allele-Specific Quantification of Structural Variations in Cancer Genomes

One of the hallmarks of cancer genome is aneuploidy, resulting in abnormal copy numbers of alleles. Structural variations (SVs) can further modify the aneuploid cancer genomes into a mixture of rearranged genomic segments with extensive range of somatic copy number alterations (CNAs). Indeed, aneuploid cancer genomes have significantly higher rate of CNAs and SVs. However, although methods have been developed to identify SVs and allele-specific copy number of genome (ASCNG) separately, no existing algorithm can simultaneously analyze SVs and ASCNG. Such integrated approach is particularly important to fully understand the complexity of cancer genomes. Here we introduce a new algorithm called Weaver to provide allele-specific quantification of SVs and CNAs in aneuploid cancer genomes. Weaver uses a probabilistic graphical model by utilizing cancer whole genome sequencing data to simultaneously estimate the digital copy number and inter-connectivity of SVs. Our simulation evaluation, comparison with single-molecule Optical Mapping analysis, and real data applications (including MCF-7, HeLa, and TCGA whole genome sequencing samples) demonstrated that Weaver is highly accurate and can greatly refine the analysis of complex cancer genome structure.

Bioinformatics

Coordinates and Intervals in Graph-based Reference Genomes

MotivationIt has been proposed that future reference genomes should be graph structures in order to better represent the sequence diversity present in a species. However, there is currently no standard method to represent genomic intervals, such as positions of genes, on graph-based reference genomes.\n\nResultsWe formalize offset-based coordinate systems on graph-based reference genomes and introduce a method for representing intervals on these reference structures. We show the advantage of our method by representing genes on a graph-based representation of the GRCh38 version of the human genome and its alternative loci for regions that are highly variable.\n\nConclusionMore complex reference genomes, containing alternative loci, require methods to represent genomic data on these structures. Our proposed notation for genomic intervals makes it possible to fully utilize the alternative loci of GRCh38 and potential future graph-based reference genomes. We illustrate our notation for genomic intervals, as well as the offset-based coordinate systems, through a web tool at: https://github.com/uio-cels/gen-graph-coords.

Bioinformatics

Re-evaluating inheritance in genome evolution: widespread transfer of LINEs between species

Transposable elements (TEs) are mobile DNA sequences, colloquially known as jumping genes because of their ability to replicate to new genomic locations. Given a vector of transfer (e.g. tick or virus), TEs can jump further: between organisms or species in a process known as horizontal transfer (HT). Here we propose that LINE-1 (L1) and Bovine-B (BovB), the two most abundant TE families in mammals, were initially introduced as foreign DNA via ancient HT events. Using a 503-genome dataset, we identify multiple ancient L1 HT events in eukaryotes and provide evidence that L1s infiltrated the mammalian lineage after the monotreme-therian split. We also extend the BovB paradigm by increasing the number of estimated transfer events compared to previous studies, finding new potential blood-sucking parasite vectors and occurrences in new lineages (e.g. bats, frog). Given that these TEs make up nearly half of the genome sequence in todays mammals, our results provide the first evidence that HT can have drastic and long-term effects on the new host genomes. This revolutionizes our perception of genome evolution to consider external factors, such as the natural introduction of foreign DNA. With the advancement of genome sequencing technologies and bioinformatics tools, we anticipate our study to be the first of many large-scale phylogenomic analyses exploring the role of HT in genome evolution.\n\nSignificance statementLINE-1 (L1) elements occupy about half of most mammalian genomes (1), and they are believed to be strictly vertically inherited (2). Mutagenic L1 insertions are thought to account for approximately 1 of every 1000 random, disease-causing insertions in humans (4-7). Our research indicates that the very presence of L1s in humans, and other therian mammals, is due to an ancient transfer event - which has drastic implications for our perception of genome evolution. Using a machina analyses over 503 genomes, we trace the origins of L1 and BovB retrotransposons across the tree of life, and provide evidence of their long-term impact on eukaryotic evolution.

evolutionary biology

Diverse genome organization following 13 independent mesopolyploid events in Brassicaceae contrasts with convergent patterns of gene retention

Hybridization and polyploidy followed by genome-wide diploidization significantly impacted the diversification of land plants. The ancient At- whole-genome duplication (WGD) preceded the diversification of crucifers (Brassicaceae). Some genera and tribes also experienced younger, mesopolyploid WGDs concealed by subsequent genome diploidization. Here we tested if multiple base chromosome numbers originated due to genome diploidization after independent mesopolyploid WGDs and how diploidization impacted post-polyploid gene retention. Sixteen species representing ten Brassicaceae tribes were analyzed by comparative chromosome painting and/or whole-transcriptome analysis of gene age distributions and phylogenetic analyses of gene duplications. Overall, we found evidence for at least 13 independent mesopolyploidies followed by different degrees of diploidization across the Brassicaceae. New mesotetraploid events were uncovered for tribes Anastaticeae, Iberideae and Schizopetaleae, and mesohexaploid WGDs for Cochlearieae and Physarieae. In contrast, we found convergent patterns of gene retention and loss among these independent WGDs. Our combined analyses of Brassicaceae genomic data indicate that the extant chromosome number variation in many plant groups, and especially polybasic but monophyletic taxa, can result from clade-specific genome duplications followed by diploidization. Our observation of parallel gene retention and loss across multiple independent WGDs provides one of the first multi-species tests that post-polyploid genome evolution is predictable.\n\nSignificance statementOur data show that multiple base chromosome numbers in some Brassicaceae clades originated due to genome diploidization following multiple independent whole-genome duplications (WGD). The parallel gene retention/loss across independent WGDs and diploidizations provides one of the first tests that post-polyploid genome evolution is predictable.

evolutionary biology

Inherited chromosomally integrated human herpesvirus 6 genomes are ancient, intact and potentially able to reactivate from telomeres

Human herpesviruses 6A and 6B (HHV6-A and HHV-6B; species Human herpesvirus 6A and Human herpesvirus 6B) have the capacity to integrate into telomeres, the essential capping structures of chromosomes that play roles in cancer and ageing. About 1% of people worldwide are carriers of chromosomally integrated HHV-6 (ciHHV-6), which is inherited as a genetic trait. Understanding the consequences of integration for the evolution of the viral genome, for the telomere and for the risk of disease associated with carrier status is hampered by a lack of knowledge about ciHHV-6 genomes. Here, we report an analysis of 28 ciHHV-6 genomes and show that they are significantly divergent from the few modern non-integrated HHV-6 strains for which complete sequences are currently available. In addition ciHHV-6B genomes in Europeans are more closely related to each other than to ciHHV-6B genomes from China and Pakistan, suggesting regional variation of the trait. Remarkably, at least one group of European ciHHV-6B carriers has inherited the same ciHHV-6B genome, integrated in the same telomere allele, from a common ancestor estimated to have existed 24,500 {+/-}10,600 years ago. Despite the antiquity of some, and possibly most, germline HHV-6 integrations, the majority of ciHHV-6B (95%) and ciHHV-6A (72%) genomes contain a full set of intact viral genes and therefore appear to have the capacity for viral gene expression and full reactivation.\n\nIMPORTANCEInheritance of HHV-6A or HHV-6B integrated into a telomere occurs at a low frequency in most populations studied to date but its characteristics are poorly understood. However, stratification of ciHHV-6 carriers in modern populations due to common ancestry is an important consideration for genome-wide association studies that aim to identify disease risks for these people. Here we present full sequence analysis of 28 ciHHV-6 genomes and show that ciHHV-6B in many carriers with European ancestry most likely originated from ancient integration events in a small number of ancestors. We propose that ancient ancestral origins for ciHHV-6A and ciHHV-6B are also likely in other populations. Moreover, despite their antiquity, all of the ciHHV-6 genomes appear to retain the capacity to express viral genes and most are predicted to be capable of full viral reactivation. These discoveries represent potentially important considerations in immune-compromised patients, in particular in organ transplantation and in stem cell therapy.

evolutionary biology

A computational framework for detecting signatures of accelerated somatic evolution in cancer genomes

By accumulation of somatic mutations, cancer genomes evolve, diverging away from the genome of the host. It remains unclear to what extent somatic evolutionary divergence is comparable across different regions of the cancer genome versus concentrated in specific genomic elements. We present a novel computational framework, SASE-mapper, to identify genomic regions that show signatures of accelerated somatic evolution (SASE) in a subset of samples in a cohort, marked by accumulation of an excess of somatic mutations compared to that expected based on local, context-aware background mutation rates in the cancer genomes. Analyzing tumor whole genome sequencing data for 365 samples from 5 cohorts we detect recurrent SASE at a genome-wide scale. The SASEs were enriched for genomic elements associated with active chromatin, and regulatory regions of several known cancer genes had SASE in multiple cohorts. Regions with SASE carried specific mutagenic signatures and often co-localized within the 3D nuclear space suggesting their common basis. A subset of SASEs was frequently associated with regulatory changes in key cancer pathways and also poor clinical outcome. While the SASE-associated mutations were not necessarily recurrent at base-pair resolution, the SASEs recurrently targeted same functional regions, with similar consequences. It is likely that regulatory redundancy and plasticity promote prevalence of SASE-like patterns in the cancer genomes.

bioinformatics

Superior ab initio Identification, Annotation and Characterisation of TEs and Segmental Duplications from Genome Assemblies.

Transposable Elements (TEs) are mobile DNA sequences that make up significant fractions of amniote genomes. However, they are difficult to detect and annotate ab initio because of their variable features, lengths and clade-specific variants. We have addressed this problem by refining and developing a Comprehensive ab initio Repeat Pipeline (CARP) to identify and cluster TEs and other repetitive sequences in genome assemblies. The pipeline begins with a pairwise alignment using krishna, a custom aligner. Single linkage clustering is then carried out to produce families of repetitive elements. Consensus sequences are then filtered for protein coding genes and then annotated using Repbase and a custom library of retrovirus and reverse transcriptase sequences. This process yields three types of family: fully annotated, partially annotated and unannotated. Fully annotated families reflect recently diverged/young known TEs present in Repbase. The remaining two types of families contain a mixture of novel TEs and segmental duplications. These can be resolved by aligning these consensus sequences back to the genome to assess copy number vs. length distribution. Our pipeline has three significant advantages compared to other methods for ab initio repeat identification: 1) we generate not only consensus sequences, but keep the genomic intervals for the original aligned sequences, allowing straightforward analysis of evolutionary dynamics, 2) consensus sequences represent low-divergence, recently/currently active TE families, 3) segmental duplications are annotated as a useful by-product. We have compared our ab initio repeat annotations for 7 genome assemblies (1 unpublished) to other methods and demonstrate that CARP compares favourably with RepeatModeler, the most widely used repeat annotation package.\n\nAuthor summaryTransposable elements (TEs) are interspersed repetitive DNA sequences, also known as jumping genes, because of their ability to replicate in to new genomic locations. TEs account for a significant proportion of all eukaryotic genomes. Previous studies have found that TE insertions have contributed to new genes, coding sequences and regulatory regions. They also play an important role in genome evolution. Therefore, we developed a novel, ab initio approach for identifying and annotating repetitive elements. The idea is simple: define a \"repeat\" as any sequence that occurs at least twice in the genome. Our ab initio method is able to identify species-specific TEs with high sensitivity and accuracy including both TEs and segmental duplications. Because of the high degree of sequence identity used in our method, the TEs we find are less diverged and may still be active. We also retain all the information that links identified repeat consensus sequences to their genome intervals, permiting direct evolutionary analysis of the TE families we identify.

bioinformatics

Idiosyncratic genome degradation in a bacterial endosymbiont of periodical cicadas

When a free-living bacterium transitions to a host-beneficial endosymbiotic lifestyle, it almost invariably loses a large fraction of its genome [1, 2]. The resulting small genomes often become unusually stable in size, structure, and coding capacity [3-5]. Candidatus Hodgkinia cicadicola (Hodgkinia), a bacterial endosymbiont of cicadas, sometimes exemplifies this genomic stability. The Hodgkinia genome has remained completely co-linear in some cicadas diverged by tens of millions of years [6, 7]. But in the long-lived periodical cicada Magicicada tredecim, the Hodgkinia genome has split into dozens of tiny, gene-sparse genomic circles that sometimes reside in distinct Hodgkinia cells [8]. Previous data suggested that other Magicicada species harbor similarly complex Hodgkinia populations, but the timing, number of origins, and outcomes of the splitting process were unknown. Here, by sequencing Hodgkinia metagenomes from the remaining six Magicicada species and two sister species, we show that all Magicicada species harbor Hodgkinia populations of at least twenty genomic circles each. We find little synteny among the 256 Hodgkinia circles analyzed except between the most closely related species. Individual gene phylogenies show that Hodgkinia first split in the common ancestor of Magicicada and its closest known relatives, but that most splitting has occurred within Magicicada and has given rise to highly variable Hodgkinia gene dosages between cicada species. These data show that Hodgkinia genome degradation has proceeded down different paths in different Magicicada species, and support a model of genomic degradation that is stochastic in outcome and likely nonadaptive for the host. These patterns mirror the genomic instability seen in some mitochondria.

evolutionary biology

Panaconda: Application of pan-synteny graph models to genome content analysis

MotivationWhole-genome alignment and pan-genome analysis are useful tools in understanding the similarities and differences of many genomes in an evolutionary context. Here we introduce the concept of pan-synteny graphs, an analysis method that combines elements of both to represent conservation and change of multiple prokaryotic genomes at an architectural level. Pan-synteny graphs represent a reference free approach for the comparison of many genomes and allows for the identification of synteny, insertion, deletion, replacement, inversion, recombination, missed assembly joins, evolutionary hotspots, and reference based scaffolding.\n\nResultsWe present an algorithm for creating whole genome multiple sequence comparisons and a model for representing the similarities and differences among sequences as a graph of syntenic gene families. As part of the pan-synteny graph creation, we first create a de Bruijn graph. Instead of the alphabet of nucleotides commonly used in genome assembly, we use an alphabet of gene families. This de Bruijn graph is then processed to create the pan-synteny graph. Our approach is novel in that it explicitly controls how regions from the same sequence and genome are aligned and generates a graph in which all sequences are fully represented as paths. This method harnesses previous computation involved in protein family calculation to speed up the creation of whole genome alignment for many genomes. We provide the software suite Panaconda, for the calculation of pan-synteny graphs given annotation input, and an implementation of methods for their layout and visualization.\n\nAvailabilityPanaconda is available at https://github.com/aswarren/pangenome_graphs and datasets used in examples are available at https://github.com/aswarren/pangenome_examples\n\nContactAndrew Warren anwarren@vt.edu

bioinformatics

Assembly-free and alignment-free sample identification using genome skims

The ability to quickly and inexpensively describe taxonomic diversity is critical in this era of rapid climate and biodiversity changes. The currently preferred molecular technique, barcoding, has been very successful, but is based on short organelle markers. Recently, an alternative genome-skimming approach has been proposed: low-pass sequencing (100Mb - several Gb per sample) is applied to voucher and/or query samples, and marker genes and/or organelle genomes are recovered computationally. The current practice of genome-skimming discards the vast majority of the data because the low coverage of genome-skims prevents assembling the nuclear genomes. In contrast, we suggest using all unassembled reads directly, but existing methods poorly support this goal. We introduce a new alignment-free tool, Skmer, to estimate genomic distances between the query and each reference genome-skim using the k-mer decomposition of reads. We test Skmer on a large set of insect and bird genomes, sub-sampled to create genome-skims. Skmer shows great accuracy in estimating genomic distances, identifying the closest match in a reference dataset, and inferring the phylogeny. The software is publicly available on https://github.com/shahab-sarmashghi/Skmer.git

ecology

MGDb: An analyzed database and a genomic resource of mango (Mangifera Indica L.) cultivars for mango research

Mango is one of the famous and fifth most important subtropical/tropical fruit crops worldwide with the production centered in India and South-East Asia. Recently, there has been a worldwide interest in mango genomics to produce tools for Marker Assisted Selection and trait association. There are no web-based analyzed genomic resources available for mango particularly. Hence a complete mango genomic resource was required for improvement in research and management of mango germplasm. In this project, we have done comparative transcriptome analysis of four mango cultivars i.e. cv. Langra, cv. Zill, cv. Shelly and cv. Kent from Pakistan, China, Israel, and Mexico respectively. The raw data is obtained through De-novo sequence assembly which generated 30,953-85,036 unigenes from RNA-Seq datasets of mango cultivars. The project is aimed to provide the scientific community and general public a mango genomic resource and allow the user to examine their data against our analyzed mango genome databases of four cultivars (cv. Langra, cv. Zill, cv. Shelly and cv. Kent). A mango web genomic resource MGdb, is based on 3-tier architecture, developed using Python, flat file database, and JavaScript. It contains the information of predicted genes of the whole genome, the unigenes annotated by homologous genes in other species, and GO (Gene Ontology) terms which provide a glimpse of the traits in which they are involved. This web genomic resource can be of immense use in the assessment of the research, development of the medicines, understanding genetics and provides useful bioinformatics solution for analysis of nucleotide sequence data. We report here worlds first web-based genomic resource particularly of mango for genetic improvement and management of mango genome.

bioinformatics

A synthesis of mapping experiments reveals extensive genomic structural diversity in the Mimulus guttatus species complex

Understanding genomic structural variation such as inversions and translocations is a key challenge in evolutionary genetics. In this paper, we tackle this challenge by developing a novel statistical approach to comparative genetic mapping. The procedure couples a Hidden Markov Model with a Genetic Algorithm to detect large-scale structural variation using low-level sequencing data from multiple genetic mapping populations. We demonstrate the method using five distinct crosses within the flowering plant genus Mimulus. The synthesis of data from these experiments is first used to correct numerous errors (misplaced sequences) in the M. guttatus reference genome. Second, we confirm and/or detect eight large inversions polymorphic within the M. guttatus species complex. Finally, we show how this method can be applied in genomic scans to improve the accuracy and resolution of Quantitative Trait Locus (QTL) mapping.\n\nAUTHOR SUMMARYGenome sequences have proved to be a critical experimental resource for genetic research in many species. However, in some species there is considerable variation in genomic organization, making a single reference genome sequence inadequate. This variation can cause issues in interpreting genomic signals, such as those coming from trait mapping. We introduce a new statistical method and computational tools that use linkage information to reorganize a single reference genome to 1) repair genome assembly errors, and 2) identify variation between individuals or populations of the same species. Using this method we can create a new genome order that improves upon the reference genome. We apply this method to five crosses among plants in the Mimulus guttatus species complex. In this system we detect eight large chromosomal inversions and improve the resolution of a trait mapping study. This work highlights the utility of our method, and indicates how others studying diverse species might use them to improve their own research.

genetics

Eight new genomes of organohalide-respiring Dehalococcoides mccartyi reveal evolutionary trends in reductive dehalogenase enzymes

BackgroundBioaugmentation is now a well-established approach for attenuating toxic groundwater and soil contaminants, particularly for chlorinated ethenes and ethanes. The KB-1 and WBC-2 consortia are cultures used for this purpose. These consortia contain organisms belonging to the Dehalococcoidia, including strains of Dehalococcoides mccartyi in KB-1 and of both D. mccartyi and Dehalogenimonas in WBC-2. These tiny anaerobic bacteria couple respiratory reductive dechlorination to growth and harbour multiple reductive dehalogenase genes (rdhA) in their genomes, the majority of which have yet to be characterized.\n\nResultsUsing a combination of Illumina mate-pair and paired-end sequencing we closed the genomes of eight new strains of Dehalococcoides mccartyi found in three related KB-1 sub-cultures that were enriched on trichloroethene (TCE), 1,2-dichloroethane (1,2-DCA) and vinyl chloride (VC), bringing the total number of genomes available in NCBI to 24. A pangenome analysis was conducted on 24 Dehalococcoides genomes and five Dehalogenimonas genomes (2 in draft) currently available in NCBI. This Dehalococcoidia pangenome generated 2875 protein families comprising of 623 core, 2203 accessory, and 49 unique protein families. In Dehalococcoides mccartyi the complement of reductive dehalogenase genes varies by strain, but what was most surprising was how the majority of rdhA sequences actually exhibit a remarkable degree of synteny across all D. mccartyi genomes. Several homologous sequences are also shared with Dehalogenimonas genomes. Nucleotide and predicted protein sequences for all reductive dehalogenases were aligned to begin to decode the evolutionary history of reductive dehalogenases in the Dehalococcoidia.\n\nConclusionsThe conserved synteny of the rdhA genes observed across Dehalococcoides genomes indicates that the major differences between strain rdhA gene complement has resulted from gene loss rather than recombination. These rdhA have a long evolutionary history and trace their origin in the Dehalococcoidia prior to the speciation of Dehalococcoides and Dehalogenimonas. The only rdhA genes suspected to have been acquired by lateral gene transfer are protein-coding rdhA that have been identified to catalyze dehalogenation of industrial pollutants. Sequence analysis suggests that evolutionary pressures resulting in new rdhA genes involve adaptation of existing dehalogenases to new substrates, mobilization of rdhA between genomes or within a genome, and to a lesser degree manipulation of regulatory regions to alter expression.

bioinformatics