Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Rapid and recent evolution of LTR retrotransposons drives rice genome evolution during the speciation of AA- genome Oryza species

The dynamics of LTR retrotransposons and their contribution to genome evolution during plant speciation have remained largely unanswered. Here, we perform a genome-wide comparison of all eight Oryza AA- genome species, and identify 3,911 intact LTR retrotransposons classified into 790 families. The top 44 most abundant LTR retrotransposon families show patterns of rapid and distinct diversification since the species split over the last ~4.8 Myr. Phylogenetic and read depth analyses of 11 representative retrotransposon families further provide a comprehensive evolutionary landscape of these changes. Compared with Ty1-copia, independent bursts of Ty3-gypsy retrotransposon expansions have occurred with the three largest showing signatures of lineage-specific evolution. The estimated insertion times of 2,213 complete retrotransposons from the top 23 most abundant families reveal divergent life-histories marked by speedy accumulation, decline and extinction that differed radically between species. We hypothesize that this rapid evolution of LTR retrotransposons not only divergently shaped the architecture of rice genomes but also contributed to the process of speciation and diversification of rice.

evolutionary biology

Comparative genomics of beetle-vectored fungal pathogens reveals a reduction in genome size and independent evolution of pathogenicity of two tree pathogens

O_LIGeosmithia morbida is an emerging fungal pathogen which serves as a paradigm for examining the evolutionary processes behind pathogenicity because it is one of two known pathogens within a genus of mostly saprophytic, beetle-associated, fungi. This pathogen causes thousand cankers disease in black walnut trees and is vectored into the host via the walnut twig beetle. G. morbida was first detected in western US and currently threatens the timber industry concentrated in eastern US.\nC_LIO_LIWe sequenced the genomes of G. morbida and two non-pathogenic Geosmithia species and compared these species to other fungal pathogens and nonpathogens to identify genes under positive selection in G. morbida that may be associated with pathogenicity.\nC_LIO_LIG. morbida possesses one of the smallest genomes among the fungal species observed in this study, and one of the smallest fungal pathogen genomes to date. The enzymatic profile is this pathogen is very similar to its relatives.\nC_LIO_LIOur findings indicate that genome reduction is an important adaptation during the evolution of a specialized lifestyle in fungal species that occupy a specific niche, such as beetle vectored tree pathogens. We also present potential genes under selection in G. morbida that could be important for adaptation to a pathogenic lifestyle.\nC_LI

evolutionary biology

A universal, genome-wide guide finder for CRISPR/Cas9 targeting in microbial genomes

BackgroundThe CRISPR/Cas system has significant potential to facilitate gene editing in a variety of bacterial species. CRISPR interference (CRISPRi) and CRISPR activation (CRISPRa) represent modifications of the CRISPR/Cas9 system utilizing a catalytically inactive Cas9 protein for transcription repression or activation, respectively. While CRISPRi and CRISPRa have tremendous potential to systematically investigate gene function in bacteria, no pan-bacterial, genome-wide tools exist for guide discovery. We have created Guide Finder: a customizable, user-friendly program that can design guides for any annotated bacterial genome.\n\nResultsGuide Finder designs guides from NGG PAM sites for any number of genes using an annotated genome and fasta file input by the user. Guides are filtered according to user-defined design parameters and removed if they contain any off-target matches. Iteration with lowered parameter thresholds allows the program to design guides for genes that did not produce guides with the more stringent parameters, a feature unique to Guide Finder. Guide Finder has been tested on a variety of diverse bacterial genomes, on average finding guides for 95% of genes. Moreover, guides designed by the program are functionally useful--focusing on CRISPRi as a potential application--as demonstrated by essential gene knockdown in two staphylococcal species.\n\nConclusionsThrough the large-scale generation of guides, this open-access software will improve accessibility to CRISPR/Cas studies for a variety of bacterial species.

microbiology

A large-scale whole-genome sequencing analysis reveals highly specific genome editing by both Cas9 and Cpf1 nucleases in rice

Targeting specificity has been an essential issue for applying genome editing systems in functional genomics, precise medicine and plant breeding. Understanding the scope of off-target mutations in Cas9 or Cpf1-edited crops is critical for research and regulation. In plants, only limited studies had used whole-genome sequencing (WGS) to test off-target effects of Cas9. However, the cause of numerous discovered mutations is still controversial. Furthermore, WGS based off-target analysis of Cpf1 has not been reported in any higher organism to date. Here, we conducted a WGS analysis of 34 plants edited by Cas9 and 15 plants edited by Cpf1 in T0 and T1 generations along with 20 diverse control plants in rice, a major food crop with a genome size of ~380 Mb. The sequencing depth ranged from 45X to 105X with reads mapping rate above 96%. Our results clearly show that most mutations in edited plants were created by tissue culture process, which caused ~102 to 148 single nucleotide variations (SNVs) and ~32 to 83 insertions/deletions (indels) per plant. Among 12 Cas9 single guide RNAs (sgRNAs) and 3 Cpf1 CRISPR RNAs (crRNAs) assessed by WGS, only one Cas9 sgRNA resulted in off-target mutations in T0 lines at sites predicted by computer programs. Moreover, we cannot find evidence for bona fide off-target mutations due to continued expression of Cas9 or Cpf1 with guide RNAs in T1 generation. Taken together, our comprehensive and rigorous analysis of WGS big data across multiple sample types suggests both Cas9 and Cpf1 nucleases are very specific in generating targeted DNA modifications and off-targeting can be avoided by designing guide RNAs with high specificity.

molecular biology

Functional genomics and programmed genome editing of omega-1 of the blood fluke Schistosoma mansoni

CRISPR/Cas9 based genome editing has yet been reported in parasitic or indeed any species of the phylum Platyhelminthes. We tested this approach by targeting omega-1 ({omega}1) of Schistosoma mansoni as a proof of principle. This secreted ribonuclease is crucial for Th2 priming and granuloma formation, providing informative immuno-pathological readouts for programmed genome editing. Schistosome eggs were either exposed to Cas9 complexed with a synthetic guide RNA (sgRNA) complementary to exon 6 of {omega}1 by electroporation or transduced with pseudotyped lentivirus encoding Cas9 and the sgRNA. Some eggs were also transduced with a single stranded oligodeoxynucleotide donor transgene that encoded six stop codons, flanked by 50 nt-long 5-and 3-microhomology arms matching the predicted Cas9-catalyzed double stranded break (DSB) within {omega}1. CRISPResso analysis of amplicons spanning the DSB revealed [~]4.5% of the reads were mutated by insertions, deletions and/or substitutions, with an efficiency for homology directed repair of 0.19% insertion of the donor transgene. Transcripts encoding {omega}1 were reduced >80% and lysates of {omega}1-edited eggs displayed diminished ribonuclease activity indicative that programmed editing mutated the {omega}1 gene. Whereas lysates of wild type eggs polarized Th2 cytokine responses including IL-4 and IL-5 in human macrophage/T cell co-cultures, diminished levels of the cytokines followed the exposure to lysates of {omega}1-mutated schistosome eggs. Following injection of schistosome eggs into the tail vein of mice, the volume of pulmonary granulomas surrounding {omega}1-mutated eggs was 18-fold smaller than wild type eggs. Programmed genome editing was active in schistosomes, Cas9-catalyzed chromosomal breakage was repaired by homology directed repair and/or non-homologous end joining, and mutation of {omega}1 impeded the capacity of schistosome eggs both to drive Th2 polarization and to provoke formation of pulmonary circumoval granulomas. Knock-out of {omega}1 and the impaired immunological phenotype showcase the novel application of programmed gene editing in and functional genomics for schistosomes.

molecular biology

Genome-wide Association Study of Clinical Features in the Schizophrenia Psychiatric Genomics Consortium: Confirmation of Polygenic Effect on Negative Symptoms

Schizophrenia is a clinically heterogeneous disorder. Proposed revisions in DSM - 5 included dimensional measurement of different symptom domains. We sought to identify common genetic variants influencing these dimensions, and confirm a previous association between polygenic risk of schizophrenia and the severity of negative symptoms. The Psychiatric Genomics Consortium study of schizophrenia comprised 8,432 cases of European ancestry with available clinical phenotype data. Symptoms averaged over the course of illness were assessed using the OPCRIT, PANSS, LDPS, SCAN, SCID, and CASH. Factor analyses of each constituent PGC study identified positive, negative, manic, and depressive symptom dimensions. We examined the relationship between the resultant symptom dimensions and aggregate polygenic risk scores indexing risk of schizophrenia. We performed genome - wide association study (GWAS) of each quantitative traits using linear regression and adjusting for significant effects of sex and ancestry. The negative symptom factor was significantly associated with polygene risk scores for schizophrenia, confirming a previous, suggestive finding by our group in a smaller sample, though explaining only a small fraction of the variance. In subsequent GWAS, we observed the strongest evidence of association for the positive and negative symptom factors, with SNPs in RFX8 on 2q11.2 (P = 6.27x10-8) and upstream of WDR72 / UNC13C on 15q21.3 (P = 7.59x10-8), respectively. We report evidence of association of novel modifier loci for schizophrenia, though no single locus attained established genome - wide significance criteria. As this may have been due to insufficient statistical power, follow - up in additional samples is warranted. Importantly, we replicated our previous finding that polygenic risk explains at least some of the variance in negative symptoms, a core illness dimension.

genetics

Whole genome bisulfite sequencing reveals a sparse, but robust pattern of DNA methylation in the Dictyostelium discoideum genome

DNA methylation, the addition of a methyl (CH3) group to a cytosine residue, is an evolutionarily conserved epigenetic mark involved in a number of different biological functions in eukaryotes, including transcriptional regulation, chromatin structural organization, cellular differentiation and development. In the slime mold Dictyostelium, previous studies have shown the existence of a DNA methyltransferase (DNMA) belonging to the DNMT2 family, but the extent and function of 5-methyl-cytosine in the genome is unclear. Here we present the whole genome DNA methylation profile of Dictyostelium discoideum using deep coverage, replicate sequencing of bisulfite converted gDNA extracted from post-starvation cells. We find an overall very low level of DNA methylation, occurring at only 462 out of the ~7.5 million (0.006%) cytosines in the genome. Despite this sparse profile, significant methylation can be detected at 51 of these sites in replicate experiments, suggesting they are robust targets for DNA methylation. These 5-methyl-cytosines are associated with a broad range of protein-coding genes, tRNA-encoding genes and retrotransposable elements. Our data provides evidence of a minimal, but functional, methylome in Dictyostelium, thereby making Dictyostelium a candidate model organism to further investigate the evolutionary function of DNA methylation.

genomics

Advanced whole genome sequencing and analysis of fetal genomes from amniotic fluid

Amniocentesis is typically performed to identify large chromosomal abnormalities within the fetus. Here we demonstrate that it is feasible to generate an accurate whole genome sequence (WGS) of a fetus from an amniotic sample. DNA from cells and the amniotic fluid were isolated and sequenced from 31 amniocenteses. Concordance of variant calls between the two DNA sources and with parental libraries was high. Two fetal genomes were found to harbor potentially detrimental variants in CHD8 and LRP1, variations in these genes have been associated with Autism Spectrum Disorder (ASD) and Keratosis pilaris atrophicans, respectively. We also discovered drug sensitivities and carrier information of fetuses for a variety of diseases. In this study, we demonstrate for the first time the sequencing of the whole genome of fetuses from amniotic fluid and show that much more information than large chromosomal abnormalities can be gained from an amniocentesis.

genomics

Whole-genome resequencing and pan-transcriptome reconstruction highlight the impact of genomic structural variation on secondary metabolism gene clusters in the grapevine Esca pathogen Phaeoacremonium minimum

The Ascomycete fungus Phaeoacremonium minimum is one of the primary causal agents of Esca, a widespread and damaging grapevine trunk disease. Variation in virulence among Pm. minimum isolates has been reported, but the underlying genetic basis of the phenotypic variability remains unknown. The goal of this study was to characterize intraspecific genetic diversity and explore its potential impact on virulence functions associated with secondary metabolism, cellular transport, and cell wall decomposition. We generated a chromosome-scale genome assembly, using single molecule real-time sequencing, and resequenced the genomes and transcriptomes of multiple isolates to identify sequence and structural polymorphisms. Numerous insertion and deletion events were found for a total of about 1 Mbp in each isolate. Structural variation in this extremely gene dense genome frequently caused presence/absence polymorphisms of multiple adjacent genes, mostly belonging to biosynthetic clusters associated with secondary metabolism. Because of the observed intraspecific diversity in gene content due to structural variation we concluded that a transcriptome reference developed from a single isolate is insufficient to represent the virulence factor repertoire of the species. We therefore compiled a pan-transcriptome reference of Pm. minimum comprising a non-redundant set of 15,245 protein-coding sequences. Using naturally infected field samples expressing Esca symptoms, we demonstrated that mapping of meta-transcriptomics data on a multi-species reference that included the Pm. minimum pan-transcriptome allows the profiling of an expanded set of virulence factors, including variable genes associated with secondary metabolism and cellular transport.

genomics

Stout camphor tree genome fills gaps in understanding of flowering plant genome and gene family evolution

We present reference-quality genome assembly and annotation for the stout camphor tree (SCT; Cinnamomum kanehirae [Laurales, Lauraceae]), the first sequenced member of the Magnoliidae comprising four orders (Laurales, Magnoliales, Canellales, and Piperales) and over 9,000 species. Phylogenomic analysis of 13 representative seed plant genomes indicates that magnoliid and eudicot lineages share more recent common ancestry relative to monocots. Two whole genome duplication events were inferred within the magnoliid lineage, one before divergence of Laurales and Magnoliales and the other within the Lauraceae. Small scale segmental duplications and tandem duplications also contributed to innovation in the evolutionary history of Cinnamomum. For example, expansion of terpenoid synthase subfamilies within the Laurales spawned the diversity of Cinnamomum monoterpenes and sesquiterpenes.

genomics

Population genomic structure and genome-wide linkage disequilibrium in farmed Atlantic salmon (Salmo salar L.) using dense SNP genotypes

Chilean Farmed Atlantic salmon (Salmo salar) populations were established with individuals of both European and North American origins. These populations are expected to be highly genetically differentiated due to evolutionary history and poor gene flow between ancestral populations from different continents. The extent and decay of linkage disequilibrium (LD) among single nucleotide polymorphism (SNP) impacts the implementation of genome-wide association studies and genomic selection and provides relevant information about demographic processes of fish populations. We assessed the population structure and characterized the extent and decay of LD in three Chilean commercial populations of Atlantic salmon with North American (NAM), Scottish (SCO) and Norwegian (NOR) origin. A total of 151 animals were genotyped using a 159K SNP Axiom(R) myDesign Genotyping Array. A total of 40K, 113K and 136 K SNP markers were used for NAM, SCO and NOR populations, respectively. The principal component analysis explained 86.7% of the genetic diversity between populations, clearly discriminating between populations of North American and European origin, and also between European populations. Admixture analysis showed that the Scottish and North American populations likely come from one ancestral population, while the Norwegian population probably originated from more than one. NAM had the lowest effective size, followed by SCO and NOR. Large differences in the LD decay were observed between populations of North American and European origin. A r2 threshold of 0.2 was estimated for marker pairs separated by 8,000 Kb, 42 and 64 Kb in the NAM, NOR and SCO populations, respectively. In this study we show that this SNP panel can be used to detect association between markers and traits of interests and also to capture high-resolution information for genome-enabled predictions. Also, we suggest the feasibility to achieve higher prediction accuracies by using a small SNP data set as was used with the NAM population.

genomics

Genome-wide Ultrabithorax binding analysis reveals highly targeted genomic loci at developmental regulators and a potential connection to Polycomb-mediated regulation

Hox homeodomain transcription factors are key regulators of animal development. They specify the identity of segments along the anterior-posterior body axis in metazoans by controlling the expression of diverse downstream targets, including transcription factors and signaling pathway components. The Drosophila melanogaster Hox factor Ultrabithorax (Ubx) directs the development of thoracic and abdominal segments and appendages, and loss of Ubx function can lead for example to the transformation of third thoracic segment appendages (e.g. halters) into second thoracic segment appendages (e.g. wings), resulting in a characteristic four-wing phenotype. Here we present a Drosophila melanogaster strain with a V5-epitope tagged Ubx allele, which we employed to obtain a high quality genome-wide map of Ubx binding sites using ChIP-seq. We confirm the sensitivity of the V5 ChIP-seq by recovering 7/8 of well-studied Ubx-dependent cis-regulatory regions. Moreover, we show that Ubx binding is predictive of enhancer activity as suggested by comparison with a genome-scale resource of in vivo tested enhancer candidates. We observed densely clustered Ubx binding sites at 12 extended genomic loci that included ANTP-C, BX-C, Polycomb complex genes, and other regulators and the clustered binding sites were frequently active enhancers. Furthermore, Ubx binding was detected at known Polycomb response elements (PREs) and was associated with significant enrichments of Pc and Pho ChIP signals in contrast to binding sites of other developmental TFs. Together, our results show that Ubx targets developmental regulators via strongly clustered binding sites and allow us to hypothesize that regulation by Ubx might involve Polycomb group proteins to maintain specific regulatory states in cooperative or mutually exclusive fashion, an attractive model that combines two groups of proteins with prominent gene regulatory roles during animal development.

Developmental Biology

Whole genome view of the consequences of a population bottleneck using 2926 genome sequences from Finland and United Kingdom

Isolated populations with enrichment of variants due to recent population bottlenecks provide a powerful resource for identifying disease-associated genetic variants and genes. As a model of an isolate population, we sequenced the genomes of 1463 Finnish individuals as part of the Sequencing Initiative Suomi (SISu) Project. We compared the genomic profiles of the 1463 Finns to a sample of 1463 British individuals that were sequenced in parallel as part of the UK10K Project. Whereas there were no major differences in the allele frequency of common variants, a significant depletion of variants in the rare frequency spectrum was observed in Finns when comparing the two populations. On the other hand, we observed >2.1 million variants that were twice as frequent among Finns compared to Britons and 800,000 variants that were more than 10 times more frequent in Finns. Furthermore, in Finns we observed a relative proportional enrichment of variants in the minor allele frequency range between 2 - 5% (p < 2.2x10-16). When stratified by their functional annotations, loss-of-function (LoF) variants showed the highest proportional enrichment in Finns (p = 0.0291). In the noncoding part of the genome, variants in conserved regions (p = 0.002) and promoters (p = 0.01) were also significantly enriched in the Finnish samples. These functional categories represent the highest a priori power for downstream association studies of rare variants using population isolates.

Genetics

Genome-wide chemical mutagenesis screens allow unbiased saturation of the cancer genome and identification of drug resistance mutations.

Drug resistance is an almost inevitable consequence of cancer therapy and ultimately proves fatal for the majority of patients. In many cases this is the consequence of specific gene mutations that have the potential to be targeted to re-sensitize the tumor. The ability to uniformly saturate the genome with point mutations without chromosome or nucleotide sequence context bias would open the door to identify all putative drug resistance mutations in cancer models. Here we describe such a method for elucidating drug resistance mechanisms using genome-wide chemical mutagenesis allied to next-generation sequencing. We show that chemically mutagenizing the genome of cancer cells dramatically increases the number of drug-resistant clones and allows the detection of both known and novel drug resistance mutations. We have developed an efficient computational process that allows for the rapid identification of involved pathways and druggable targets. Such a priori knowledge would greatly empower serial monitoring strategies for drug resistance in the clinic as well as the development of trials for drug resistant patients.

Genetics

A genome-wide haplotype association analysis of major depressive disorder identifies two genome-wide significant haplotypes

Genome-wide association studies using genotype data have had limited success in the identification of variants associated with major depressive disorder (MDD). Haplotype data provide an alternative method for detecting associations between variants in weak linkage disequilibrium with genotyped variants and a given trait of interest. A genome-wide haplotype association study for MDD was undertaken utilising a family-based population cohort, Generation Scotland: Scottish Family Health Study (n = 18 773), as a discovery cohort with UK Biobank used as a population-based cohort replication cohort (n = 25 035). Fine mapping of haplotype boundaries was used to account for overlapping haplotypes potentially tagging the same causal variant. Within the discovery cohort, two haplotypes exceeded genome-wide significance (P < 5 x 10-8) for an association with MDD. One of these haplotypes was nominally significant in the replication cohort (P < 0.05) and was located in 6q21, a region which has been previously associated with bipolar disorder, a psychiatric disorder that is phenotypically and genetically correlated with MDD. Several haplotypes with P < 10-7 in the discovery cohort were located within gene coding regions associated with diseases that are comorbid with MDD. Using such haplotypes to highlight regions for sequencing may lead to the identification of the underlying causal variants.

Genetics

Network analysis links genome-wide phenotypic and transcriptional stress responses in a bacterial pathogen with a large pan-genome.

BackgroundBacteria modulate subcellular processes to handle stressful environments. Genome-wide profiling of gene expression (RNA-Seq) and fitness (Tn-Seq) allows two views of the same genetic network underlying these responses. However, it remains unclear how they combine, enabling a bacterium to overcome a perturbation.\n\nResultsHere we generate RNA-Seq and Tn-Seq profiles in three strains of S. pneumoniae in response to stress defined by different levels of nutrient depletion. These profiles show that genes that change their expression and/or become phenotypically important come from a diverse set of functional categories, and genes that are phenotypically important tend to be highly expressed. Surprisingly, we find that expression and fitness changes rarely occur on the same gene, which we confirmed by over 140 validation experiments. To rationalize these unexpected results we built the first genome-scale metabolic model of S. pneumoniae showing that differential expression and phenotypic importance actually correlate between nearest neighbors, although they are distinctly partitioned into small subnetworks. Moreover, a meta-analysis of 234 S. pneumoniae gene expression studies reveals that essential genes and phenotypically important subnetworks rarely change expression, indicating that they are shielded from transcriptional fluctuations and that a clear distinction exists between transcriptional and phenotypic response networks.\n\nConclusionsWe present a genome-wide computational/experimental approach that contextualizes changes that occur on transcriptomic and phenomic levels in response to stress. Importantly, this highlights the need to connect disparate response networks, for instance in antibiotic target identification, where preferred targets are phenotypically important genes that would be overlooked by transcriptomic analyses alone.

Systems Biology

Population Genomics of the Foothill Yellow-Legged Frog (Rana boylii) and RADseq Parameter Choice for Large-Genome Organisms

Genomic data are useful for attaining high resolution in population genetic studies and have become increasingly available for answering questions in biological conservation. We analyzed RADseq data for the protected foothill yellow-legged frog (Rana boylii) throughout its native range in California and Oregon, including many of the same localities included in an earlier study based on mitochondrial DNA. We recovered five primary clades that correspond to geographic regions within California and Oregon, with better resolution and more spatially consistent patterns than the previous study, confirming the increased resolving power of genomic approaches compared to single-locus analyses. Bayesian clustering, PCA and population differentiation with admixture analyses all indicated that approximately half the range of R. boylii consists of a single, relatively uniform population, while regions in the Sierra Nevada and Central Coast Range of California are deeply differentiated genetically. Additionally, a major methodological challenge for large genome organisms, including many amphibians, is deciding on sequence similarity clustering thresholds for population genetic analyses using RADseq data, and we develop a novel set of metrics that allow researchers to set a sequence similarity threshold that maximizes the separation of paralogous regions while minimizing the oversplitting of naturally occurring allelic variation within loci.

evolutionary biology

The Complete Chloroplast Genome of Dendrobium nobile, an endangered medicinal orchid from Northeast India and its comparison with related chloroplast genomes of Dendrobium species.

The medicinal orchid genus Dendrobium belonging to the Orchidaceae family is the largest genus comprising about 800-1500 species. To better illustrate the species status in the genus Dendrobium, a comparative analysis of 33 newly sequenced chloroplast genomes retrieved from NCBI Refseq database was compared with that of the first complete chloroplast genome of D. nobile from north-east India based on next-generation sequencing methods (Illumina HiSeq 2500-PE150). Our results provide comparative chloroplast genomic information for taxonomical identification, alignment-free phylogenomic inference and other statistical features of Dendrobium plastomes, which can also provide valuable information on their mutational events and sequence divergence.

evolutionary biology