Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,225 records · Page 68Linked to original sources

Sensitive information leakage from functional genomics data: Theoretical quantifications & practical file formats for privacy preservation

The generation of functional genomics datasets is surging, as they provide insight into gene regulation and organismal phenotypes (e.g., genes upregulated in cancer). The intention of functional genomics experiments is not necessarily to study genetic variants, yet they pose privacy concerns due to their use of next-generation sequencing. Moreover, there is a great incentive to share raw reads for better analyses and general research reproducibility. Thus, we need new modes of sharing beyond traditional controlled-access models. Here, we develop a data-sanitization procedure allowing raw functional genomics reads to be shared while minimizing privacy leakage, thus enabling principled privacy-utility trade-offs. It works with traditional Illumina-based assays and newer technologies such as 10x single-cell RNA-sequencing. The procedure depends on quantifying the privacy leakage in reads by statistically linking study participants to known individuals. We carried out these linkages using data from highly accurate reference genomes and more realistic environmental samples.

bioinformatics

Age-related late-onset disease heritability patterns and implications for genome-wide association studies

BackgroundGenome-wide association studies and other computational biology techniques are gradually discovering the causal gene variants that contribute to late-onset human diseases. After more than a decade of genome-wide association study efforts, these can account for only a fraction of the heritability implied by familial studies, the so-called \"missing heritability\" problem.\n\nMethodsComputer simulations of polygenic late-onset diseases in an aging population have quantified the risk allele frequency decrease at older ages caused by individuals with higher polygenic risk scores becoming ill proportionately earlier. This effect is most prominent for diseases characterized by high cumulative incidence and high heritability, examples of which include Alzheimers disease, coronary artery disease, cerebral stroke, and type 2 diabetes.\n\nResultsThe incidence rate for late-onset diseases grows exponentially for decades after early onset ages, guaranteeing that the cohorts used for genome-wide association studies overrepresent older individuals with lower polygenic risk scores, whose disease cases are disproportionately due to environmental causes such as old age itself. This mechanism explains the decline in clinical predictive power with age and the lower discovery power of familial studies of heritability and genome-wide association studies. It also explains the relatively constant-with-age heritability found for late-onset diseases of lower prevalence, exemplified by cancers.\n\nConclusionsFor late-onset polygenic diseases showing high cumulative incidence together with high initial heritability, rather than using relatively old age-matched cohorts, study cohorts combining the youngest possible cases with the oldest possible controls may significantly improve the discovery power of genome-wide association studies.

genetics

Switchable genome editing via genetic code expansion

Multiple applications of genome editing by CRISPR-Cas9 necessitate stringent regulation and Cas9 variants have accordingly been generated whose activity responds to small ligands, temperature or light. However, these approaches are often impracticable, for example in clinical therapeutic genome editing in situ or gene drives in which environmentally-compatible control is paramount. With this in mind, we have developed heritable Cas9-mediated mammalian genome editing that is acutely controlled by the cheap lysine derivative, Lys(Boc) (BOC). Genetic code expansion permitted non-physiological BOC incorporation such that Cas9 (Cas9BOC) was expressed in a full-length, active form in cultured somatic cells only after BOC exposure. Stringently BOC-dependent, heritable editing of transgenic and native genomic loci occurred when Cas9BOC was expressed at the onset of mouse embryonic development from cRNA or Cas9BOC transgenic females. The tightly controlled Cas9 editing system reported here promises to have broad applications and is a first step towards purposed, spatiotemporal gene drive regulation over large geographical ranges.

developmental biology

Developing a network view of type 2 diabetes risk pathways through integration of genetic, genomic and functional data

Genome wide association studies (GWAS) have identified several hundred susceptibility loci for Type 2 Diabetes (T2D). One critical, but unresolved, issue concerns the extent to which the mechanisms through which these diverse signals influencing T2D predisposition converge on a limited set of biological processes. However, the causal variants identified by GWAS mostly fall into non-coding sequence, complicating the task of defining the effector transcripts through which they operate. Here, we describe implementation of an analytical pipeline to address this question. First, we integrate multiple sources of genetic, genomic, and biological data to assign positional candidacy scores to the genes that map to T2D GWAS signals. Second, we introduce genes with high scores as seeds within a network optimization algorithm (the asymmetric prize-collecting Steiner Tree approach) which uses external, experimentally-confirmed protein-protein interaction (PPI) data to generate high confidence subnetworks. Third, we use GWAS data to test the T2D-association enrichment of the \"non-seed\" proteins introduced into the network, as a measure of the overall functional connectivity of the network. We find: (a) non-seed proteins in the T2D protein-interaction network so generated (comprising 705 nodes) are enriched for association to T2D (p=0.0014) but not control traits; (b) stronger T2D-enrichment for islets than other tissues when we use RNA expression data to generate tissue-specific PPI networks; and (c) enhanced enrichment (p=3.9xl0-5) when we combine analysis of the islet-specific PPI network with a focus on the subset of T2D GWAS loci which act through defective insulin secretion. These analyses reveal a pattern of non-random functional connectivity between causal candidate genes atT2D GWAS loci, and highlight the products of genes including YWHAG, SMAD4 or CDK2 as contributors to T2D-relevant islet dysfunction. The approach we describe can be applied to other complex genetic and genomic data sets, facilitating integration of diverse data types into disease-associated networks.\n\nAuthor summaryWe were interested in the following question: as we discover more and more genetic variants associated with a complex disease, such as type 2 diabetes, will the biological pathways implicated by those variants proliferate, or will the biology converge onto a more limited set of aetiological processes? To address this, we first took the 1895 genes that map to ~100 type 2 diabetes association signals, and pruned these to a set of 451 for which combined genetic, genomic and biological evidence assigned the strongest candidacy with respect to type 2 diabetes pathogenesis. We then sought to maximally connect these genes within a curated protein-protein interaction network. We found that proteins brought into the resulting diabetes interaction network were themselves enriched for diabetes association signals as compared to appropriate control proteins. Furthermore, when we used tissue-specific RNA abundance data to filter the generic protein-protein network, we found that the enrichment for type 2 diabetes association signals was enhanced within a network filtered for pancreatic islet expression, particularly when we selected the subset of diabetes association signals acting through reduced insulin secretion. Our data demonstrate convergence of the biological processes involved in type 2 diabetes pathogenesis and highlight novel contributors.

bioinformatics

Miniscule differences between the sex chromosomes in the giant genome of a salamander, Ambystoma mexicanum

In the Mexican axolotl (Ambystoma mexicanum) sex is known to be determined by a single Mendelian factor, yet the sex chromosomes of this model salamander do not exhibit morphological differentiation that is typical of many vertebrate taxa that possess a single sex-determining locus. Differentiated sex chromosomes are thought to evolve rapidly in the context of a Mendelian sex-determining gene and, therefore, undifferentiated chromosomes provide an exceptional opportunity to reconstruct early events in sex chromosome evolution. Whole chromosome sequencing, whole genome resequencing (48 individuals from a backcross of axolotl and tiger salamander) and in situ hybridization were used to identify a homomorphic chromosome that carries an A. mexicanum sex determining factor and identify sequences that are present only on the W chromosome. Altogether, these sequences cover ~300 kb, or roughly 1/100,000th of the ~32 Gb genome. Notably, these W-specific sequences also contain a recently duplicated copy of the ATRX gene: a known component of mammalian sex-determining pathways. This gene (designated ATRW) is one of the few functional (non-repetitive) genes in the chromosomal segment and maps to the tip of chromosome 9 near the marker E24C3, which was previously found to be linked to the sex-determining locus. These analyses provide highly predictive markers for diagnosing sex in A. mexicanum and identify ATRW as a strong candidate for the primary sex determining locus or alternately a strong candidate for a recently acquired, sexually antagonistic gene.\n\nAUTHOR SUMMARYSex chromosomes are thought to follow fairly stereotypical evolutionary trajectories that result in differentiation of sex-specific chromosomes. In the salamander A. mexicanum (the axolotl), sex is determined by a single Mendelian locus, yet the sex chromosomes are essentially undifferentiated, suggesting that these sex chromosomes have recently acquired a sex locus and are in the early stages of differentiating. Although Mendelian sex determination was first reported for the axolotl more than 70 years ago, no sex-specific sequences have been identified for this important model species. Here, we apply new technologies and approaches to identify and validate a tiny region of female-specific DNA within the gigantic genome of the axolotl (1/100,000th of the genome). This region contains a limited number of genes, including a duplicate copy of the ATRX gene which, has been previously shown to contribute to mammalian sex determination. Our analyses suggest that this gene, which we refer to as ATRW, evolved from a recent duplication and presents a strong candidate for the primary sex determining factor of the axolotl, or alternately a recently evolved sexually antagonistic gene.

evolutionary biology

Whole genome sequencing of antimicrobial-resistant Shigella sonnei associated with infection acquired from international and domestic setting reveals gene allelic variants that predict Global Lineages

Shigella spp. are a major cause of gastroenteritis worldwide, and S. sonnei is the most common species isolated within the United States. Recently, advancements in technology have made whole genome sequencing (WGS) readily available, and as such, laboratories are moving to implement WGS in outbreak analysis, surveillance, and antimicrobial resistance (AMR) monitoring of major foodborne pathogens. Accordingly, our study examined a collection of 22 antimicrobial resistant S. sonnei isolates from patients who either acquired the infections within the United States or when travelling to international locations between 2009 to 2014. We applied WGS to investigate both the relatedness of these isolates and the genetic determinants of AMR to address the phenotypic differences seen in previous observations. We analyzed the phylogeny of these strains and observed segmentation based on the previously described Global Lineages of S. sonnei. Following these results, 17 gene sequences with lineage specific single nucleotide polymorphisms (SNPs) were identified and developed into a lineage prediction test to determine the Global Lineage of uncharacterized S. sonnei, which accurately predicted phylogenetic segmentation and additionally showed specificity for S. sonnei genomes (97% accuracy, 38/39 genomes). Lastly, to determine differences between either the international or domestic isolates or between the Global Lineages, the AMR determinants were identified. We found a variety of AMR determinants within the genomes, and while the international and domestic S. sonnei carried similar resistance determinants, differences between Global Lineages were observed.

microbiology

An ancient genome duplication in the speciose reef-building coral genus, Acropora

Whole-genome duplication (WGD) has been recognized as a significant evolutionary force in the origin and diversification of vertebrates, plants, and other organisms. Acropora, one of the most speciose reef-building coral genera, responsible for creating spectacular but increasingly threatened marine ecosystems, is suspected to have originated by polyploidy, yet there is no genetic evidence to support this hypothesis. Using comprehensive phylogenomic and comparative genomic approaches, we analyzed five Acropora genomes and an Astreopora genome (Scleractinia: Acroporidae) to show that a WGD event likely occurred between 27.9 and 35.7 Million years ago (Mya) in the most recent common ancestor of Acropora, concurrent with a massive worldwide coral extinction. We found that duplicated genes became highly enriched in gene regulation functions, some of which are involved in stress responses. The different functional clusters of duplicated genes are related to the divergence of gene expression patterns during development. Some gene duplications of proteinaceous toxins were generated by WGD in Acropora compared with other Cnidarian species. Collectively, this study provides evidence for an ancient WGD event in corals and it helps to explain the origin and diversification of Acropora.

evolutionary biology

A cell cycle-coordinated nuclear compartment for Polymerase II transcription encompasses the earliest gene expression before global genome activation

Most metazoan embryos commence development with rapid cleavages without zygotic gene expression and their genome activation is delayed until the mid-blastula transition (MBT). However, a set of genes escape global repression during the extremely fast cell cycles, which lack gap phases and their transcription is activated before the MBT. Here we describe the formation and the spatio-temporal dynamics of a distinct transcription compartment, which encompasses the earliest detectable transcription during the first wave of genome activation. Simultaneous 4D imaging of expression of pri-miR430 and zinc finger genes by a novel, native transcription imaging approach reveals a pair of shared transcription compartments regulated by homolog chromosome organisation. These nuclear compartments carry the majority of nascent RNAs and transcriptionally active Polymerase II, are depleted of compact chromatin and represent the main sites for detectable transcription before MBT. We demonstrate that transcription occurs in the S-phase of the cleavage cycles and that the gradual slowing of these cell cycles are permissive to transcription before global genome activation. We propose that the demonstrated transcription compartment is part of the regulatory architecture of nucleus organisation, and provides a transcriptionally competent, supporting environment to facilitate early escape from the general nuclear repression before global genome activation.

developmental biology

Kipoi: accelerating the community exchange and reuse of predictive models for genomics

Advanced machine learning models applied to large-scale genomics datasets hold the promise to be major drivers for genome science. Once trained, such models can serve as a tool to probe the relationships between data modalities, including the effect of genetic variants on phenotype. However, lack of standardization and limited accessibility of trained models have hampered their impact in practice. To address this, we present Kipoi, a collaborative initiative to define standards and to foster reuse of trained models in genomics. Already, the Kipoi repository contains over 2,000 trained models that cover canonical prediction tasks in transcriptional and post-transcriptional gene regulation. The Kipoi model standard grants automated software installation and provides unified interfaces to apply and interpret models. We illustrate Kipoi through canonical use cases, including model benchmarking, transfer learning, variant effect prediction, and building new models from existing ones. By providing a unified framework to archive, share, access, use, and build on models developed by the community, Kipoi will foster the dissemination and use of machine learning models in genomics.

bioinformatics

The genomic and molecular basis of response to selection for longer limbs in mice

Evolutionary studies are often limited by missing data that are critical to understanding the history of selection. Selection experiments, which reproduce rapid evolution under controlled conditions, are excellent tools to study how genomes evolve under strong selection. Here we present a genomic dissection of the Longshanks selection experiment, in which mice were selectively bred over 20 generations for longer tibiae relative to body mass, resulting in 13% longer tibiae in two replicate lines. We synthesized evolutionary theory, genome sequences and molecular genetics to understand the selection response and found that it involved both polygenic adaptation and discrete loci of major effect, with the strongest loci likely to be selected in parallel between replicates. We show that selection may favor de-repression of bone growth through inactivation of two limb enhancers of an inhibitor, Nkx3-2. Our integrative genomic analyses thus show that it is possible to connect individual base-pair changes to the overall selection response.

evolutionary biology

Population size history from short genomic scaffolds: how short is too short?

The Pairwise Sequentially Markov Coalescent (PSMC), and its extension PSMC', model past population sizes from a single diploid genome. Both models have been widely applied, even to organisms with scaffold-level genome reference assemblies of limited contiguity. However it is unclear how PSMC and PSMC' perform on short scaffolds. We evaluated psmc and msmc, implementations of the PSMC and PSMC' models respectively, on simulated genomes with low contiguity, and compared results to those from fully contiguous data. Simulations with scaffolds from 100 Mb to 10 kb revealed that psmc maintains high accuracy down to lengths of 100 kb, while msmc is accurate down to 1 Mb. The discrepancy is not due to differing models, but stems from an implementation detail of msmc--homozygous tracts at the ends of scaffolds are discarded, making msmc unreliable for low contiguity genomes. We recommend excluding data that are aligned to shorter scaffolds when undertaking demographic inference.

genetics

A Swiftian Voyage from Brobdingnag to Lilliput: Freshwater Planctomycetes drifting towards the poles of the genome size spectrum

Freshwater environments teem with microbes. Currently, our apprehension of evolutionary ecology of freshwater bacteria is hampered by the difficulty to establish organism models for the most representative clades. To circumvent the bottlenecks inherent to the cultivation-based techniques, we applied ecogenomics approaches in order to unravel the evolutionary history and the processes that drive genome architecture in hallmark freshwater lineages from Planctomycetes phylum. The evolutionary history inferences showed that sediment/soil Planctomycetes transitioned to aquatic environments were, through processes mostly associated with reductive genome evolution, gave rise to new freshwater-specific clades. The most successful lineage was found to simultaneously have the most specialized lifestyle (increased regulatory genetic circuits; metabolism tuned for mineralization of proteinaceous sinking aggregates; psychrotrophic behavior) and to harbor the smallest genomes, highlighting a genomic architecture shaped by niche-directed evolution.

microbiology

Identification of deleterious and regulatory genomic variations in known asthma loci

BackgroundCandidate gene and genome-wide association studies have identified hundreds of asthma risk loci. The majority of associated variants, however, are not known to have any biological function and are believed to represent markers rather than true causative mutations. We hypothesized that many of these associated markers are in linkage disequilibrium (LD) with the elusive causative variants.\n\nMethodsWe compiled a comprehensive list of 447 asthma-associated variants previously reported in candidate gene and genome-wide association studies. Next, we identified all sequence variants located within the 304 unique genes using whole-genome sequencing data from the 1000 Genomes Project. Then, we calculated the LD between known asthma variants and the sequence variants within each gene. LD variants identified were then annotated to determine those that are potentially deleterious and/or functional (i.e. coding or regulatory effects on the encoded transcript or protein).\n\nResultsWe identified 10,048 variants in LD (r2 > 0.6) with known asthma variants. Annotations of these LD variants revealed that several have potentially deleterious effects including frameshift, alternate splice site, stop-lost, and missense. Moreover, 24 of the LD variants have been reported to regulate gene expression as expression quantitative trait loci (eQTLs).\n\nConclusionsThis study is proof of concept that many of the genetic loci previously associated with complex diseases such as asthma are not causative but represent markers of disease, which are in LD with the elusive causative variants. We hereby report a number of potentially deleterious and regulatory variants that are in LD with the reported asthma loci. These reported LD variants could account for the original association signals with asthma and represent the true causative mutations at these loci.

bioinformatics

Optimized Cas9 expression systems for highly efficient Arabidopsis genome editing facilitate isolation of complex alleles in a single generation

Genetic resources for the model plant Arabidopsis comprise mutant lines defective in almost any single gene in reference accession Columbia. However, gene redundancy and/or close linkage often render it extremely laborious or even impossible to isolate a desired line lacking a specific function or set of genes from segregating populations. Therefore, we here evaluated strategies and efficiencies for the inactivation of multiple genes by Cas9-based nucleases and multiplexing. In first attempts, we succeeded in isolating a mutant line carrying a 70 kb deletion, which occurred at a frequency of ~1.6% in the T2 generation, through PCR-based screening of numerous individuals. However, we failed to isolate a line lacking Lhcb1 genes, which are present in five copies organized at two loci in the Arabidopsis genome. To improve efficiency of our Cas9-based nuclease system, regulatory sequences controlling Cas9 expression levels and timing were systematically compared. Indeed, use of DD45 and RPS5a promoters improved efficiency of our genome editing system by approximately 25-30-fold in comparison to the previous ubiquitin promoter. Using an optimized genome editing system with RPS5a promoter-driven Cas9, putatively quintuple mutant lines lacking detectable amounts of Lhcb1 protein represented approximately 30% of T1 transformants. These results show how improved genome editing systems facilitate the isolation of complex mutant alleles, previously considered impossible to generate, at high frequency even in a single (T1) generation.

plant biology

Prediction and analysis of skin cancer progression using genomics profiles of patients

Metastatic state of the Skin Cutaneous Melanoma (SKCM) has led to high mortality rate worldwide. Previously, various studies have revealed the association of the metastatic melanoma with the diminished survival rate in comparison to primary tumors. Thus, prediction of melanoma at primary tumor state is crucial to employ optimal therapeutic strategy for prolonged survival of patients. The RNA, miRNA and methylation data of The Cancer Genome Atlas (TCGA) cohort of SKCM is comprehensively analysed to recognize key genomic features that can categorize various states of metastatic tumors from primary tumors with high precision. Subsequently, various prediction models were developed using filtered genomic features implementing various machine learning techniques to classify these primary tumors from metastatic tumors. The SVC model (with class weight and RBF kernel) developed using 17 mRNA features achieved maximum MCC 0.73 with sensitivity, specificity and accuracy 89.19%, 90.48% and 89.47% respectively on independent validation dataset. Our study reveals that gene expression based features performs better than features obtained from miRNA profiling and epigenomic profiling. Our analysis shows that the expression of genes C7, MMP3, KRT14, KRT17, MASP1, and miRNA hsa-mir-205 and hsa-mir-203a are among the key genomic features that may substantially contribute to the oncogenesis of melanoma even on the basis of simple expression threshold. The major prediction models and analysis modules to predict metastatic and primary tumor samples of SKCM are available from a webserver, CancerSPP (http://webs.iiitd.edu.in/raghava/cancerspp/).

bioinformatics

NucBreak: Location of structural errors in a genome assembly by using paired-end Illumina reads

BackgroundAdvances in whole genome sequencing strategies have provided the opportunity for genomic and comparative genomic analysis of a vast variety of organisms. The analysis results are highly dependent on the quality of the genome assemblies used. Assessment of the assembly accuracy may significantly increase the reliability of the analysis results and is therefore of great importance.\n\nResultsHere, we present a new tool called NucBreak aimed at detecting structural errors in assemblies, including insertions, deletions, duplications, inversions, and different inter-and intra-chromosomal rearrangements. NucBreak analyses the alignments of reads properly mapped to an assembly and exploits information about the alternative read alignments. We have compared NucBreak with other existing assembly accuracy assessment tools, namely Pilon, REAPR, and FRCbam as well as with several structural variant detection tools, including BreakDancer, Lumpy, and Wham, by using both simulated and real datasets.\n\nConclusionsThe benchmarking results have shown that NucBreak in general predicts assembly errors of different types and sizes with relatively high sensitivity and with higher precision than the other tools. Such a balance between sensitivity and precision makes NucBreak a good alternative to the existing assembly accuracy assessment tools and SV detection tools. NucBreak is freely available at https://github.com/uio-bmi/NucBreak under the MPL license.

bioinformatics

Endogenous viral elements are widespread in arthropod genomes and commonly give rise to piRNAs

Arthropod genomes contain sequences derived from integrations of DNA and non-retroviral RNA viruses. These sequences, known as endogenous viral elements (EVEs), have been acquired over the course of evolution and have been proposed to serve as a record of past viral infection. Recent evidence indicates that EVEs can function as templates for the biogenesis of PIWI-interacting RNAs (piRNAs) in some mosquito species and cell lines, raising the possibility that EVEs may function as a source of immunological memory in these organisms. However, whether EVEs are capable of acting as templates for piRNA production in other arthropod species is unknown. Here we used publically available genome assemblies and small RNA sequencing datasets to characterize the repertoire and function of EVEs across 48 arthropod genomes. We found that EVEs are widespread in arthropod genomes and primarily correspond to unclassified ssRNA viruses and viruses belonging to the Rhabdoviridae and Parvoviridae families. Additionally, EVEs were enriched in piRNA clusters in a majority of species and we found that production of primary piRNAs from EVEs is common, particularly for EVEs located within piRNA clusters. While we found evidence suggesting that piRNAs mapping to a number of EVEs are produced via the ping-pong cycle, potentially pointing towards a role for EVE-derived piRNAs during viral infection, limited nucleotide identity between currently described viruses and EVEs identified here likely limits the extent to which this process plays a role during infection with known viruses.

microbiology

OMA standalone: orthology inference among public and custom genomes and transcriptomes

Genomes and transcriptomes are now typically sequenced by individual labs, but analysing them often remains challenging. One essential step in many analyses lies in identifying orthologs--corresponding genes across multiple species--but this is far from trivial. The OMA (Orthologous MAtrix) database is a leading resource for identifying orthologs among publicly available, complete genomes. Here, we describe the OMA pipeline available as a standalone program for Linux and Mac. When run on a cluster, it has native support for the LSF, SGE, PBS Pro, and Slurm job schedulers and can scale up to thousands of parallel processes. Another key feature of OMA standalone is that users can combine their own data with existing public data by exporting genomes and pre-computed alignments from the OMA database, which currently contains over 2100 complete genomes. We compare OMA standalone to other methods in the context of phylogenetic tree inference, by inferring a phylogeny of the Lophotrochozoa, a challenging clade within the Protostomes. We also discuss other potential applications of OMA standalone, including identifying gene families having undergone duplications/losses in specific clades, and identifying potential drug targets in non-model organisms. OMA Standalone is available at http://omabrowser.org/standalone under the permissible open source Mozilla Public License Version 2.0.

bioinformatics