Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,135 records · Page 63Linked to original sources

Evaluating the clinical validity of gene-disease associations: an evidence-based framework developed by the Clinical Genome Resource

With advances in genomic sequencing technology, the number of reported gene-disease relationships has rapidly expanded. However, the evidence supporting these claims varies widely, confounding accurate evaluation of genomic variation in a clinical setting. Despite the critical need to differentiate clinically valid relationships from less well-substantiated relationships, standard guidelines for such evaluation do not currently exist. The NIH-funded Clinical Genome Resource (ClinGen) has developed a framework to define and evaluate the clinical validity of gene-disease pairs across a variety of Mendelian disorders. In this manuscript we describe a proposed framework to evaluate relevant genetic and experimental evidence supporting or contradicting a gene-disease relationship, and the subsequent validation of this framework using a set of representative gene-disease pairs. The framework provides a semi-quantitative measurement for the strength of evidence of a gene-disease relationship which correlates to a qualitative classification: \"Definitive\", \"Strong\", \"Moderate\", \"Limited\", \"No Reported Evidence\" or \"Conflicting Evidence.\" Within the ClinGen structure, classifications derived using this framework are reviewed and confirmed or adjusted based on clinical expertise of appropriate disease experts. Detailed guidance for utilizing this framework and access to the curation interface is available on our website. This evidence-based, systematic method to assess the strength of gene-disease relationships will facilitate more knowledgeable utilization of genomic variants in clinical and research settings.

genetics

The Sequence of 1504 Mutants in the Model Rice Variety Kitaake Facilitates Rapid Functional Genomic Studies

The availability of a whole-genome sequenced mutant population and the cataloging of mutations of each line at a single-nucleotide resolution facilitates functional genomic analysis. To this end, we generated and sequenced a fast-neutron-induced mutant population in the model rice cultivar Kitaake (Oryza sativa L. ssp. japonica), which completes its life cycle in 9 weeks. We sequenced 1,504 mutant lines at 45-fold coverage and identified 91,513 mutations affecting 32,307 genes, 58% of all rice genes. We detected an average of 61 mutations per line. Mutation types include single base substitutions, deletions, insertions, inversions, translocations, and tandem duplications. We observed a high proportion of loss-of-function mutations. Using this mutant population, we identified an inversion affecting a single gene as the causative mutation for the short-grain phenotype in one mutant line with a small segregating population. This result reveals the usefulness of the resource for efficient identification of genes conferring specific phenotypes. To facilitate public access to this genetic resource, we established an open access database called KitBase that provides access to sequence data and seed stocks, enabling rapid functional genomic studies of rice.\n\nOne-sentence summaryWe have sequenced 1,504 mutant lines generated in the short life cycle rice variety Kitaake (9 weeks) and established a publicly available database, enabling rapid functional genomic studies of rice.

plant biology

Conflicting evolutionary histories of the mitochondrial and nuclear genomes in New World Myotis

The rapid diversification of Myotis bats into more than 100 species is one of the most extensive mammalian radiations available for study. Efforts to understand relationships within Myotis have primarily utilized mitochondrial markers and trees inferred from nuclear markers lacked resolution. Our current understanding of relationships within Myotis is therefore biased towards a set of phylogenetic markers that may not reflect the history of the nuclear genome. To resolve this, we sequenced the full mitochondrial genomes of 37 representative Myotis, primarily from the New World, in conjunction with targeted sequencing of 3,648 ultraconserved elements (UCEs). We inferred the phylogeny and explored the effects of concatenation and summary phylogenetic methods, as well as combinations of markers based on informativeness or levels of missing data, on our results. Of the 294 phylogenies generated from the nuclear UCE data, all are significantly different from phylogenies inferred using mitochondrial genomes. Even within the nuclear data, quartet frequencies indicate that around half of all UCE loci conflict with the estimated species tree. Several factors can drive such conflict, including incomplete lineage sorting, introgressive hybridization, or even phylogenetic error. Despite the degree of discordance between nuclear UCE loci and the mitochondrial genome and among UCE loci themselves, the most common nuclear topology is recovered in one quarter of all analyses with strong nodal support. Based on these results, we re-examine the evolutionary history of Myotis to better understand the phenomena driving their unique nuclear, mitochondrial, and biogeographic histories.

evolutionary biology

LASSIM - a network inference toolbox for genome-wide mechanistic modeling

Recent technological advancements have made time-resolved, quantitative, multi-omics data available for many model systems, which could be integrated for systems pharmacokinetic use. Here, we present large-scale simulation modeling (LASSIM), which is the first general mathematical tool for performing large-scale inference using mechanistically defined ordinary differential equations (ODE) for gene regulatory networks (GRNs). LASSIM integrates structural knowledge about regulatory interactions and non-linear equations with multiple steady states and dynamic response expression datasets. The rationale behind LASSIM is that biological GRNs can be simplified using a limited subset of core genes that are assumed to regulate all other gene transcription events in the network. LASSIM models are built in two steps, where each step can integrate multiple data-types, and the method is implemented as a general-purpose toolbox using the PyGMo Python package to make the most of multicore computers and high performance clusters, and is available at https://gitlab.com/Gustafsson-lab/lassim. As a method, LASSIM first infers a non-linear ODE system of the pre-specified core genes. Second, LASSIM optimizes the parameters that models the regulation of peripheral genes by core-system genes in parallel. We showed the usefulness of this method by applying LASSIM to infer a large-scale nonlinear model of naive Th2 differentiation, made possible by integrating Th2 specific bindings, time-series and six public and six novel siRNA-mediated knock-down experiments. ChIP-seq showed significant overlap for all tested transcription factors. Next, we performed novel time-series measurements of total T-cells during differentiation towards Th2 and verified that our LASSIM model could monitor those data significantly better than comparable models that used the same Th2 bindings. In summary, the LASSIM toolbox opens the door to a new type of model-based data analysis that combines the strengths of reliable mechanistic models with truly systems-level data. We exemplified the advantage by inferring the first mechanistically motivated genome-wide model of the Th2 transcription regulatory system, which plays an important role in the progression of immune related diseases.\n\nAuthor summaryThere are excellent methods to mathematically model time-resolved biological data on a small scale using accurate mechanistic models. Despite the rapidly increasing availability of such data, mechanistic models have not been applied on a genome-wide level due to excessive runtimes and the non-identifiability of model parameters. However, genome-wide, mechanistic models could potentially answer key clinical questions, such as finding the best drug combinations to induce an expression change from a disease to a healthy state.\n\nWe present LASSIM, which is a toolbox built to infer parameters within mechanistic models on a genomic scale. This is made possible due to a property shared across biological systems, namely the existence of a subset of master regulators, here denoted the core system. The introduction of a core system of genes simplifies the inference into small solvable subproblems, and implies that all main regulatory actions on peripheral genes come from a small set of regulator genes. This separation allows substantial parts of computations to be solved in parallel, i.e. permitting the use of a computer cluster, which substantially reduces the time required for the computation to finish.

systems biology

Rapid genome recoding by iterative recombineering of synthetic DNA

Genome recoding will provide a deeper understanding of genetics and transform biotechnology. We bypass the reliance of previous genome recoding methods on site-specific enzymes and demonstrate a rapid recombineering based strategy for writing genomes by Stepwise Integration of Rolling Circle Amplified Segments (SIRCAS). We installed the largest number of codon substitutions in a single organism yet published, creating a strain of Salmonella typhimurium with 1557 leucine codon changes across 200 kb of the genome.

synthetic biology

Comparison of methods that use whole genome data to estimate the heritability and genetic architecture of complex traits.

Heritability, h2, is a foundational concept in genetics, critical to understanding the genetic basis of complex traits. Recently-developed methods that estimate heritability from genotyped SNPs, h2 SNP, explain substantially more genetic variance than genome-wide significant loci, but less than classical estimates from twins and families. However, h2SNP estimates have yet to be comprehensively compared under a range of genetic architectures, making it difficult to draw conclusions from sometimes conflicting published estimates. Here, we used thousands of real whole genome sequences to simulate realistic phenotypes under a variety of genetic architectures, including those from very rare causal variants. We compared the performance of ten methods across different types of genotypic data (commercial SNP array positions, whole genome sequence variants, and imputed variants) and under differing causal variant frequencies, levels of stratification, and relatedness thresholds. These results provide guidance in interpreting past results and choosing optimal approaches for future studies. We then chose two methods (GREML-MS and GREML-LDMS) that best estimated overall h2SNP and the causal variant frequency spectra to six phenotypes in the UK Biobank using imputed genome-wide variants. Our results suggest that as imputation reference panels become larger and more diverse, estimates of the frequency distribution of causal variants will become increasingly unbiased and the vast majority of trait narrow-sense heritability will be accounted for.

genetics

The Sentieon Genomics Tools - A fast and accurate solution to variant calling from next-generation sequence data

In the past six years worldwide capacity for human genome sequencing has grown by more than five orders of magnitude, with costs falling by nearly two orders of magnitude over the same period [1], [2]. The rapid expansion in the production of next-generation sequence data and the use of these data in a wide range of new applications has created a need for improved computational tools for data processing. The Sentieon Genomics tools provide an optimized reimplementation of the most accurate pipelines for calling variants from next-generation sequence data, resulting in more than a 10-fold increase in processing speed while providing identical results to best practices pipelines. Here we demonstrate the consistency and improved performance of Sentieons tools relative to BWA, GATK, MuTect, and MuTect2 through analysis of publicly available human exome, low-coverage genome, and high-depth genome sequence data.

bioinformatics

Sexual antagonism exerts evolutionarily persistent genomic constraints on sexual differentiation in Drosophila melanogaster

The evolution of sexual dimorphism is constrained by a shared genome, leading to sexual antagonism where different alleles at given loci are favoured by selection in males and females. Despite its wide taxonomic incidence, we know little about the identity, genomic location and evolutionary dynamics of antagonistic genetic variants. To address these deficits, we use sex-specific fitness data from 202 fully sequenced hemiclonal D. melanogaster fly lines to perform a genome-wide association study of sexual antagonism. We identify ~230 chromosomal clusters of candidate antagonistic SNPs. In contradiction to classic theory, we find no clear evidence that the X chromosome is a hotspot for sexually antagonistic variation. Characterising antagonistic SNPs functionally, we find a large excess of missense variants but little enrichment in terms of gene function. We also assess the evolutionary persistence of antagonistic variants by examining extant polymorphism in wild D. melanogaster populations. Remarkably, antagonistic variants are associated with multiple signatures of balancing selection across the D. melanogaster distribution range, indicating widespread and evolutionarily persistent (>10,000 years) genomic constraints. Based on our results, we propose that antagonistic variation accumulates due to constraints on the resolution of sexual conflict over protein coding sequences, thus contributing to the long-term maintenance of heritable fitness variation.

evolutionary biology

GenomicDataCommons: a Bioconductor Interface to the NCI Genomic Data Commons

The National Cancer Institute (NCI) Genomic Data Commons (Grossman et al. 2016, https://gdc.cancer.gov/) provides the cancer research community with an open and unified repository for sharing and accessing data across numerous cancer studies and projects via a high-performance data transfer and query infrastructure. The Bioconductor project (Huber et al. 2015) is an open source and open development software project built on the R statistical programming environment (R Core Team 2016). A major goal of the Bioconductor project is to facilitate the use, analysis, and comprehension of genomic data. The GenomicDataCommons Bioconductor package provides basic infrastructure for querying, accessing, and mining genomic datasets available from the GDC. We expect that Bioconductor developer and bioinformatics community will build on the GenomicDataCommons package to add higher-level functionality and expose cancer genomics data to many state-of-the-art bioinformatics methods available in Bioconductor.\n\nAvailabilityhttps://github.com/Bioconductor/GenomicDataCommons (& soon in Bioconductor).

bioinformatics

Draft genome of the Heterotardigrade Milnesium tardigradum sheds light on ecdysozoan evolution

Tardigrades are among the most stress tolerant animals and survived even unassisted exposure to space in low earth orbit. Still, the adaptations leading to these unusual physiological features remain unclear. Even the phylogenetic position of this phylum within the Ecdysozoa is unclear. Complete genome sequences might help to address these questions as genomic adaptations can be revealed and phylogenetic reconstructions can be based on new markers. Here, we present a first draft genome of a species from the family Milnesiidae, namely Milnesium tardigradum. We consistently place M. tardigradum and the two previously sequenced Hypsibiidae species, Hypsibius dujardini and Ramazzottius varieornatus, as sister group of the nematodes with the arthropods as outgroup. Based on this placement, we identify a massive gene loss thus far attributed to the nematodes which predates their split from the tardigrades. We provide a comprehensive catalog of protein domain expansions linked to stress response and show that previously identified tardigrade-unique proteins are erratically distributed across the genome of M. tardigradum. We further suggest alternative pathways to cope with high stress levels that are yet unexplored in tardigrades and further promote the phylum Tardigrada as a rich source of stress protection genes and mechanisms.

evolutionary biology

karyoploteR: An R/Bioconductor Package To Plot Customizable Linear Genomes Displaying Arbitrary Data

MotivationData visualization is a crucial tool for data exploration, analysis and interpretation. For the visualization of genomic data there lacks a tool to create customizable non-circular plots of whole genomes from any species.\n\nResultsWe have developed karyoploteR, an R/Bioconductor package to create linear chromosomal representations of any genome with genomic annotations and experimental data plotted along them. Plot creation process is inspired in R base graphics, with a main function creating karyoplots with no data and multiple additional functions, including custom functions written by the end-user, adding data and other graphical elements. This approach allows the creation of highly customizable plots from arbitrary data with complete freedom on data positioning and representation.\n\nAvailabilitykaryoploteR is released under Artistic-2.0 License. Source code and documentation are freely available through Bioconductor (http://www.bioconductor.org/packages/karyoploteR)\n\nContactbgel@igtp.cat

bioinformatics

Phylogenetic Conflict In Bears Identified By Automated Discovery Of Transposable Element Insertions In Low Coverage Genomes

Compared to sequence analyses, phylogenetic reconstruction from transposable elements (TEs) offers an additional perspective to study evolutionary processes. However, detecting phylogenetically informative TE insertions requires tedious experimental work, limiting the power of phylogenetic inference. Here, we analyzed the genomes of seven bear species using high throughput sequencing data to detect thousands of TE insertions. The newly developed pipeline for TE detection called TeddyPi (TE detection and discovery for Phylogenetic Inference) obtained 150,513 high-quality TE insertions in the genomes of ursine and tremarctine bears. By integrating different TE insertion callers and using a stringent filtering approach, the TeddyPi pipeline produced highly reliable TE insertion calls, which were confirmed by extensive in vitro validation experiments. Screening for single nucleotide substitutions in the flanking regions of the TEs show that these substitutions correlate with the phylogenetic signal from the TE insertions. Our phylogenomic analyses show that TEs are a major driver of genomic variation in bears and enabled phylogenetic reconstruction of a well-resolved species tree, even with strong signals for incomplete lineage sorting and introgression. The analyses show that the Asiatic black, sun and sloth bear form a monophyletic clade. TeddyPi is open source and can be adapted to various TE and structural variation callers. The pipeline makes it easy to confidently extract thousands of TE insertions even from low coverage genomes of non-model organisms, opening new possibilities for biologists to study phylogenies, evolutionary processes as well as rates and patterns of (retro-)transposition and structural variation.

evolutionary biology

K-mer Similarity, Networks Of Microbial Genomes And Taxonomic Rank

Alignment-free (AF) methods have recently been adopted to infer phylogenetic trees. However, the evolutionary relationships among microbes, impacted by common phenomena such as lateral genetic transfer and rearrangement, cannot be adequately captured in a strictly tree-like structure. Bacterial and archaeal genomes consist of highly conserved regions, e.g. ribosomal RNA genes (commonly used as phylogenetic markers), more-variable regions and extrachromosomal elements, i.e. plasmids (that contain genes critical under a selective condition e.g. antibiotic resistance). The impact of these elements on genome-scale inference of microbial phylogeny remains little known. Here, using an AF approach, we inferred phylogenomic networks of microbial life based on 2785 completely sequenced bacterial and archaeal genomes, and systematically assessed the impact of ribosomal RNA genes and plasmid sequences in this network. Our results indicate that k-mer similarity can correlate with taxonomic rank of microbes. Using a relational database approach, we linked the implicated k-mers to annotated genomic regions (thus functions), and defined core functions in specific phyletic groups and genera. We found that, in most phyla, highly conserved functions are often related to Amino acid metabolism and transport, and Energy production and conversion. Our findings indicate that AF phylogenomics can be used to infer reticulate relationships in a scalable manner and provide new perspective into microbial biology and evolution.

evolutionary biology

Integrated Computational Guide Design, Execution, And Analysis Of Arrayed And Pooled CRISPR Genome Editing Experiments

CRISPR genome editing experiments offer enormous potential for the evaluation of genomic loci using arrayed single guide RNAs (sgRNAs) or pooled sgRNA libraries. Numerous computational tools are available to help design sgRNAs with optimal on-target efficiency and minimal off-target potential. In addition, computational tools have been developed to analyze deep sequencing data resulting from genome editing experiments. However, these tools are typically developed in isolation and oftentimes not readily translatable into laboratory-based experiments. Here we present a protocol that describes in detail both the computational and benchtop implementation of an arrayed and/or pooled CRISPR genome editing experiment. This protocol provides instructions for sgRNA design with CRISPOR, experimental implementation, and analysis of the resulting high-throughput sequencing data with CRISPResso. This protocol allows for design and execution of arrayed and pooled CRISPR experiments in 4-5 weeks by non-experts as well as computational data analysis in 1-2 days that can be performed by both computational and non-computational biologists alike.

molecular biology

Enhancing Multiplex Genome Editing by Natural Transformation (MuGENT) via inactivation of ssDNA exonucleases

Recently, we described a method for multiplex genome editing by natural transformation (MuGENT). Mutant constructs for MuGENT require large arms of homology (>2000 bp) surrounding each genome edit, which necessitates laborious in vitro DNA splicing. In Vibrio cholerae, we uncover that this requirement is due to cytoplasmic ssDNA exonucleases, which inhibit natural transformation. In ssDNA exonuclease mutants, one arm of homology can be reduced to as little as 40 bp while still promoting integration of genome edits at rates of ~50% without selection in cis. Consequently, editing constructs are generated in a single PCR reaction where one homology arm is oligonucleotide encoded. To further enhance editing efficiencies, we also developed a strain for transient inactivation of the mismatch repair system. As a proof-of-concept, we used these advances to rapidly mutate 10 high-affinity binding sites for the nucleoid occlusion protein SlmA and generated a duodecuple mutant of 12 diguanylate cyclases in V. cholerae. Whole genome sequencing revealed little to no off-target mutations in these strains. Finally, we show that ssDNA exonucleases inhibit natural transformation in Acinetobacter baylyi. Thus, rational removal of ssDNA exonucleases may be broadly applicable for enhancing the efficacy and ease of MuGENT in diverse naturally transformable species.

microbiology

Redox-Dependent Condensation Of Mycobacterial Genome By WhiB4

Oxidative stress response in bacteria is generally mediated through coordination between the regulators of oxidant-remediation systems (e.g. OxyR, SoxR) and nucleoid condensation (e.g. Dps, Fis). However, these genetic factors are either absent or rendered nonfunctional in the human pathogen Mycobacterium tuberculosis (Mtb). Therefore, how Mtb organizes genome architecture and regulates gene expression to counterbalance oxidative imbalance during infection is not known. Here, we report that an intracellular redox-sensor, WhiB4, dynamically links genome condensation and oxidative stress response in Mtb. Disruption of WhiB4 affects the expression of genes involved in maintaining redox homeostasis, central carbon metabolism (CCM), respiration, cell wall biogenesis, DNA repair and protein quality control under oxidative stress. Notably, disulfide-linked oligomerization of WhiB4 in response to oxidative stress activates the proteins ability to condense DNA in vitro and in vivo. Further, overexpression of WhiB4 led to hypercondensation of nucleoids, redox imbalance and increased susceptibility to oxidative stress, whereas WhiB4 disruption reversed this effect. In accordance with the findings in vitro, ChIP-Seq data demonstrated non-specific binding of WhiB4 to GC-rich regions of the Mtb genome. Lastly, data indicate that WhiB4 deletion affected the expression of only a fraction of genes preferentially bound by the protein, suggesting its indirect effect on gene expression. We propose that WhiB4 is a novel redox-dependent nucleoid condensing protein that structurally couples Mtbs response to oxidative stress with genome organization and transcription.\n\nSignificance StatementMycobacterium tuberculosis (Mtb) needs to adapt in response to oxidative stress encountered inside human phagocytes. In other bacteria, condensation state of nucleoids modulates gene expression to coordinate oxidative stress response. However, this relation remains elusive in Mtb. We performed molecular dissection of a mechanism controlled by an intracellular redox sensor, WhiB4, in organizing both chromosomal structure and selective expression of adaptive traits to counter oxidative stress in Mtb. Using high-resolution sequencing, transcriptomics, imaging, and redox biosensor, we describe how WhiB4 modulates nucleoid condensation, global gene expression, and redox-homeostasis. WhiB4 over-expression hypercondensed nucleoids and perturbed redox homeostasis whereas WhiB4 disruption had an opposite effect. Our study discovered an empirical role for WhiB4 in integrating redox signals with nucleoid condensation in Mtb.

microbiology

Shared Genetic Architecture Of Asthma With Allergic Diseases: A Genome-wide Cross Trait Analysis Of 112,000 Individuals From UK Biobank

Clinical and epidemiological data suggest that asthma and allergic diseases are associated. And may share a common genetic etiology. We analyzed genome-wide single-nucleotide polymorphism (SNP) data for asthma and allergic diseases in 35,783 cases and 76,768 controls of European ancestry from the UK Biobank. Two publicly available independent genome wide association studies (GWAS) were used for replication. We have found a strong genome-wide genetic correlation between asthma and allergic diseases (rg = 0.75, P = 6.84x10-62). Cross trait analysis identified 38 genome-wide significant loci, including novel loci such as D2HGDH and GAL2ST2. Computational analysis showed that shared genetic loci are enriched in immune/inflammatory systems and tissues with epithelium cells. Our work identifies common genetic architectures shared between asthma and allergy and will help to advance our understanding of the molecular mechanisms underlying co-morbid asthma and allergic diseases.

genetics

Genome And Epigenome Engineering CRISPR Toolkit For Probing In Vivo cis-Regulatory Interactions In The Chicken Embryo

CRISPR-Cas9 genome engineering has revolutionised all aspects of biological research, with epigenome engineering transforming gene regulation studies. Here, we present a highly efficient toolkit enabling genome and epigenome engineering in the chicken embryo, and demonstrate its utility by probing gene regulatory interactions mediated by neural crest enhancers. First, we optimise efficient guide-RNA expression from novel chick U6-mini-vectors, provide a strategy for rapid somatic gene knockout and establish protocol for evaluation of mutational penetrance by targeted next generation sequencing. We show that CRISPR/Cas9-mediated disruption of transcription factors causes a reduction in their cognate enhancer-driven reporter activity. Next, we assess endogenous enhancer function using both enhancer deletion and nuclease-deficient Cas9 (dCas9) effector fusions to modulate enhancer chromatin landscape, thus providing the first report of epigenome engineering in a developing embryo. Finally, we use the synergistic activation mediator (SAM) system to activate an endogenous target promoter. The novel genome and epigenome engineering toolkit developed here enables manipulation of endogenous gene expression and enhancer activity in chicken embryos, facilitating high-resolution analysis of gene regulatory interactions in vivo.\n\nSummary StatementWe present an optimised toolkit for efficient genome and epigenome engineering using CRISPR in chicken embryos, with a particular focus on probing gene regulatory interactions during neural crest development.\n\nList of AbbreviationsGenome Engineering (GE), Epigenome Engineering (EGE), single guide RNA (sgRNA), Neural Crest (NC), Transcription Factor (TF), Next Generation Sequencing (NGS), somite stage (ss), Hamburger Hamilton (HH).

developmental biology