Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

A filter-flow perspective of hematogenous metastasis offers a non-genetic paradigm for personalized cancer therapy

Translational RelevanceSince the discovery of circulating tumor cells (CTC), we have struggled for ways to use them to inform treatment. The only currently accepted method for this is a more is worse paradigm by which clinicians measure CTC burden before and after treatment to assess efficacy. Research efforts are currently focused almost entirely on genetic classification of these cells, which has yet to bear any fruit translationally. We suggest that we should shift the focus of our investigation to one driven by a physical sciences perspective. Specifically, by understanding the vascular system as a network of interconnected organs and capillary beds as filters that capture CTCs. By ascertaining the distribution of CTCs in this network for individual patients, information about the existence of subclinical metastatic disease, and therefore metastatic propensity, will come to light, and allow for better staging, prognostication and rational use of organ-directed therapy in the setting of oligometastatic disease.\n\nAbstractO_ST_ABSPurposeC_ST_ABSResearch into mechanisms of hematogenous metastasis has largely become genetic in focus, attempting to understand the molecular basis of seed-soil relationships. However, preceding this biological mechanism is the physical process of dissemination of circulating tumour cells (CTCs) in the circulatory network. We utilize a novel, network perspective of hematogenous metastasis and a large dataset on metastatic patterns to shed new light on this process.\n\nExperimental DesignThe metastatic efficiency index (MEI), previously suggested by Weiss, quantifies the process of hematogenous metastasis by taking the ratio of metastatic incidence for a given primary-target organ pair and the relative blood flow between the two sites. In this paper we extend the methodology by taking into account the reduction in CTC number that occurs in capillary beds and a novel network model of CTC flow.\n\nResultsBy applying this model to a dataset of metastatic incidence, we show that the MEI depends strongly on the assumptions of micrometastatic lesions in the lung and liver. Utilizing this framework we can represent different configurations of metastatic disease and offer a rational method for identifying patients with oligometastatic disease for inclusion in future trials.\n\nConclusionsWe show that our understanding of the dynamics of CTC flow is significantly lacking, and that this specifically precludes our ability to predict metastatic patterns in individual patients. Our formalism suggests an opportunity to go a step further in metastatic disease characterization by including the distribution of CTCs at staging, offering a rational method of trial design for oligometastatic disease.

Cancer Biology

Universality and predictability in molecular quantitative genetics

Molecular traits, such as gene expression levels or protein binding affinities, are increasingly accessible to quantitative measurement by modern high-throughput techniques. Such traits measure molecular functions and, from an evolutionary point of view, are important as targets of natural selection. We review recent developments in evolutionary theory and experiments that are expected to become building blocks of a quantitative genetics of molecular traits. We focus on universal evolutionary characteristics: these are largely independent of a traits genetic basis, which is often at least partially unknown. We show that universal measurements can be used to infer selection on a quantitative trait, which determines its evolutionary mode of conservation or adaptation. Furthermore, universality is closely linked to predictability of trait evolution across lineages. We argue that universal trait statistics extends over a range of cellular scales and opens new avenues of quantitative evolutionary systems biology.

Evolutionary Biology

High Genetic Diversity and Adaptive Potential of Two Simian Hemorrhagic Fever Viruses in a Wild Primate Population

Key biological properties such as high genetic diversity and high evolutionary rate enhance the potential of certain RNA viruses to adapt and emerge. Identifying viruses with these properties in their natural hosts could dramatically improve disease forecasting and surveillance. Recently, we discovered two novel members of the viral family Arteriviridae: simian hemorrhagic fever virus (SHFV)-krc1 and SHFV-krc2, infecting a single wild red colobus (Procolobus rufomitratus tephrosceles) in Kibale National Park, Uganda. Nearly nothing is known about the biological properties of SHFVs in nature, although the SHFV type strain, SHFV-LVR, has caused devastating outbreaks of viral hemorrhagic fever in captive macaques. Here we detected SHFV-krc1 and SHFV-krc2 in 40% and 47% of 60 wild red colobus tested, respectively. We found viral loads in excess of 106 - 107 RNA copies per milliliter of blood plasma for each of these viruses. SHFV-krc1 and SHFV-krc2 also showed high genetic diversity at both the inter- and intra-host levels. Analyses of synonymous and non-synonymous nucleotide diversity across viral genomes revealed patterns suggestive of positive selection in SHFV open reading frames (ORF) 5 (SHFV-krc2 only) and 7 (SHFV-krc1 and SHFV-krc2). Thus, these viruses share several important properties with some of the most rapidly evolving, emergent RNA viruses.

Microbiology

SINGLE NUCLEOTIDE POLYMORPHISMS SHED LIGHT ON CORRELATIONS BETWEEN ENVIRONMENTAL VARIABLES AND ADAPTIVE GENETIC DIVERGENCE AMONG POPULATIONS IN ONCORHYNCHUS KETA

Identifying the genetic and ecological basis of adaptation is of immense importance in evolutionary biology. In our study, we applied a panel of 58 biallelic single nucleotide polymorphisms (SNPs) for the economically and culturally important salmonid Oncorhynchus keta. Samples included 4164 individuals from 43 populations ranging from Coastal Western Alaska to southern British Colombia and northern Washington. Signatures of natural selection were detected by identifying seven outlier loci using two independent approaches: one based on outlier detection and another based on environmental correlations. Evidence of divergent selection at two candidate SNP loci, Oke_RFC2-168 and Oke_MARCKS-362, indicates significant environmental correlations, particularly with the number of frost-free days (NFFD). Important associations found between environmental variables and outlier loci indicate that those environmental variables could be the major driving forces of allele frequency divergence at the candidate loci. NFFD, in particular, may play an important adaptive role in shaping genetic variation in O. keta. Correlations between divergent selection and local environmental variables will help shed light on processes of natural selection and molecular adaptation to local environmental conditions.

Evolutionary Biology

Genetic drift suppresses bacterial conjugation in spatially structured populations

Conjugation is the primary mechanism of horizontal gene transfer that spreads antibiotic resistance among bacteria. Although conjugation normally occurs in surface-associated growth (e.g., biofilms), it has been traditionally studied in well-mixed liquid cultures lacking spatial structure, which is known to affect many evolutionary and ecological processes. Here we visualize spatial patterns of gene transfer mediated by F plasmid conjugation in a colony of Escherichia coli growing on solid agar, and we develop a quantitative understanding by spatial extension of traditional mass-action models. We found that spatial structure suppresses conjugation in surface-associated growth because strong genetic drift leads to spatial isolation of donor and recipient cells, restricting conjugation to rare boundaries between donor and recipient strains. These results suggest that ecological strategies, such as enforcement of spatial structure and enhancement of genetic drift, could complement molecular strategies in slowing the spread of antibiotic resistance genes.

Biophysics

Comparison of the theoretical and real-world evolutionary potential of a genetic circuit.

With the development of next-generation sequencing technologies, many large scale experimental efforts aim to map genotypic variability among individuals. This natural variability in populations fuels many fundamental biological processes, ranging from evolutionary adaptation and speciation to the spread of genetic diseases and drug resistance. An interesting and important component of this variability is present within the regulatory regions of genes. As these regions evolve, accumulated mutations lead to modulation of gene expression, which may have consequences for the phenotype. A simple model system where the link between genetic variability, gene regulation and function can be studied in detail is missing. In this article we develop a model to explore how the sequence of the wild-type lac promoter dictates the fold change in gene expression. The model combines single-base pair resolution maps of transcription factor and RNA polymerase binding energies with a comprehensive thermodynamic model of gene regulation. The model was validated by predicting and then measuring the variability of lac operon regulation in a collection of natural isolates. We then implement the model to analyze the sensitivity of the promoter sequence to the regulatory output, and predict the potential for regulation to evolve due to point mutations in the promoter region.

Evolutionary Biology

Mapping migration in a songbird using high-resolution genetic markers

Neotropical migratory birds are declining across the Western Hemisphere, but conservation efforts have been hampered by the inability to assess where migrants are most limited - the breeding grounds, migratory stopover sites, or wintering areas. A major challenge has been the lack of an efficient, reliable, and broadly applicable method for connecting populations across the annual cycle. Here we show how high-resolution genetic markers can be used to identify populations of a migratory bird, the Wilsons warbler (Cardellina pusilla), at fine enough spatial scales to facilitate assessing regional drivers of demographic trends. By screening 1626 samples using 96 single nucleotide polymorphisms (SNPs) selected from a large pool of candidates ([~]450,000), we identify novel region-specific migratory routes and timetables of migration along the Pacific Flyway. Our results illustrate that high-resolution genetic markers are more reliable, accurate, and amenable to high throughput screening than previously described tracking techniques, making them broadly applicable to large-scale monitoring and conservation of migratory organisms.

Ecology

MicroRNAs provide the first evidence of genetic link between diapause and aging in vertebrates

Diapause and aging are controlled by overlapping genetic mechanisms in C.elegans and these include microRNAs (miRNAs). Here, we investigated miRNA regulation in embryos of annual killifish that naturally undergo diapause to overcome desiccation of their habitats. We compared miRNA expression in diapausing and non-diapausing embryos in three independent lineages of killifish. We identified 13 miRNAs with similar regulation in all three lineages. One of these is miR-430, which is known as key regulator of early embryonic development in fish. We further tested whether this regulation overlaps with the aging-dependent regulation of miRNAs in one annual species: Nothobranchius furzeri. We found that miR-101a and miR-18a are regulated in the same direction during diapause and aging. These results provide the first evidence that overlapping genetic networks control diapause and aging in vertebrates and suggest that diapause mimics aging to some extent.

Genomics

Concurrent origins of the genetic code and the homochirality of life, and the origin and evolution of biodiversity. Part I: Observations and explanations

The post-genomic era has brought opportunities to bridge traditionally separate fields on early history of life. New methods promote a deeper understanding of the origin of biodiversity. Relative stabilities of base triplexes are able to regulate base substitutions in triplex DNAs. We constructed a roadmap based on such a regulation to explain concurrent origins of the genetic code and the homochirality of life. Based on the recruitment order of codons in the roadmap and the complete genome sequences, we reconstructed the three-domain tree of life. The Phanerozoic biodiversity curve has been reconstructed based on genomic, climatic and eustatic data; this result supports tectonic cause of mass extinctions. Our results indicate that chirality played a crucial role in the origin and evolution of life. Here is Part I of my two-part series paper; technical details are in Part II of this paper (see \"Concurrent origins of the genetic code and the homochirality of life, and the origin and evolution of biodiversity. Part II: Technical appendix\" on bioRxiv).

Evolutionary Biology

Efficient genotype compression and analysis of large genetic variation datasets

The economy of human genome sequencing has catalyzed ambitious efforts to interrogate the genomes of large cohorts in search of new insight into the genetic basis of disease. This manuscript introduces Genotype Query Tools (GQT) as a new indexing strategy and toolset that addresses an analytical bottleneck by enabling interactive analyses based on genotypes, phenotypes and sample relationships. Speed improvements are achieved by operating directly on a compressed genotype index without decompression. GQTs data compression ratios increase favorably with cohort size and relative analysis performance improves in kind. We demonstrate substantial performance improvements over state-of-the-art tools using datasets from the 1000 Genomes Project (46 fold), the Exome Aggregation Consortium (443 fold), and simulated datasets of up to 100,000 genomes (218 fold). Furthermore, we show that this indexing strategy facilitates population and statistical genetics measures such as principal component analysis and burden tests. Based on its computational efficiency and by complementing existing toolsets, GQT provides a flexible framework for current and future analyses of massive genome datasets.

Genomics

Characterizing and Prototyping Genetic Networks with Cell-Free Transcription-Translation Reactions

A central goal of synthetic biology is to engineer cellular behavior by engineering synthetic gene networks for a variety of biotechnology and medical applications. The process of engineering gene networks often involves an iterative design-build-test cycle, whereby the parts and connections that make up the network are built, characterized and varied until the desired network function is reached. Many advances have been made in the design and build portions of this cycle. However, the slow process of in vivo characterization of network function often limits the timescale of the testing step. Cell-free transcription-translation (TX-TL) systems offer a simple and fast alternative to performing these characterizations in cells. Here we provide an overview of a cell-free TX-TL system that utilizes the native Escherichia coli TX-TL machinery, thereby allowing a large repertoire of parts and networks to be characterized. As a way to demonstrate the utility of cell-free TX-TL, we illustrate the characterization of two genetic networks: an RNA transcriptional cascade and a protein regulated incoherent feed-forward loop. We also provide guidelines for designing TX-TL experiments to characterize new genetic networks. We end with a discussion of current and emerging applications of cell free systems.\n\nAbbreviations

Synthetic Biology

Are Genetic Interactions Influencing Gene Expression Evidence for Biological Epistasis or Statistical Artifacts?

The importance of epistasis - or statistical interactions between genetic variants - to the development of complex disease in humans has long been controversial. Genome-wide association studies of statistical interactions influencing human traits have recently become computationally feasible and have identified many putative interactions. However, several factors that are difficult to address confound the statistical models used to detect interactions and make it unclear whether statistical interactions are evidence for true molecular epistasis. In this study, we investigate whether there is evidence for epistasis regulating gene expression after accounting for technical, statistical, and biological confounding factors that affect interaction studies. We identified 1,119 (FDR=5%) interactions within cis-regulatory regions that regulate gene expression in human lymphoblastoid cell lines, a tightly controlled, largely genetically determined phenotype. Approximately half of these interactions replicated in an independent dataset (363 of 803 tested). We then performed an exhaustive analysis of both known and novel confounders, including ceiling/floor effects, missing genotype combinations, haplotype effects, single variants tagged through linkage disequilibrium, and population stratification. Every replicated interaction could be explained by at least one of these confounders, and replication in independent datasets did not protect against this issue. Assuming the confounding factors provide a more parsimonious explanation for each interaction, we find it unlikely that cis-regulatory interactions contribute strongly to human gene expression. As this calls into question the relevance of interactions for other human phenotypes, the analytic framework used here will be useful for protecting future studies of epistasis against confounding.

Genomics

Evolution of color phenotypes in two distantly related species of stick insect: different ecological regimes acting on similar genetic architectures

Recurrent (e.g. parallel or convergent) evolution is widely cited as evidence for natural selections central role in evolution but can also highlight constraints affecting evolution. Here we describe the evolution of green and melanistic color phenotypes in two species of stick insect: Timema podura and T. cristinae. We show that similar color phenotypes of these species (1) cluster in phenotypic space and (2) confer crypsis on different plant microhabitats. We then use genome-wide association mapping to determine the genetic architecture of color in T. podura, and compare this to previous results in T. cristinae. In both species, color is under simple genetic control, dominance relationships of melanistic and green alleles are the same, and SNPs associated with color phenotypes colocalize to the same genomic region. These results differ from those of typical parallel phenotypes because the form of selection acting on color differs between species: a balance of multiple sources of selection acting within host species maintains the color polymorphism in T. cristinae whereas T. podura color phenotypes are under divergent selection between hosts. Our results highlight how different adaptive landscapes can result in the evolution of similar phenotypic variation, and suggest the same genomic region is involved.

Evolutionary Biology

Combining niche modelling, land use-change, and genetic information to assess the conservation status of Pouteria splendens populations in Central Chile

BackgroundPouteria splendens (lucumo chileno) is an endemic shrub to the coastal areas of Central Chile classified as Endangered and Rare by the Chilean threatened species list, but as Lower Risk (LR) by IUCN. Based in historical records some authors have hypothesized that P. splendens originally formed a large metapopulation, but due to habitat loss and fragmentation these populations have been reduced to two main areas separated by 100 km, neither of both currently protected by the Chilean system of protected areas. Knowledge about this species is scarce and no studies have provided evidence to support the large metapopulation hypothesis. This gap of knowledge limits our availability to gauge the real urgency to conserve remaining P. splendens populations, which can generate tragic consequences in light of the increasing land-use change and climatic change that are facing these populations. In this study we combined niche modelling, land-use information, future climatic scenarios, and conservation genetics techniques, to test the hypothesis of a potential original large metapopulation, evaluate the role of land-use change in population decline, assess the threats this species may face in the future, and combine the generated information to re-assess its conservation status using the IUCN criteria.\n\nResultsOur results show that locations with P. splendens are fewer than described in the literature. Results from the niche modelling and genetic analyses support the hypothesis of an originally large metapopulation that was recently reduced and fragmented by anthropogenic land-use change. Future climate change could increase the range of suitable habitats for P. splendens towards inland areas; however the high level of fragmentation of these new areas is expected to preclude colonization processes.\n\nConclusionsBased on our results we recommend urgent actions towards the conservation of this species, including (1) re-evaluating its current IUCN conservation status and reclassifying it as Endangered (EN), and (2) take immediate actions to develop strategies that effectively protect the remaining populations.

Ecology

Diverse phenotypic and genetic responses to short-term selection in evolving Escherichia coli populations

Beneficial mutations fuel adaptation by altering phenotypes that enhance the fit of organisms to their environment. However, the phenotypic effects of mutations often depend on ecological context, making the distribution of effects across multiple environments essential to understanding the true nature of beneficial mutations. Studies that address both the genetic basis and ecological consequences of adaptive mutations remain rare. Here, we characterize the direct and pleiotropic fitness effects of a collection of 21 first-step beneficial mutants derived from naive and adapted genotypes used in a long-term experimental evolution of Escherichia coli. Whole-genome sequencing was used to identify most beneficial mutations. In contrast to previous studies, we find diverse fitness effects of mutations selected in a simple environment and few cases of genetic parallelism. The pleiotropic effects of these mutations were predominantly positive but some mutants were highly antagonistic in alternative environments. Further, the fitness effects of mutations derived from the adapted genotypes were dramatically reduced in nearly all environments. These findings suggest that many beneficial variants are accessible from a single point on the fitness landscape, and the fixation of alternative beneficial mutations may have dramatic consequences for niche breadth reduction via metabolic erosion.

Evolutionary Biology

Analysis of protein-coding genetic variation in 60,706 humans

Large-scale reference data sets of human genetic variation are critical for the medical and functional interpretation of DNA sequence changes. Here we describe the aggregation and analysis of high-quality exome (protein-coding region) sequence data for 60,706 individuals of diverse ethnicities generated as part of the Exome Aggregation Consortium (ExAC). The resulting catalogue of human genetic diversity contains an average of one variant every eight bases of the exome, and provides direct evidence for the presence of widespread mutational recurrence. We show that this catalogue can be used to calculate objective metrics of pathogenicity for sequence variants, and to identify genes subject to strong selection against various classes of mutation; we identify 3,230 genes with near-complete depletion of truncating variants, 72% of which have no currently established human disease phenotype. Finally, we demonstrate that these data can be used for the efficient filtering of candidate disease-causing variants, and for the discovery of human \"knockout\" variants in protein-coding genes.

Genomics

XMRF: An R package to Fit Markov Networks to High-Throughput Genetics Data

MotivationTechnological advances in medicine have led to a rapid proliferation of high-throughput \"omics\" data. Tools to mine this data and discover disrupted disease networks are needed as they hold the key to understanding complicated interactions between genes, mutations and aberrations, and epi-genetic markers.\n\nResultsWe developed an R software package, XMRF, that can be used to fit Markov Networks to various types of high-throughput genomics data. Encoding the models and estimation techniques of the recently proposed exponential family Markov Random Fields (Yang et al., 2012), our software can be used to learn genetic networks from RNA-sequencing data (counts via Poisson graphical models), mutation and copy number variation data (categorical via Ising models), and methylation data (continuous via Gaussian graphical models).\n\nAvailabilityXMRF is available from the CRAN Project and Github at: https://github.com/zhandong/XMRF

Bioinformatics

Expression weighted cell type enrichments reveal genetic and cellular nature of major brain disorders

The cell types that trigger the primary pathology in many brain diseases remain largely unknown. One route to understanding the primary pathological cell type for a particular disease is to identify the cells expressing susceptibility genes. Although this is straightforward for monogenic conditions where the causative mutation may alter expression of a cell type specific marker, methods are required for the common polygenic disorders. We developed the Expression Weighted Cell Type Enrichment (EWCE) method that uses single cell transcriptomes to generate the probability distribution associated with a gene list having an average level of expression within a cell type. Following validation, we applied EWCE to human genetic data from cases of epilepsy, Schizophrenia, Autism, Intellectual Disability, Alzheimers disease, Multiple Sclerosis and anxiety disorders. Genetic susceptibility primarily affected microglia in Alzheimers and Multiple Sclerosis; was shared between interneurons and pyramidal neurons in Autism and Schizophrenia; while intellectual disabilities and epilepsy were attributable to a range of cell-types, with the strongest enrichment in interneurons. We hypothesized that the primary cell type pathology could trigger secondary changes in other cell types and these could be detected by applying EWCE to transcriptome data from diseased tissue. In Autism, Schizophrenia and Alzheimers disease we find evidence of pathological changes in all of the major brain cell types. These findings give novel insight into the cellular origins and progression in common brain disorders. The methods can be applied to any tissue and disorder and have applications in validating mouse models.

Neuroscience