Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

Genetically Targeted Ratiometric and Activated pH Indicator Complexes (TRApHIC) for Receptor Trafficking

Fluorescent protein based pH sensors are a useful tool for measuring protein trafficking through pH changes associated with endo-and exocytosis. However, commonly used pH sensing probes are ubiquitously expressed with their protein of interest throughout the cell, hindering the ability to focus on specific trafficking pools of proteins. We developed a family of excitation-ratiometric, activatable pH responsive tandem dyes, consisting of a pH sensitive Cy3 donor linked to a fluorogenic malachite green acceptor. These cell-excluded dyes are targeted and activated upon binding to a genetically expressed fluorogen activating protein, and are suitable for selective labeling of surface proteins for analysis of endocytosis and recycling in live cells using both confocal and superresolution microscopy. Quantitative profiling of the endocytosis and recycling of tagged {beta}2-adrenergic receptor (B2AR) at a single vesicle level revealed differences among B2AR agonists, consistent with more detailed pharmacological profiling.

cell biology

Neither pulled nor pushed: Genetic drift and front wandering uncover a new class of reaction-diffusion waves

Short AbstractTraveling waves describe diverse natural phenomena from crystal growth in physics to range expansions in biology. Two classes of waves exist with very different properties: pulled and pushed. Pulled waves are driven by high growth rates at the expansion edge, where the number of organisms is small and fluctuations are large. In contrast, fluctuations are suppressed in pushed waves because the region of maximal growth is shifted towards the population bulk. Although it is commonly believed that expansions are either pulled or pushed, we found an intermediate class of waves with bulk-driven growth, but exceedingly large fluctuations. These waves are unusual because their properties are controlled by both the leading edge and the bulk of the front.\n\nLong AbstractEpidemics, flame propagation, and cardiac rhythms are classic examples of reaction-diffusion waves that describe a switch from one alternative state to another. Only two types of waves are known: pulled, driven by the leading edge, and pushed, driven by the bulk of the wave. Here, we report a distinct class of semi-pushed waves for which both the bulk and the leading edge contribute to the dynamics. These hybrid waves have the kinetics of pushed waves, but exhibit giant fluctuations similar to pulled waves. The transitions between pulled, semi-pushed, and fully-pushed waves occur at universal ratios of the wave velocity to the Fisher velocity. We derive these results in the context of a species invading a new habitat by examining front diffusion, rate of diversity loss, and fluctuation-induced corrections to the expansion velocity. All three quantities decrease as a power law of the population density with the same exponent. We analytically calculate this exponent taking into account the fluctuations in the shape of the wave front. For fully-pushed waves, the exponent is -1 consistent with the central limit theorem. In semi-pushed waves, however, the fluctuations average out much more slowly, and the exponent approaches 0 towards the transition to pulled waves. As a result, a rapid loss of genetic diversity and large fluctuations in the position of the front occur even for populations with cooperative growth and other forms of an Allee effect. The evolutionary outcome of spatial spreading in such populations could therefore be less predictable than previously thought.

ecology

Sensitive and specific post-call filtering of genetic variants in xenograft and primary tumors

MotivationTumor genome sequencing offers great promise for guiding research and therapy, but spurious variant calls can arise from multiple sources. Mouse contamination can generate many spurious calls when sequencing patient-derived xenografts (PDXs). Paralogous genome sequences can also generate spurious calls when sequencing any tumor. We developed a BLAST-based algorithm, MAPEX, to identify and filter out spurious calls from both these sources.\n\nResultsWhen calling variants from xenografts, MAPEX has similar sensitivity and specificity to more complex algorithms. When applied to any tumor, MAPEX also automatically flags calls that potentially arise from paralogous sequences. Our implementation, mapexr, runs quickly and easily on a desktocomputer. MAPEX is thus a useful addition to almost any pipeline for calling genetic variants in tumors.

bioinformatics

Recent African gene flow responsible for excess of old rare genetic variation in Great Britain

Population genomic studies can reveal the allele frequencies at millions of SNPs, with the numbers of observed low frequency SNPs increasing as more genomes are sequenced. Rare alleles tend to be younger than common alleles and are especially useful for studying demographic history, selection and heritability1,2. However, allele frequency can be a poor proxy for allele age, as genetic drift and natural selection can lead to alleles that are both rare and old. In order to allow joint assessments of allele frequency and allele age, a new estimator of allele age was developed that can be applied to variants of the lowest observed frequencies (singletons). By examining the geographic and age distribution of very rare variants in a large genomic sample from the UK3, we identify new evidence of gene flow from Africa into the ancestors of the modern UK population. A substantial proportion of variants with observed frequencies as low as 1.4 x 10-4 are orders of magnitude older than can be explained without African gene flow and are found at much higher frequencies within modern African populations. We estimate that African populations contributed approximately 1.2% of the UK gene pool and did so approximately 400 years ago. These findings are relevant both to our understanding of human history and to the nature of rare variation segregating within populations: a variant that is rare because it is a recent mutation in the direct ancestor of the population will have had a very different evolutionary history than an ancient one that has persisted at high frequencies in a diverged population and only recently arrived through migration.

genomics

A Yeast Global Genetic Screen Reveals that Metformin Induces an Iron Deficiency-Like State

We report here a simple and global strategy to map out gene functions and target pathways of drugs, toxins or other small molecules based on \"homomer dynamics\" Protein-fragment Complementation Assays (hdPCA). hdPCA measures changes in self-association (homomerization) of over 3,500 yeast proteins in yeast grown under different conditions. hdPCA complements genetic interaction measurements while eliminating confounding effects of gene ablation. We demonstrate that hdPCA accurately predicts the effects of two longevity and health-span-affecting drugs, immunosuppressant rapamycin and type II diabetes drug metformin, on cellular pathways. We also discovered an unsuspected global cellular response to metformin that resembles iron deficiency. This discovery opens a new avenue to investigate molecular mechanisms for the prevention or treatments of diabetes, cancers and other chronic diseases of aging.

systems biology

Filling Gaps in Bacterial Amino Acid Biosynthesis Pathways with High-throughput Genetics

For many bacteria with sequenced genomes, we do not understand how they synthesize some amino acids. This makes it challenging to reconstruct their metabolism, and has led to speculation that bacteria might be cross-feeding amino acids. We studied heterotrophic bacteria from 10 different genera that grow without added amino acids even though an automated tool predicts that the bacteria have gaps in their amino acid synthesis pathways. Across these bacteria, there were 11 gaps in their amino acid biosynthesis pathways that we could not fill using current knowledge. Using genome-wide mutant fitness data, we identified novel enzymes that fill 9 of the 11 gaps and hence explain the biosynthesis of methionine, threonine, serine, or histidine by bacteria from six genera. We also found that the sulfate-reducing bacterium Desulfovibrio vulgaris synthesizes homocysteine (which is a precursor to methionine) by using DUF39, NIL/ferredoxin, and COG2122 proteins, and that homoserine is not an intermediate in this pathway. Our results suggest that most free-living bacteria can likely make all 20 amino acids and illustrate how high-throughput genetics can uncover previously-unknown amino acid biosynthesis genes.

microbiology

Genetic load and mutational meltdown in cancer cell populations

ABSRACTLarge and non-recombining genomes are prone to accumulating deleterious mutations faster than natural selection can purge (Mullers ratchet). A possible consequence would then be the extinction of small populations. Relative to most single-cell organisms, cancer cells, with large and non-recombining genomes, could be particularly susceptible to such \"mutational meltdown\". Curiously, deleterious mutations in cancer cells are rarely noticed despite the strong signals in cancer genome sequences. Here, by monitoring single-cell clones from HeLa cell lines, we characterize deleterious mutations that retard cell proliferation. The main mutational events are copy number variations (CNVs), which happen at an extraordinarily high rate of 0.29 events per cell division. The average fitness reduction, estimated to be 18% per mutation, is also very high. HeLa cell populations therefore have very substantial genetic load and, at this level, natural population would likely experience mutational meltdown. We suspect that HeLa cell populations may avoid extinction only after the population size becomes large. Because CNVs are common in most cell lines and cancer tissues, the observations hint at cancer cells vulnerability, which could be exploited by therapeutic strategies.

evolutionary biology

Genetic Adaptation of C. elegans to Environment Changes I: Multigenerational Analysis of the Transcriptome

Abrupt environment changes can elicit an array of genetic effects. However, many of these effects can be overlooked by functional genomic studies conducted in static laboratory conditions. We studied the transcriptomic responses of Caenorhabditis elegans under single generation exposures to drastically different culturing conditions. In our experimental scheme, P0 worms were maintained on terrestrial environments (agar plates), F1 in aquatic cultures, and F2 back to terrestrial environments. The laboratory N2 strain and the wild isolate AB1 strain were utilized to examine how the genotype contributes to the transcriptome dynamics. Significant variations were found in the gene expressions between the \"domesticated\" laboratory strain and the wild isolate in the different environments. The results showed that 20% - 27% of the transcriptional responses to the environment changes were transmitted to the subsequent generation. In aquatic conditions, the domesticated strain showed differential gene expression particularly for the genes functioning in the reproductive system and the cuticle development. In accordance with the transcriptomic responses, phenotypic abnormalities were detected in the germline and cuticle of the domesticated strain. Further studies showed that distinct groups of genes are exclusively expressed under specific environmental conditions, and many of these genes previously lacked supporting biochemical evidence.

genomics

The common origin of symmetry and structure in genetic sequences

When exploring statistical properties of genetic sequences two main features stand out: the existence of non-random structures at various scales (e.g., long-range correlations) and the presence of symmetries (e.g., Chargaff parity rules). In the last decades, numerous studies investigated the origin and significance of each of these features separately. Here we show that both symmetry and structure have to be considered as the outcome of the same biological processes, whose cumulative effect can be quantitatively measured on extant genomes. We present a novel analysis (based on a minimal model) that not only explains and reproduces previous observations but also predicts the existence of a nested hierarchy of symmetries emerging at different structural scales. Our genome-wide analysis of H. Sapiens confirms the theoretical predictions.

genomics

Measuring Selection Across HIV Gag: Combining Physico-Chemistry and Population Genetics

We present physico-chemical based model grounded in population genetics. Our model predicts the stationary probability of observing an amino acid residue at a given site. Its predictions are based on the physico-chemical properties of the inferred optimal residue at that site and the sensitivity of the proteins functionality to deviation from the physico-chemical optimum at that site. We contextualize our physico-chemical model by comparing our model fit and parameters it to the more general, but less biologically meaningful entropy based metric: site sensitivity or 1/E. We show mathematically that our physico-chemical model is a more restricted form of the entropy model and how 1/E is proportional to the log-likelihood of a parameter-wise saturated model. Next, we fit both our physico-chemical and entropy models to sequences for subtype Cs Gag poly-protein in the LANL HIV database. Comparing our models site sensitivity parameters G' to 1/E we find they are highly correlated. We also compare the ability of G', 1/E, and other indirect measures of HIV fitness to empirical in vitro and in vivo measures. We find G' does a slightly better job predicting empirical fitness measures of in vivo viral escape time and in vitro spreading rates. While our predictive gain is modest, our model can be modified to test more complex or alternative biological hypotheses. More generally, because of its explicit biological formulation, our model can be easily extended to test for stabilizing vs. diversifying selection. We conjecture that our model could also be extended include epistasis in a more realistic manner than Ising models, while requiring many fewer parameters than Potts models.

evolutionary biology

Genetic analysis of de novo variants reveals sex differences in complex and isolated congenital diaphragmatic hernia and indicates MYRF as a candidate gene

Congenital diaphragmatic hernia (CDH) is one of the most common and lethal birth defects. Previous studies using exome sequencing support a significant contribution of coding de novo variants in complex CDH cases with additional anomalies and likely gene-disrupting (LGD) variants in isolated CDH cases. To further investigate the genetic architecture of CDH, we performed exome or genome sequencing in 283 proband-parent trios. Combined with data from previous studies, we analyzed a total of 357 trios, including 148 complex and 209 isolated cases. Complex and isolated cases both have a significant burden of deleterious de novo coding variants (1.7~fold, p= 1.2x10-5 for complex, 1.5~fold, p= 9.0x10-5 for isolated). Strikingly, in isolated CDH, almost all of the burden is carried by female cases (2.1~fold, p=0.004 for likely gene disrupting and 1.8~fold, p= 0.0008 for damaging missense variants); whereas in complex CDH, the burden is similar in females and males. Additionally, de novo LGD variants in complex cases are mostly enriched in genes highly expressed in developing diaphragm, but distributed in genes with a broad range of expression levels in isolated cases. Finally, we identified a new candidate risk gene MYRF (4 de novo variants, p-value=2x10-10), a transcription factor intolerant of mutations. Patients with MYRF mutations have additional anomalies including congenital heart disease and genitourinary defects, likely representing a novel syndrome.

genomics

Enhanced Astrocyte Responses are Driven by a Genetic Risk Allele Associated with Multiple Sclerosis

Epigenetic annotation studies of genetic risk variants for multiple sclerosis (MS) implicate dysfunctional lymphocytes in MS susceptibility; however, the role of central nervous system (CNS) cells remains unclear. We investigated the effect of the risk variant, rs7665090G, located near NFKB1, on astrocytes. We demonstrated that chromatin is accessible at the risk locus, a prerequisite for its impact on astroglial function. The risk variant was associated with increased NF-{kappa}B signaling and target gene expression driving lymphocyte recruitment in cultured human astrocytes and astrocytes within MS lesions, and with increased lesional lymphocytic infiltrates. In MS patients, the risk genotype was associated with increased lesion volumes on MRI. Thus, we established that the rs7665090G variant perturbs astrocyte function resulting in increased CNS access for peripheral immune cells. MS may thus result from variant-driven dysregulation of the peripheral immune system and the CNS, where perturbed CNS cell function aids in establishing local autoimmune inflammation.\n\nOne Sentence SummaryThe NF-{kappa}B relevant multiple sclerosis risk variant, rs7665090G, drives astrocyte responses that promote lesion formation.

neuroscience

Genotype fingerprints enable fast and private comparison of genetic testing results for research and direct-to-consumer applications

As genetic testing expands out of the research laboratory into medical practice as well as the direct-to-consumer market, the efficiency with which the resulting genotype data can be compared between individuals is of increasing importance.\n\nWe present a method for summarizing personal genotypes, yielding genotype fingerprints that can be derived from any single nucleotide polymorphism (SNP)-based assay and readily compared to estimate relatedness. The resulting fingerprints remain comparable as chip designs evolve to higher marker densities. We demonstrate that they support applications including distinguishing genotypes of closely related individuals by relationship type, distinguishing closely related individuals from individuals from the same background population, identification of individuals in known background populations, and de novo identification of subpopulations within a large cohort in a high-throughput manner.\n\nAn important feature of genotype fingerprints is that, while fingerprints do not preserve anonymity, they summarize individual marker data in a way that prevents phenotype prediction. Genotype fingerprints are therefore well-suited to public sharing for ancestry determination purposes, without revealing personal health risk status.

bioinformatics

Origins of the current outbreak of multidrug resistant malaria in Southeast Asia: a retrospective genetic study

BackgroundAntimalarial failure is rapidly spreading across parts of Southeast Asia where dihydroartemisinin-piperaquine (DHA-PPQ) is used as first line treatment. The first published reports came from western Cambodia in 2013. Here we analyse genetic changes in the Plasmodium falciparum population of western Cambodia in the six years prior to that.\n\nMethodsWe analysed genome sequence data on 1492 P. falciparum samples from Southeast Asia, including 464 collected in western Cambodia between 2007 and 2013. Different epidemiological origins of resistance were identified by haplotypic analysis of the kelch13 artemisinin resistance locus and the plasmepsin 2-3 piperaquine resistance locus.\n\nFindingsWe identified over 30 independent origins of artemisinin resistance, of which the O_SCPCAPKELC_SCPCAP1 lineage accounted for 91% of DHA-PPQ-resistant parasites. In 2008, O_SCPCAPKELC_SCPCAP1 combined with O_SCPCAPPLAC_SCPCAP1, the major lineage associated with piperaquine resistance. By 2012, the O_SCPCAPKELC_SCPCAP1/O_SCPCAPPLAC_SCPCAP1 co-lineage had reached over 60% frequency in western Cambodia and had spread to northern Cambodia.\n\nInterpretationThe O_SCPCAPKELC_SCPCAP1/O_SCPCAPPLAC_SCPCAP1 co-lineage emerged in the same year that DHA-PPQ became the first line antimalarial drug in western Cambodia and spread aggressively thereafter, displacing other artemisinin-resistant parasite lineages. These findings have significant implications for management of the global health risk associated with the current outbreak.\n\nFundingWellcome Trust, Bill & Melinda Gates Foundation, Medical Research Council, UK Department for International Development, and Intramural Research Program of the US National Institute of Allergy and Infectious Diseases, National Institutes of Health.

evolutionary biology

Identification and prioritisation of causal variants in human genetic disorders from exome or whole genome sequencing data

With genome sequencing entering the clinics as diagnostic tool to study genetic disorders, there is an increasing need for bioinformatics solutions that enable precise causal variant identification in a timely manner.\n\nBackgroundWorkflows for the identification of candidate disease-causing variants perform usually the following tasks: i) identification of variants; ii) filtering of variants to remove polymorphisms and technical artifacts; and iii) prioritization of the remaining variants to provide a small set of candidates for further analysis.\n\nMethodsHere, we present a pipeline designed to identify variants and prioritize the variants and genes from trio sequencing or pedigree-based sequencing data into different tiers.\n\nResultsWe show how this pipeline was applied in a study of patients with neurodevelopmental disorders of unknown cause, where it helped to identify the causal variants in more than 35% of the cases.\n\nConclusionsClassification and prioritization of variants into different tiers helps to select a small set of variants for downstream analysis.

bioinformatics

INFERNO - INFERring the molecular mechanisms of NOncoding genetic variants

The majority of variants identified by genome-wide association studies (GWAS) reside in the noncoding genome, where they affect regulatory elements including transcriptional enhancers. We propose INFERNO (INFERring the molecular mechanisms of NOncoding genetic variants), a novel method which integrates hundreds of diverse functional genomics data sources with GWAS summary statistics to identify putatively causal noncoding variants underlying association signals. INFERNO comprehensively infers the relevant tissue contexts, target genes, and downstream biological processes affected by causal variants. We apply INFERNO to schizophrenia GWAS data, recapitulating known schizophrenia-associated genes including CACNA1C and discovering novel signals related to transmembrane cellular processes.

bioinformatics

A genetically encoded Ca2+ indicator based on circularly permutated sea anemone red fluorescent protein

Genetically-encoded calcium ion (Ca2+) indicators (GECIs) are indispensable tools for measuring Ca2+ dynamics and neuronal activities in vitro and in vivo. Red fluorescent protein (RFP)-based GECIs enable multicolor visualization with blue or cyan-excitable fluorophores and combined use with blue or cyan-excitable optogenetic actuators. Here we report the development, structure, and validation of a new red fluorescent Ca2+ indicator, K-GECO1, based on a circularly permutated RFP derived from the sea anemone Entacmaea quadricolor. We characterized the performance of K-GECO1 in cultured HeLa cells, dissociated neurons, stem cell derived cardiomyocytes, organotypic brain slices, zebrafish spinal cord in vivo, and mouse brain in vivo.

biophysics

Comprehensive analysis of mobile genetic elements in the gut microbiome reveals a phylum-level niche-adaptive gene pool

Mobile genetic elements (MGEs) drive extensive horizontal transfer in the gut microbiome. This transfer could benefit human health by conferring new metabolic capabilities to commensal microbes, or it could threaten human health by spreading antibiotic resistance genes to pathogens. Despite their biological importance and medical relevance, MGEs from the gut microbiome have not been systematically characterized. Here, we present a comprehensive analysis of chromosomal MGEs in the gut microbiome using a method called Split Read Insertion Detection (SRID) that enables the identification of the exact mobilizable unit of MGEs. Leveraging the SRID method, we curated a database of 5600 putative MGEs encompassing seven MGE classes called ImmeDB (Intestinal microbiome mobile element database) (https://immedb.mit.edu/). We observed that many MGEs carry genes that confer an adaptive advantage to the gut environment including gene families involved in antibiotic resistance, bile salt detoxification, mucus degradation, capsular polysaccharide biosynthesis, polysaccharide utilization, and sporulation. We find that antibiotic resistance genes are more likely to be spread by conjugation via integrative conjugative elements or integrative mobilizable elements than transduction via prophages. Additionally, we observed that horizontal transfer of MGEs is extensive within phyla but rare across phyla. Taken together, our findings support a phylum level niche-adaptive gene pools in the gut microbiome. ImmeDB will be a valuable resource for future fundamental and translational studies on the gut microbiome and MGE communities.

bioinformatics