Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,423 records · Page 79Linked to original sources

The Genetic Chain Rule for Probabilistic Kinship Estimation

Accurate kinship predictions using DNA forensic samples has utility for investigative leads, remains identification, identifying relationships between individuals of interest, etc. High throughput sequencing (HTS) of STRs and single nucleotide polymorphisms (SNPs) is enabling the characterization of larger numbers of loci. Large panels of SNP loci have been proposed for improved mixture analysis of forensic samples. While multiple kinship prediction approaches have been established, we present an approach focusing on these large HTS SNP panels for predicting degree of kinship predictions. Formulas for first degree relatives can be multiplied (chained) together to model extended kinship relationships. Predictions are made using these formulations by calculating log likelihood ratios and selecting the maximum likelihood across the possible relationships. With a panel of 30,000 SNPs evaluated on an in silico dataset, this method can resolve parents from siblings and distinguish 1st, 2nd, and 3rd degree relatives from each other and unrelated individuals.

bioinformatics

Genetic Locus Modulating IOP

PurposeIntraocular pressure (IOP) is the primary risk factor for developing glaucoma. The present study examines genomic contribution to the normal regulation of IOP in the mouse.\n\nMethodsThe BXD recombinant inbred (RI) strain set was used to identify genomic loci modulating IOP. We measured the IOP from 532 eyes from 34 different strains. The IOP data will be subjected to conventional quantitative trait analysis using simple and composite interval mapping along with epistatic interactions to define genomic loci modulating normal IOP.\n\nResultsThe analysis defined one significant quantitative trait locus (QTL) on Chr.8 (100 to 106 Mb). The significant locus was further examined to define candidate genes that modulate normal IOP. There are only two good candidate genes within the 6 Mb over the peak, Cdh8 (Cadherin 8) and Cdh11 (Cadherin 11). Expression analysis on gene expression and immunohistochemistry indicate that Cdh11 is the best candidate for modulating the normal levels of IOP.\n\nConclusionsWe have examined the genomic regulation of IOP in the BXD RI strain set and found one significant QTL on Chr. 8. Within this QTL that are two potential candidates for modulating IOP with the most likely gene being Cdh11.

bioinformatics

The response to selection in Glycoside Hydrolase Family 13 structures: A comparative quantitative genetics approach

The Glycoside Hydrolase Family 13 (GH13) is both evolutionary diverse and relevant to many industrial applications. Its members perform the hydrolysis of starch into smaller carbohydrates. Members of the family have been bioengineered to improve catalytic function under industrial environments. We introduce a framework to analyze the response to selection of GH13 protein structures given some phylogenetic and simulated dynamic information. We found that the TIM-barrel is not selectable since it is under purifying selection. We also show a method to rank important residues with higher inferred response to selection. These residues can be altered to effect change in properties. In this work, we define fitness as inferred thermodynamic stability. We show that under the developed framework, residues 112Y, 122K, 124D, 125W, and 126P are good candidates to increase the stability of the truncated protein 4E2O. Overall, this paper demonstrate the feasibility of a framework for the analysis of protein structures for any other fitness landscape.

bioinformatics

Massively parallel single cell lineage tracing using CRISPR/Cas9 induced genetic scars

A key goal of developmental biology is to understand how a single cell transforms into a full-grown organism consisting of many different cell types. Single-cell RNA-sequencing (scRNA-seq) has become a widely-used method due to its ability to identify all cell types in a tissue or organ in a systematic manner 1-3. However, a major challenge is to organize the resulting taxonomy of cell types into lineage trees revealing the developmental origin of cells. Here, we present a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA-seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes, we reconstruct developmental lineage trees in zebrafish larvae and adult fish. In future analyses, LINNAEUS (LINeage tracing by Nuclease-Activated Editing of Ubiquitous Sequences) can be used as a systematic approach for identifying the lineage origin of novel cell types, or of known cell types under different conditions.

systems biology

The population genetics of spatial sorting

In most systems, dispersal occurs despite clear fitness costs to dispersing individuals. Theory posits that spatial heterogeneity in habitat quality pushes dispersal rates to evolve towards zero, while temporal heterogeneity in habitat quality favours non-zero dispersal rates. One aspect of dispersal evolution that has received a great deal of recent attention is a process known as spatial sorting, which has been referred to as a \"shy younger sibling\" of natural selection. More precisely, spatial sorting is the process whereby variation in dispersal ability is sorted along density clines and will, in nature, often be a transient phenomenon. Despite this transience, spatial sorting is likely a general mechanism behind non-zero dispersal in spatiotemporally varying environments. While generally transient, spatial sorting is persistent on invasion fronts, where its effect cannot be ignored, causing rapid evolution of traits related to dispersal. Spatial sorting is described in several elegant models, yet these models require a high level of mathematical sophistication and are not accessible to most evolutionary biologists or their students. Here, we frame spatial sorting in terms of the classic haploid and diploid models of natural selection. We show that, on an invasion front, spatial sorting can be conceptualized precisely as selection operating through space rather than (as with natural selection) time, and that genotypes can be viewed as having both spatial and temporal aspects of fitness. The resultant model is strikingly similar to classic models of natural selection. This similarity renders the model easy to understand (and to teach), but also suggests that many established theoretical results around natural selection could apply equally to spatial sorting.

evolutionary biology

The ancestral animal genetic toolkit revealed by diverse choanoflagellate transcriptomes

The changes in gene content that preceded the origin of animals can be reconstructed by comparison with their sister group, the choanoflagellates. However, only two choanoflagellate genomes are currently available, providing poor coverage of their diversity. We sequenced transcriptomes of 19 additional choanoflagellate species to produce a comprehensive reconstruction of the gains and losses that shaped the ancestral animal gene repertoire. We find roughly 1,700 gene families with origins on the animal stem lineage, of which only a core set of 36 are conserved across animals. We find more than 350 gene families that were previously thought to be animal-specific actually evolved before the animal-choanoflagellate divergence, including Notch and Delta, Toll-like receptors, and glycosaminoglycan hydrolases that regulate animal extracellular matrix (ECM). In the choanoflagellate Salpingoeca helianthica, we show that a glycosaminoglycan hydrolase modulates rosette colony size, suggesting a link between ECM regulation and morphogenesis in choanoflagellates and animals.\n\nData AvailabilityRaw sequencing reads: NCBI BioProject PRJNA419411 (19 choanoflagellate transcriptomes), PRJNA420352 (S. rosetta polyA selection test)\n\nTranscriptome assemblies, annotations, and gene families: https://dx.doi.org/10.6084/m9.figshare.5686984\n\nProtocols: https://dx.doi.org/10.17504/protocols.io.kwscxee

evolutionary biology

Efficiently controlling for case-control imbalance and sample relatedness in large-scale genetic association studies

In genome-wide association studies (GWAS) for thousands of phenotypes in large biobanks, most binary traits have substantially fewer cases than controls. Both of the widely used approaches, linear mixed model and the recently proposed logistic mixed model, perform poorly - producing large type I error rates - in the analysis of phenotypes with unbalanced case-control ratios. Here we propose a scalable and accurate generalized mixed model association test that uses the saddlepoint approximation (SPA) to calibrate the distribution of score test statistics. This method, SAIGE, provides accurate p-values even when case-control ratios are extremely unbalanced. It utilizes state-of-art optimization strategies to reduce computational time and memory cost of generalized mixed model. The computation cost linearly depends on sample size, and hence can be applicable to GWAS for thousands of phenotypes by large biobanks. Through the analysis of UK Biobank data of 408,961 white British European-ancestry samples for >1400 binary phenotypes, we show that SAIGE can efficiently analyze large sample data, controlling for unbalanced case-control ratios and sample relatedness.

genomics

Module analysis captures pancancer (epi)genetically deregulated cancer driver genes for smoking and antiviral response

The availability of increasing volumes of multi-omics profiles across many cancers promises to improve our understanding of the regulatory mechanisms underlying cancer. The main challenge is to integrate these multiple levels of omics profiles and especially to analyze them across many cancers. Here we present AMARETTO, an algorithm that addresses both challenges in three steps. First, AMARETTO identifies potential cancer driver genes through integration of copy number, DNA methylation and gene expression data. Then AMARETTO connects these driver genes with co-expressed target genes that they control, defined as regulatory modules. Thirdly, we connect AMARETTO modules identified from different cancer sites into a pancancer network to identify cancer driver genes. Here we applied AMARETTO in a pancancer study comprising eleven cancer sites and confirmed that AMARETTO captures hallmarks of cancer. We also demonstrated that AMARETTO enables the identification of novel pancancer driver genes. In particular, our analysis led to the identification of pancancer driver genes of smoking-induced cancers and antiviral interferon-modulated innate immune response.\n\nSoftware availabilityAMARETTO is available as an R package at https://bitbucket.org/gevaertlab/pancanceramaretto\n\nHighlightsO_LIWe present an algorithm for pancancer identification of cancer driver genes based on multiomics data fusion\nC_LIO_LIGPX2 is a novel driver gene in smoking induced cancers and validated using knockdown of GPX2 in the A549 cell line.\nC_LIO_LIOAS2 is a novel driver gene defining cancers with an antiviral signature supported by increased infiltration of tumor-associated macrophages.\nC_LI\n\nResearch in contextWe present an algorithm that combines multiple sources of molecular data to identify novel genes that are involved in cancer development. We applied this algorithm on multiple cancers in a combined fashion and identified a network of pancancer driver genes. We highlighted two genes in detail GPX2 and OAS2. We showed that GPX2 is an important cancer gene in smoking induced cancers, and validated our predictions using experimental data where GPX2 was inactivated in a lung cancer cell line. Similarly we showed that OAS2 is an important cancer driver gene in cancers that show an antiviral signature.

bioinformatics

The rust fungus Melampsora larici-populina expresses a conserved genetic program and distinct sets of secreted protein genes during infection of its two host plants, larch and poplar

Mechanims required for broad spectrum or specific host colonization of plant parasites are poorly understood. As a perfect illustration, heteroecious rust fungi require two alternate host plants to complete their life cycle. Melampsora larici-populina infects two taxonomically unrelated plants, larch on which sexual reproduction is achieved and poplar on which clonal multiplication occurs leading to severe epidemics in plantations. High-depth RNA sequencing was applied to three key developmental stages of M. larici-populina infection on larch: basidia, pycnia and aecia. Comparative transcriptomics of infection on poplar and larch hosts was performed using available expression data. Secreted protein was the only significantly over-represented category among differentially expressed M. larici-populina genes in basidia, pycnia and aecia compared together, highlighting their probable involvement in the infection process. Comparison of fungal transcriptomes in larch and poplar revealed a majority of rust genes commonly expressed on the two hosts and a fraction exhibiting a host-specific expression. More particularly, gene families encoding small secreted proteins presented striking expression profiles that highlight probable candidate effectors specialized on each host. Our results bring valuable new information about the biological cycle of rust fungi and identify genes that may contribute to host specificity.

microbiology

Genetic and cellular sensitivity of Caenorhabditis elegans to the chemotherapeutic agent cisplatin

Cisplatin and derivatives are commonly used as chemotherapeutic agents. Although the cytotoxic action of cisplatin on cancer cells is very efficient, clinical oncologists need to deal with two major difficulties: (i) the onset of resistance to the drug, and (ii) the cytotoxic effect in patients. Here we use Caenorhabditis elegans to investigate factors influencing the response to cisplatin in multicellular organisms. In this hermaphroditic model organism, we observed that sperm failure is a major cause in cisplatin-induced infertility. RNA-seq data indicate that cisplatin triggers a systemic stress response in which DAF-16/FOXO and SKN-1/Nrf2, two conserved transcription factors, are key regulators. We determined that inhibition of the DNA-damage induced apoptotic pathway does not confer cisplatin protection to the animal. However, mutants for the pro-apoptotic BH3-only gene ced-13 are sensitive to cisplatin, suggesting a protective role of the intrinsic apoptotic pathway. Finally, we demonstrate that our system can also be used to identify mutations providing resistance to cisplatin and therefore potential biomarkers of innate cisplatin-refractory patients. We show that mutants for the redox regulator trxr-1, ortholog of the mammalian Thioredoxin-Reductase-1 TrxR1, display cisplatin resistance and that such resistance relies on a single selenocysteine residue.

cancer biology

Genetic dissection of cyclic pyranopterin monophosphate biosynthesis in plant mitochondria

Mitochondria play a key role in the biosynthesis of two metal cofactors, iron-sulfur (FeS) clusters and molybdenum cofactor (Moco). The two pathways intersect at several points, but a scarcity of mutants has hindered studies to better understand these links. We screened a collection of sirtinol-resistant Arabidopsis thaliana mutants for lines with decreased activities of cytosolic FeS enzymes and Moco enzymes. We identified a new mutant allele of ATM3, encoding the ATP-binding cassette Transporter of the Mitochondria 3 (systematic name ABCB25), confirming the previously reported role of ATM3 in both FeS cluster and Moco biosynthesis. We also identified a mutant allele in CNX2, Cofactor of Nitrate reductase and Xanthine dehydrogenase 2, encoding GTP 3',8-cyclase, the first step in Moco biosynthesis which is localized in the mitochondria. A single nucleotide polymorphism in cnx2-2 leads to substitution of Arg88 with Gln in the N-terminal FeS cluster-binding motif. cnx2-2 plants are small and chlorotic, with severely decreased Moco enzyme activities, but they performed better than a cnx2-1 knockout mutant, which could only survive with ammonia as nitrogen source. Measurement of cyclic pyranopterin monophosphate (cPMP) levels by LC-MS/MS showed that this Moco intermediate was below the limit of detection in both cnx2-1 and cnx2-2, and accumulated more than 10-fold in seedlings mutated in the downstream gene CNX5. Interestingly, atm3-1 mutants had less cPMP than wild type, correlating with previous reports of a similar decrease in nitrate reductase activity. Taken together, our data functionally characterise CNX2 and suggest that ATM3 is indirectly required for cPMP synthesis.

biochemistry

Extensive sex differences at the initiation of genetic recombination

Homologous recombination in meiosis is initiated by programmed DNA double strand breaks (DSBs) and DSB repair as a crossover is essential to prevent chromosomal abnormalities in gametes. Sex differences in recombination have been previously observed by analyses of recombination end-products. To understand when and how sex differences are established, we built genome-wide maps of meiotic DSBs in both male and female mice. We found that most recombination initiates at sex-biased DSB hotspots. Local context, the choice of DSB targeting pathway and sex-specific patterns of DNA methylation give rise to these differences. Sex differences are not limited to the initiation stage, as the rate at which DSBs are repaired as crossovers appears to differ between the sexes in distal regions. This uneven repair patterning may be linked to the higher aneuploidy rate in females. Together, these data demonstrate that sex differences occur early in meiotic recombination.

genomics

Identification of Subsets of Genetic Alterations in KRAS-mutant Lung Cancer using Association Rule Mining

BackgroundLung cancer is the leading cause of all cancer death accounting for 1 out of 4 cancer-related death in both men and women. KRAS mutations occur in ~ 25% of patients with lung cancer, and the presence of these mutations is associated with poor prognosis. Efforts to directly target KRAS or associated downstream MAPK or the PI3K/AKT/mTOR pathways have seen little or no benefits. One probable reason for the lack of progress in targeting KRAS-mutant tumors is the co-occurrence of other cell survival pathways and mechanisms.\n\nMethod and resultsTo identify other potential cell survival pathways in subsets of KRAS-mutant tumors, I performed unsupervised machine learning on somatic mutations in metastatic lung cancer from 725 patient samples. I identified 67 other genes that were mutated in at least 10% of the samples with KRAS alterations. This gene list was enriched with genes involved in the MAPK, AKT and STAT3 pathways, cell-cell adhesion, DNA repair, chromatin remodeling, and the Wnt/beta-catenin pathway. I also identified 160 overlapping subsets of 3 or more genes that code for oncogenic or oncosuppressive proteins that were mutated in at least 10% of KRAS-mutant tumors.\n\nConclusionsIn this study, I identified genes that are co-mutated in KRAS-mutant lung cancer. I also identify subpopulations of KRAS-mutant lung cancer based on the set of genes that were also altered in the tumor samples. The design of research models that captures these subsets of KRAS-mutant tumors would enhance our understanding of the disease and facilitate personalized treatment for lung cancer patients with KRAS alterations.

cancer biology

Self-limiting population genetic control with sex-linked genome editors

In male heterogametic species the Y chromosome is transmitted solely from fathers to sons, and is selected for based only on its impacts on male fitness. This fact can be exploited to develop efficient pest control strategies that use Y-linked editors to disrupt the fitness of female descendants. In simple \"strategic\" population models we show that Y-linked editors can be substantially more efficient than other self-limiting strategies and, while not as efficient as gene drive approaches, are expected to have less impact on non-target populations with which there is some gene flow. Efficiency can be further augmented by simultaneously releasing an autosomal X-shredder construct, in either the same or different males. Y-linked editors may be attractive option to consider when efficient control of a species is desired in some locales but not others.

synthetic biology

epiTAD: a web application for visualizing high throughput chromosome conformation capture data in the context of genetic epidemiology

The increasing availability of public data resources coupled with advancements in genomic technology has created greater opportunities for researchers to examine the genome on a large and complex scale. To meet the need for integrative genome wide exploration, we present epiTAD. This web-based tool enables researchers to compare genomic structures and annotations across multiple databases and platforms in an interactive manner in order to facilitate in silico discovery. epiTAD can be accessed at https://apps.gerkelab.com/epiTAD/.

genomics

Spatial Capture-Recapture for Categorically Marked Populations with An Application to Genetic Capture-Recapture

Recently introduced unmarked spatial capture-recapture (SCR), spatial mark-resight (SMR), and 2-flank spatial partial identity models (SPIM) extend the domain of SCR to populations or observation systems that do not always allow for individual identity to be determined with certainty. For example, some species do not have natural marks that can reliably produce individual identities from photographs, and some methods of observation produce partial identity samples as is the case with remote cameras that sometimes produce single flank photographs. These models share the feature that they probabilistically resolve the uncertainty in individual identity using the spatial location where samples were collected. Spatial location is informative of individual identity in spatially structured populations with home range sizes smaller than the extent of the trapping array because a latent identity sample is more likely to have been produced by an individual living near the trap where it was recorded than an individual living further away from the trap. Further, the level of information about individual identity that a spatial location contains is determined by two key ecological concepts, population density and home range size. The number of individuals that could have produced a latent or partial identity sample increases as density and home range size increase because more individual home ranges will overlap any given trap. We show this uncertainty can be quantified using a metric describing the expected magnitude of uncertainty in individual identity for any given population density and home range size, the Identity Diversity Index (IDI). We then show that the performance of latent and partial identity SCR models varies as a function of this index and produces imprecise and biased estimates in many high IDI scenarios when data are sparse. We then extend the unmarked SCR model to incorporate partially identifying covariates which reduce the level of uncertainty in individual identity, increasing the reliability and precision of density estimates, and allowing reliable density estimation in scenarios with higher IDI values and with more sparse data. We illustrate the performance of this \"categorical SPIM\" via simulations and by applying it to a black bear data set using microsatellite loci as categorical covariates, where we reproduce the full data set estimates with only slightly less precision using fewer loci than necessary for confident individual identification. The categorical SPIM offers an alternative to using probability of identity criteria for classifying genotypes as unique, shifting the \"shadow effect\", where more than one individual in the population has the same genotype, from a source of bias to a source of uncertainty. We discuss the difficulties that real world data sets pose for latent identity SCR methods, most importantly, individual heterogeneity in detection function parameters, and argue that the addition of partial identity information reduces these concerns. We then discuss how the categorical SPIM can be applied to other wildlife sampling scenarios such as remote camera surveys, where natural or researcher-applied partial marks can be observed in photographs. Finally, we discuss how the categorical SPIM can be added to SMR, 2-flank SPIM, or other future latent identity SCR models.

ecology

Genetic recombination of poliovirus facilitates subversion of host barriers to infection

The contribution of RNA recombination to viral fitness and pathogenesis is poorly defined. Here, we isolate a recombination-deficient, poliovirus variant and find that, while recombination is detrimental to virus replication in tissue culture, recombination is important for pathogenesis in infected animals. Notably, recombination-defective virus exhibits severe attenuation following intravenous inoculation that is associated with a significant reduction in population size during intra-host spread. Because the impact of high mutational loads manifests most strongly at small population sizes, our data suggest that the repair of mutagenized genomes is an essential function of recombination and that this function may drive the long-term maintenance of recombination in viral species despite its associated fitness costs.\n\nSignificance StatementRNA recombination is a widespread but poorly understood feature of RNA virus replication. For poliovirus, recombination is involved in the emergence of neurovirulent circulating vaccine-derived poliovirus, which has hampered global poliovirus eradication efforts. This emergence illustrates the power of recombination to drive major adaptive change; however, it remains unclear if these adaptive events represent the primary role of recombination in virus survival. Here, we identify a viral mutant with a reduced rate of recombination and find that recombination also plays a central role in the spread of virus within animal hosts. These results highlight a novel approach for improving the safety of live attenuated vaccines and further our understanding of the role of recombination in virus pathogenesis and evolution.

microbiology

ACE: A Workbench using Evolutionary Genetic Algorithms for analysing association in TCGA Data

Modern methods in generating molecular data have dramatically scaled in recent years, allowing researchers to efficiently acquire large volumes of information. However, this has increased the challenge of recognising interesting patterns within the data. Atlas Correlation Explorer (ACE) is a user-friendly workbench for seeking associations between attributes in the cancer genome atlas (TCGA) database. It allows any combination of clinical and genomic data streams to be selected for searching, and highlights significant correlations within the chosen data. It is based on an evolutionary algorithm which is capable of producing results for very large searches in a short time.

bioinformatics