Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Demonstration of de novo chemotaxis in E. coli using a real-time, quantitative, and digital-like approach

Chemotaxis is the movement of an organism in response to an external chemical stimulus. This system enables bacteria to sense their immediate environment and adapt to changes in its chemical composition. Bacterial chemotaxis is mediated by chemoreceptors, membrane proteins that bind an effector and transduce the signal to the downstream proteins. From a synthetic biology perspective, the natural chemotactic repertoire is of little use since bacterial chemoreceptors have evolved to sense specific ligands that either benefit or harm the cell. Here we demonstrate that using a combined computational design approach together with a quantitative, real-time, and digital detection approach, we can rapidly design, manufacture, and characterize a synthetic chemoreceptor in E. coli for histamine (a ligand for which there are no known chemoreceptors). First, we employed a computational protocol that uses the Rosetta bioinformatics software together with high threshold filters to design mutational variants to the native Tar ligand binding domain that target histamine. Second, we tested different ligand-chemoreceptors pairs with a novel chemotaxis assay, based on optical reflectance interferometry of porous silicon (PSi) optical transducers, enabling label-free quantification of chemotaxis by monitoring real-time changes in the optical readout (expressed as the effective optical thickness, EOT). We found that different ligands can be characterized by an individual set of fingerprints in our assay. Namely, a binary, digital-like response in EOT change (i.e. positive or negative) that differentiates between attractants and repellants, the amplitude of change of EOT response, and the rate by which steady state in EOT change is reached. Using this assay, we were able to positively identify and characterize a single mutational chemoreceptor variant for histamine that mediated chemotaxis comparably to the natural Tar-aspartate system. Our results demonstrate the possibility of not only expanding the natural chemotaxis repertoire, but also provide a new quantitative assay by which to characterize the efficacy of the chemotactic response.

synthetic biology

Single molecule, full-length transcript sequencing provides insight into the extreme metabolism of ruby-throated hummingbird Archilochus colubris

Hummingbirds can support their high metabolic rates exclusively by oxidizing ingested sugars, which is unsurprising given their sugar-rich nectar diet and use of energetically expensive hovering flight. However, they cannot rely on dietary sugars as a fuel during fasting periods, such as during the night, at first light, or when undertaking long-distance migratory flights, and must instead rely exclusively on onboard lipids. This metabolic flexibility is remarkable both in that the birds can switch between exclusive use of each fuel type within minutes and in that de novo lipogenesis from dietary sugar precursors is the principle way in which fat stores are built, sometimes at exceptionally high rates, such as during the few days prior to a migratory flight. The hummingbird hepatopancreas is the principle location of de novo lipogenesis and likely plays a key role in fuel selection, fuel switching, and glucose homeostasis. Yet understanding how this tissue, and the whole organism, achieves and moderates high rates of energy turnover is hampered by a fundamental lack of information regarding how genes coding for relevant enzymes differ in their sequence, expression, and regulation in these unique animals. To address this knowledge gap, we generated a de novo transcriptome of the hummingbird liver using PacBio full-length cDNA sequencing (Iso-Seq), yielding a total of 8.6Gb of sequencing data, or 2.6M reads from 4 different size fractions. We analyzed data using the SMRTAnalysis v3.1 Iso-Seq pipeline, including classification of reads and clustering of isoforms (ICE) followed by error-correction (Arrow). With COGENT, we clustered different isoforms into gene families to generate de novo gene contigs. We performed orthology analysis to identify closely related sequences between our transcriptome and other avian and human gene sets. We also aligned our transcriptome against the Calypte anna genome where possible. Finally, we closely examined homology of critical lipid metabolic genes between our transcriptome data and avian and human genomes. We confirmed high levels of sequence divergence within hummingbird lipogenic enzymes, suggesting a high probability of adaptive divergent function in the hepatic lipogenic pathways. Our results have leveraged cutting-edge technology and a novel bioinformatics pipeline to provide a compelling first direct look at the transcriptome of this incredible organism.

genomics

mRNA And Long Non-Coding RNA Expression Profiles In Rats Reveal Inflammatory Features In Sepsis-Associated Encephalopathy

BackgroundSepsis-associated encephalopathy (SAE) is related to cognitive sequelae in patients in the intensive care unit (ICU) and can have serious impacts on quality of life after recovery. Although various pathogenic pathways are involved in SAE development, little is known concerning the global role of long non-coding RNAs (lncRNAs) in SAE.\n\nMethodsHerein, we employed transcriptome sequencing approaches to characterize the effects of lipopolysaccharide (LPS) on lncRNA expression patterns in brain tissue isolated from Sprague-Dawley (SD) rats with and without SAE. We performed high-throughput transcriptome sequencing after LPS was intraperitoneally injected and predicted targets and functions using bioinformatics tools. Subsequently, we explored the results in detail according to Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) analyses.\n\nResultsLncRNAs were differentially expressed in brain tissue after LPS treatment. After 6 h of LPS exposure, expression of 400 lncRNAs were significantly changed, including an increase in 316 lncRNAs and a decrease in 84 lncRNAs. In addition, 155 mRNAs were differentially expressed, with 84 up-regulated and 71 down-regulated. At 24 h post-treatment, expression of 117 lncRNAs and 57 mRNAs was consistently elevated, while expression of 79 lncRNAs and 21 mRNAs was decreased (change > 1.5-fold; p < 0.05). We demonstrated for the first time that differentially expressed lncRNAs were predicted to be enriched in a post-chaperonin tubulin folding pathway (GO : 007023), which is closely related to the key step in the tubulin folding process.\n\nInterestingly, the predicted pathway (KEGG 04360: axon guidance) was significantly changed under the same conditions. These results reveal that LPS might influence the construction and polarization of microtubules, which exert predominant roles in synaptogenesis and related biofunctions in the rodent central nervous system (CNS).\n\nConclusionsAn inventory of LPS-modulated expression profiles from the rodent CNS is an important step toward understanding the function of mRNAs, including lncRNAs, and suggests that microtubule malformation and dysfunction may be involved in SAE pathogenesis.

genomics

MTAG: Multi-Trait Analysis of GWAS

We introduce Multi-Trait Analysis of GWAS (MTAG), a method for joint analysis of summary statistics from GWASs of different traits, possibly from overlapping samples. We apply MTAG to summary statistics for depressive symptoms (Neff = 354,862), neuroticism (N = 168,105), and subjective well-being (N = 388,538). Compared to 32, 9, and 13 genome-wide significant loci in the single-trait GWASs (most of which are themselves novel), MTAG increases the number of loci to 64, 37, and 49, respectively. Moreover, association statistics from MTAG yield more informative bioinformatics analyses and increase variance explained by polygenic scores by approximately 25%, matching theoretical expectations.

genomics

Trichoderma reesei Complete Genome Sequence, Repeat-Induced Point Mutation And Partitioning Of CAZyme Gene Clusters

Trichoderma reesei (Ascomycota, Pezizomycotina) QM6a is a model fungus for a broad spectrum of physiological phenomena, including plant cell wall degradation, industrial production of enzymes, light responses, conidiation, sexual development, polyketide biosynthesis and plant-fungal interactions. The genomes of QM6a and its high-enzyme producing mutants have been sequenced by second-generation-sequencing methods and are publicly available from the Joint Genome Institute (JGI). While these genome sequences have offered useful information for genomic and transcriptomic studies, their limitations and especially their short read lengths make them poorly suited for some particular biological problems, including assembly, genome-wide determination of chromosome architecture and genetic modification or engineering. We integrated Pacific Biosciences and Illumina sequencing platforms for the highest-quality genome assembly yet achieved, revealing seven telomere-to-telomere chromosomes (34,922,528 bp; 10877 genes) with 1630 newly-predicted genes and >1.5 Mb of new sequences. Most new sequences are located on AT-rich blocks, including 7 centromeres, 14 subtelomeres and 2329 interspersed AT-rich blocks. The seven QM6a centromeres separately consist of 24 conserved repeats and 37 putative centromere-encoded genes. These findings open up a new perspective for future centromere and chromosome architecture studies. Next, we demonstrate that sexual crossing readily induced cytosine-to-thymine point mutations on both tandem and unlinked duplicated sequences. We also show by bioinformatic analysis that Trichoderma reesei has evolved a robust repeat-induced point mutation (RIP) system to accumulate AT-rich sequences, with longer AT-rich blocks having more RIP mutations. The widespread distribution of AT-rich blocks correlates genome-wide partitions with gene clusters, explaining why clustering of genes has been reported to not influence gene expression in Trichoderma reesei. Compartmentation of ancestral gene clusters by AT-rich blocks might promote flexibilities that are evolutionarily advantageous in this fungus soil habitats and other natural environments. Our analyses, together with the complete genome sequence, provide a better blueprint for biotechnological and industrial applications.

genomics

Predicting Novel Metabolic Pathways Through Subgraph Mining

The ability to predict pathways for biosynthesis of metabolites is very important in metabolic engineering. It is possible to mine the repertoire of biochemical transformations from reaction databases, and apply the knowledge to predict reactions to synthesize new molecules. However, this usually involves a careful understanding of the mechanism and the knowledge of the exact bonds being created and broken. There is clearly a need for a method to rapidly predict reactions for synthesizing new molecules, which relies only on the structures of the molecules, without demanding additional information such as thermodynamics or hand-curated information such as atom-atom mapping, which are often hard to obtain accurately.\n\nWe here describe a robust method based on subgraph mining, to predict a series of biochemical transformations, which can convert between two (even previously unseen) molecules. We first describe a reliable method based on subgraph edit distance to map reactants and products, using only their chemical structures. Having mapped reactants and products, we identify the reaction centre and its neighbourhood, the reaction signature, and store this in a reaction rule network. This novel representation enables us to rapidly predict pathways, even between previously unseen molecules. We also propose a heuristic that predominantly recovers natural biosynthetic pathways from amongst hundreds of possible alternatives, through a directed search of the reaction rule network, enabling us to provide a reliable ranking of the different pathways. Our approach scales well, even to databases with > 100,000 reactions. A Java-based implementation of our algorithms is available at https://github.com/RamanLab/ReactionMiner\n\nCCS CONCEPTS*Information systems [->]Data mining; *Applied computing [->]Bioinformatics;

systems biology

Uncovering The Repertoire Of Endogenous Flaviviral Elements In Aedes Mosquito Genomes

Endogenous viral elements derived from non-retroviral RNA viruses were described in various animal genomes. Whether they have a biological function such as host immune protection against related viruses is a field of intense study. Here, we investigated the repertoire of endogenous flaviviral elements (EFVEs) in Aedes mosquitoes, the vectors of arboviruses such as dengue and chikungunya viruses. Previous studies identified three EFVEs from Ae. albopictus and one from Ae. aegypti cell lines. However, in-depth characterization of EFVEs in wild-type mosquito populations and individuals in vivo has not been performed. We detected the full-length DNA sequence of the previously described EFVEs and their respective transcripts in several Ae. albopictus and Ae. aegypti populations from geographically distinct areas. However, EFVE-derived proteins were not detected by mass spectrometry. Using deep sequencing, we detected the production of piRNA-like small RNAs in antisense orientation, targeting the EFVEs and their flanking regions in vivo. The EFVEs were integrated in repetitive regions of the mosquito genomes, and their flanking sequences varied among mosquito populations from different geographical regions. We bioinformatically predicted several new EFVEs from a Vietnamese Ae. albopictus population and observed variation in the occurrence of those elements among mosquito populations. Phylogenetic analysis of an Ae. aegypti EFVE suggested that it integrated prior to the global expansion of the species and subsequently diverged among and within populations. Together, this study revealed substantial structural and nucleotide diversity of flaviviral integrations in Aedes genomes. Unraveling this diversity will help to elucidate the potential biological function of these EFVEs.\n\nImportanceEndogenous viral elements (EVEs) are whole or partial viral sequences integrated in host genomes. Interestingly, some EVEs have important functions for host fitness and antiviral defense. Because mosquitoes also have EVEs in their genomes, we decided to thoroughly characterized them to lay the foundation of the potential use of these EVEs to manipulate the mosquito antiviral response. Here, we focused on EVEs related to the Flavivirus genus, to which dengue and Zika viruses belong, in Aedes mosquito individuals from geographically distinct areas. We showed the existence in vivo of flaviviral EVEs previously identified in mosquito cell lines and we detected new ones. We showed that EVEs have evolved differently in each mosquito population. They produced transcripts and small RNAs, but not proteins, suggesting a function at the RNA level. Our study uncovers the diverse repertoire of flaviviral EVEs in Aedes mosquito populations and suggests a role in the host antiviral system.

microbiology

A Novel Ultra High-Throughput 16S rRNA Amplicon Sequencing Library Preparation Method On The Illumina HiSeq Platform

BackgroundAdvances in sequencing technologies and bioinformatics have made the analysis of microbial communities almost routine. Nonetheless, the need remains to improve on the techniques used for gathering such data, including increasing throughput while lowering cost, and benchmarking the techniques so that potential sources of bias can be better characterized.\n\nResultsWe present a triple-index amplicon sequencing strategy that uses a two-stage PCR protocol. The strategy was extensively benchmarked through analysis of a mock community in order to assess biases introduced by sample indexing, number of PCR cycles, and template concentration. We further evaluated the method through re-sequencing of a standardized environmental sample. Finally, we evaluated our protocol on a set of fecal samples from a small cohort of healthy adults, demonstrating good performance in a realistic experimental setting. Between-sample variation was mainly related to batch effects, such as DNA extraction, while sample indexing was also a significant source of bias. PCR cycle number strongly influenced chimera formation and affected relative abundance estimates of species with high GC content. Libraries were sequenced using the Illumina HiSeq and MiSeq platforms to demonstrate that this protocol is highly scalable to sequence thousands of samples at a very low cost.\n\nConclusionsHere, we provide the most comprehensive study of performance and bias inherent to a 16S rRNA gene amplicon sequencing method to date. Triple-indexing greatly reduces the number of long custom DNA oligos required for library preparation, while the inclusion of variable length heterogeneity spacers minimizes the need for PhiX spike-in. This design results in a significant cost reduction of highly multiplexed amplicon sequencing. The biases we characterize highlight the need for highly standardized protocols. Reassuringly, we find that the biological signal is a far stronger structuring factor than the various sources of bias.

microbiology

Thermodynamic Projection Of The Antibody Interaction Network: The Binding Fountain Energy Landscape

The complexity of the adaptive immune system in humans is comparable to that of the central nervous system in terms of cell numbers, cellular diversity and the network of interactions between components. While the application of molecular biological methods and bioinformatics has brought about an ever deepening and sharpening static description of the molecular and cellular components of the system, a unifying theoretical understanding of the laws governing the dynamics of the system is still lacking.\n\nWe have recently developed a quantitative model for the description of antibody homeostasis as defined by the dimensions of antigen concentration, antigen-antibody interaction affinity and antibody concentration. In this paper we develop the concept of a novel thermodynamic representation of multiple molecular interactions in a system, the fountain energy landscape of binding. We show that the hypersurface of the binding fountain corresponds to the antibody-antigen interaction network projected onto an energy landscape defined by conformational entropy and free energy of binding. We demonstrate that thymus independent and thymus dependent antibody responses show distinct patterns of changes in the energy landscape. Overall, the binding fountain energy landscape concept allows a systems biological, thermodynamic perception of the functioning of the clonal humoral immune system.

biophysics

Causal Analyses, Statistical Efficiency And Phenotypic Precision Through Recall-By-Genotype Study Design

Genome-wide association studies have been useful in identifying common genetic variants related to a variety of complex traits and diseases; however, they are often limited in their ability to inform about underlying biology. Whilst bioinformatics analyses, studies of cells, animal models and applied genetic epidemiology have provided some understanding of genetic associations or causal pathways, there is a need for new genetic studies that elucidate causal relationships and mechanisms in a cost-effective, precise and statistically efficient fashion. We discuss the motivation for and the characteristics of the Recall-by-Genotype (RbG) study design, an approach that enables genotype-directed deep-phenotyping and improvement in drawing causal inferences. Specifically, we present RbG designs using single and multiple variants and discuss the inferential properties, analytical approaches and applications of both. We consider the efficiency of the RbG approach, the likely value of RbG studies for the causal investigation of disease aetiology and the practicalities of incorporating genotypic data into population studies in the context of the RbG study design. Finally, we provide a catalogue of the UK-based resources for such studies, an online tool to aid the design of new RbG studies and discuss future developments of this approach.

genetics

Abnormal Social Behaviors And Dysfunction Of Autism-Related Genes Associated With Daily Agonistic Interactions In Mice

BackgroundThe ability of people to communicate with each other is a necessary component of social behavior and the normal development of individuals who live in a community. An apparent decline in sociability may be the result of a negative social environment or the development of affective and neurological disorders, including autistic spectrum disorders. The behavior of these humans may be characterized by the deterioration of socialization, low communication, and repetitive and restricted behaviors. This study aimed to analyze changes in the social behaviors of male mice induced by daily agonistic interactions and investigate the involvement of genes, related with autistic spectrum disorders in the process of the impairment of social behaviors.\n\nMethodsAbnormal social behavior is induced by repeated experiences of aggression accompanied by wins (winners) or chronic social defeats (losers) in daily agonistic interactions in male mice. The collected brain regions (the midbrain raphe nuclei, ventral tegmental area, striatum, hippocampus, and hypothalamus) were sequenced at JSC Genoanalytica (http://genoanalytica.ru/, Moscow, Russia). The Cufflinks program was used to estimate the gene expression levels. Bioinformatic methods were used for the analysis of differentially expressed genes in male mice.\n\nResultsThe losers exhibited an avoidance of social contacts toward unfamiliar conspecific, immobility and low communication on neutral territory. The winners demonstrated aggression and hyperactivity in this condition. The exploratory activity (rearing) and approaching behavior time towards the partner were decreased, and the number of episodes of repetitive self-grooming behavior was increased in both social groups. These symptoms were similar to the symptoms observed in animal models of autistic spectrum disorders. In an analysis of the RNA-Seq database of the whole transcriptome in the brain regions of the winners and losers, we identified changes in the expression of the following genes, which are associated with autism in humans: Tph2, Maoa, Slc6a4, Htr7,Gabrb3, Nrxn1, Nrxn2, Nlgn1, Nlgn2, Nlgn3, Shank2, Shank3, Fmr1, Ube3a, Pten, Cntn3, Foxp2, Oxtr, Reln, Cadps2, Pcdh10, Ctnnd2, En2, Arx, Auts2, Mecp2, and Ptchd1.Common and specific changes in the expression of these genes in different brain regions were identified in the winners and losers.\n\nConclusionsThis research demonstrates for the first time that abnormalities in social behaviors that develop under a negative social environment in adults may be associated with alterations in expression of genes, related with autism in the brain.

neuroscience

Effects of the glucocorticoid drug prednisone on urinary proteome and candidate biomarkers

Urine is a good source of biomarkers for clinical proteomics studies. However, one challenge in the use of urine biomarkers is that outside factors can affect the urine proteome. Prednisone is a commonly prescribed glucocorticoid used to treat various diseases in the clinic. To evaluate the possible impact of glucocorticoid drugs on the urine proteome, specifically disease biomarkers, this study investigated the effects of prednisone on the rat urine proteome. Urine samples were collected from control rats and prednisone-treated rats after drug administration. The urinary proteome was analyzed using liquid chromatography-tandem mass spectrometry (LC-MS/MS), and proteins were identified using label-free proteome quantification. Differentially expressed proteins and their human orthologs were analyzed with bioinformatics methods. A total of 523 urinary proteins were identified in rat urine. Using label-free quantification, 27 urinary proteins showed expression changes after prednisone treatment. A total of 16 proteins and/or their human orthologs have been previously annotated as disease biomarkers. After functional analysis, we found that the pharmacological effects of prednisone were reflected in the urine proteome. Thus, urinary proteomics has the potential to be a powerful drug efficacy monitoring tool in the clinic. Meanwhile, alteration of the urine proteome due to prednisone treatment should be considered in future disease biomarker studies.

molecular biology

Comparative Genomics Shows That Viral Integrations Are Abundant And Express piRNAs In The Arboviral Vectors Aedes aegypti And Aedes albopictus

BackgroundArthropod-borne viruses (arboviruses) transmitted by mosquito vectors cause many important emerging or resurging infectious diseases in humans including dengue, chikungunya and Zika. Understanding the co-evolutionary processes among viruses and vectors is essential for the development of novel transmission-blocking strategies. Arboviruses form episomal viral DNA fragments upon infection of mosquito cells and adults. Additionally, sequences from insect-specific viruses and arboviruses have been found integrated into mosquito genomes.\n\nResultsWe used a bioinformatic approach to analyze the presence, abundance, distribution, and transcriptional activity of integrations from 425 non-retroviral viruses, including 133 arboviruses, across the presently available 22 mosquito genome sequences. Large differences in abundance and types of viral integrations were observed in mosquito species from the same region. Viral integrations are unexpectedly abundant in the arboviral vector species Aedes aegypti and Ae. albopictus, but are [~]10-fold less abundant in all other mosquitoes analysed. Additionally, viral integrations are enriched in piRNA clusters of both the Ae. aegypti and Ae. albopictus genomes and, accordingly, they express piRNAs, but not siRNAs.\n\nConclusionsDifferences in number of viral integrations in the genomes of mosquito species from the same geographic area support the conclusion that integrations of viral sequences is not dependent on viral exposure, but that lineage-specific interactions exits. Viral integrations are abundant in Ae. aegypti and Ae. albopictus, and represent a thus far unappreciated component of their genomes. Additionally, the genome locations of viral integrations and their production of piRNAs indicate a functional link between viral integrations and the piRNA pathway. These results greatly expand the breadth and complexity of small RNA-mediated regulation and suggest a role for viral integrations in antiviral defense in these two mosquito species.

genomics

Single-Virion Sequencing Of Lamivudine Treated HBV Populations Reveal Population Evolution Dynamics And Demographic History

Viral populations are complex, dynamic, and fast evolving. The evolution of groups of closely related viruses in a competitive environment is termed quasispecies. To fully understand the role that quasispecies play in viral evolution, characterizing the trajectories of viral genotypes in an evolving population is the key. In particular, long-range haplotype information for thousands of individual viruses is critical; yet generating this information is non-trivial. Popular deep sequencing methods generate relatively short reads that do not preserve linkage information, while third generation sequencing methods have higher error rates that make detection of low frequency mutations a bioinformatics challenge. Here we applied BAsE-Seq, an Illumina-based single-virion sequencing technology, to eight samples from four chronic hepatitis B (CHB) patients - once before antiviral treatment and once after viral rebound due to resistance. We obtained 248-8,796 single-virion sequences per sample, which allowed us to find evidence for both hard and soft selective sweeps. We were also able to reconstruct population demographic history that was independently verified by clinically collected data. We further verified four of the samples independently on PacBio and Illumina sequencers. Overall, we showed that single-virion sequencing yields insight into viral evolution and population dynamics in an efficient and high throughput manner. We believe that single-virion sequencing is widely applicable to the study of viral evolution in the context of drug resistance, differentiating between soft or hard selective sweeps, and the reconstruction of intra-host viral population demographic history.

evolutionary biology

Mapping And Phasing Of Structural Variation In Patient Genomes Using Nanopore Sequencing

Structural genomic variants form a common type of genetic alteration underlying human genetic disease and phenotypic variation. Despite major improvements in genome sequencing technology and data analysis, the detection of structural variants still poses challenges, particularly when variants are of high complexity. Emerging long-read single-molecule sequencing technologies provide new opportunities for detection of structural variants. Here, we demonstrate sequencing of the genomes of two patients with congenital abnormalities using the ONT MinION at 11x and 16x mean coverage, respectively. We developed a bioinformatic pipeline - NanoSV - to efficiently map genomic structural variants (SVs) from the long-read data. We demonstrate that the nanopore data are superior to corresponding short-read data with regard to detection of de novo rearrangements originating from complex chromothripsis events in the patients. Additionally, genome-wide surveillance of SVs, revealed 3,253 (33%) novel variants that were missed in short-read data of the same sample, the majority of which are duplications < 200bp in size. Long sequencing reads enabled efficient phasing of genetic variations, allowing the construction of genome-wide maps of phased SVs and SNVs. We employed read-based phasing to show that all de novo chromothripsis breakpoints occurred on paternal chromosomes and we resolved the long-range structure of the chromothripsis. This work demonstrates the value of long-read sequencing for screening whole genomes of patients for complex structural variants.

genomics

TelNet - a database for human and yeast genes involved in telomere maintenance

The ends of linear chromosomes, the telomeres, comprise repetitive DNA sequences that are protected by the shelterin protein complex. Cancer cells need to extend these telomere repeats for their unlimited proliferation, either by reactivating the reverse transcriptase telomerase or by using the alternative lengthening of telomeres (ALT) pathway. The different telomere maintenance (TM) mechanisms appear to involve hundreds of proteins but their telomere repeat length related activities are only partly understood. Currently, a database that integrates information on TM relevant genes is missing. To provide a reference for studies that dissect TM features, we here introduce the TelNet database at http://www.cancertelsys.org/telnet/. It offers a comprehensive compilation of more than 2,000 human and over 1,100 yeast genes linked to telomere maintenance. These genes were annotated in terms of TM mechanism, associated specific functions and orthologous genes, a TM significance score and information from peer-reviewed literature. This TM information can be retrieved via different search and view modes and evaluated for a set of genes on a statistics page. With these features TelNet can be integrated into the annotation of genes identified from bioinformatics analysis pipelines to determine possible connections with TM networks as illustrated by an exemplary application. We anticipate that TelNet will be a helpful resource for researchers that study TM processes.

cancer biology

Reassessment Of Lesion-Associated Gene And Variant Pathogenicity In Focal Human Epilepsies

PurposeIncreasing availability of surgically resected brain tissue from Focal Cortical Dysplasia and low-grade epilepsy-associated tumor patients fostered large-scale genetic examination. However, assessment of germline and somatic variant pathogenicity remains difficult.\n\nMethodsHere, we critically reevaluated the pathogenicity for all neuropathology-associated variants reported to date in the PubMed and ClinVar databases, including 12 disease-related genes and 88 neuropathology-associated missense variants. We (1) assessed evolutionary gene constraint using the pLI and missense z scores, (2) applied guidelines by the American College of Medical Genetics and Genomics (ACMG), and (3) predicted pathogenicity by using PolyPhen-2, CADD, and GERP.\n\nResultsConstraint analysis classified only seven out of 12 genes to be likely disease-associated, while 35 (40%) of those 88 variants were classified as being variants of unknown significance (VUS) and 53 (60%) as being likely pathogenic (LPII). Pathogenicity prediction yielded discrimination between neuropathology-associated variants (LPII and VUS) and rare variant scores obtained from individuals present in the Genome Aggregation Database (gnomAD).\n\nConclusionWe conclude that several VUS are likely disease-associated and will be reclassified by future molecular evidence. In summary, interpretation of lesion-associated gene variants remains complex while the application of current ACMG guidelines including bioinformatic pathogenicity prediction will help improving interpretation and prediction.

genetics

Deep Sequencing: Intra-Terrestrial Metagenomics Illustrates The Potential Of Off-Grid Nanopore DNA Sequencing

Genetic and genomic analysis of nucleic acids from environmental samples has helped transform our perception of the Earths subsurface as a major reservoir of microbial novelty. Many of the microbial taxa living in the subsurface are under-represented in culture-dependent investigations. In this regard, metagenomic analyses of subsurface environments exemplify both the utility of metagenomics and its power to explore microbial life in some of the most extreme and inaccessible environments on Earth. Hitherto, the transfer of microbial samples to home laboratories for DNA sequencing and bioinformatics is the standard operating procedure for exploring microbial diversity. This approach incurs logistical challenges and delays the characterization of microbial biodiversity. For selected applications, increased portability and agility in metagenomic analysis is therefore desirable. Here, we describe the implementation of sample extraction, metagenomic library preparation, nanopore DNA sequencing and taxonomic classification using a portable, battery-powered, suite of off-the-shelf tools (the \"MetageNomad\") to sequence ochreous sediment microbiota while within the South Wales Coalfield. While our analyses were frustrated by short read lengths and a limited yield of DNA, within the assignable reads, Proteobacterial (-, {beta}-, {gamma}-Proteobacteria) taxa dominated, followed by members of Actinobacteria, Firmicutes and Bacteroidetes, all of which have previously been identified in coals. Further to this, the fungal genus Candida was detected, as well as a methanogenic archaeal taxon. To the best of our knowledge, this application of the MetageNomad represents an initial effort to conduct metagenomics within the subsurface, and stimulates further developments to take metagenomics off the beaten track.

microbiology