Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Axon-seq decodes the motor axon transcriptome and its modulation in response to ALS

Spinal motor axons traverse large distances to innervate target muscles, thus requiring local control of cellular events for proper functioning. To interrogate axon-specific processes we developed Axon-seq, a refined method incorporating microfluidics, RNA-seq and bioinformatic-QC. We show that the axonal transcriptome is distinct from somas and contains fewer genes. We identified 3,500-5,000 transcripts in mouse and human stem cell-derived spinal motor axons, most of which are required for oxidative energy production and ribogenesis. Axons contained transcription factor mRNAs, e.g. Ybx1, with implications for local functions. As motor axons degenerate in amyotrophic lateral sclerosis (ALS), we investigated their response to the SOD1G93A mutation, identifying 121 ALS-dysregulated transcripts. Several of these are implicated in axonal function, including Nrp1, Dbn1 and Nek1, a known ALS-causing gene. In conclusion, Axon-seq provides an improved method for RNA-seq of axons, increasing our understanding of peripheral axon biology and identifying novel therapeutic targets in motor neuron disease.

neuroscience

Loss-of-function in IRF2BPL is associated with neurological phenotypes

The Interferon Regulatory Factor 2 Binding Protein Like (IRF2BPL) gene encodes a member of the IRF2BP family of transcriptional regulators. Currently the biological function of this gene is obscure, and the gene has not been associated with a Mendelian disease. Here we describe seven individuals affected with neurological symptoms who carry damaging heterozygous variants in IRF2BPL. Five cases carrying nonsense variants in IRF2BPL resulting in a premature stop codon display severe neurodevelopmental regression, hypotonia, progressive ataxia, seizures, and a lack of coordination. Two additional individuals, both with missense variants, display global developmental delay and seizures and a relatively milder phenotype than those with nonsense alleles. The bioinformatics signature for IRF2BPL based on population genomics is consistent with a gene that is intolerant to variation. We show that the IRF2BPL ortholog in the fruit fly, called pits (protein interacting with Ttk69 and Sin3A), is broadly expressed including the nervous system. Complete loss of pits is lethal early in development, whereas partial knock-down with RNA interference in neurons leads to neurodegeneration, revealing requirement for this gene in proper neuronal function and maintenance. The nonsense variants in IRF2BPL identified in patients behave as severe loss-of-function alleles in this model organism, while ectopic expression of the missense variants leads to a range of phenotypes. Taken together, IRF2BPL and pits are required in the nervous system in humans and flies, and their loss leads to a range of neurological phenotypes in both species.

genetics

Prefrontal co-expression of schizophrenia risk genes is associated with treatment response in patients

Gene co-expression networks are relevant to functional and clinical translation of schizophrenia (SCZ) risk genes. We hypothesized that SCZ risk genes may converge into coexpression pathways which may be associated with gene regulation mechanisms and with response to treatment in patients with SCZ. We identified gene co-expression networks in two prefrontal cortex post-mortem RNA sequencing datasets (total N=688) and replicated them in four more datasets (total N=227). We identified and replicated (all p-values<.001) a single module enriched for SCZ risk loci (13 risk genes in 10 loci). In silico screening of potential regulators of the SCZ risk module via bioinformatic analyses identified two transcription factors and three miRNAs associated with the risk module. To translate post-mortem information into clinical phenotypes, we identified polymorphisms predicting co-expression and combined them to obtain an index approximating module co-expression (Polygenic Co-expression Index: PCI). The PCI-co-expression association was successfully replicated in two independent brain transcriptome datasets (total N=131; all p-values<.05). Finally, we tested the association between the PCI and short-term treatment response in two independent samples of patients with SCZ treated with olanzapine (total N=167). The PCI was associated with treatment response in the positive symptom domain in both clinical cohorts (all p-values<.05).\n\nIn summary, our findings in a large sample of human post-mortem prefrontal cortex show that coexpression of a set of genes enriched for schizophrenia risk genes is relevant to treatment response. This co-expression pathway may be co-regulated by transcription factors and miRNA associated with it.

neuroscience

Autonomous functionality of an upstream open reading frame in polycistronic mammalian mRNAs

Upstream open reading frames (uORFs) are established as cis-acting elements for eukaryotic translation of annotated ORFs (anORFs) located on the same mRNAs. Here, we identified a mammalian uORF with functions that are independent from anORF translation regulation. Bioinformatics screening using ribosome profiling data of human and mouse brains yielded 308 neurologically vital genes from which anORF and uORFs are polycistronically translated in both species. Among them, Arhgef9 contains a uORF named SPICA, which is highly conserved among vertebrates and stably translated only in specific brain regions of mice. Disruption of SPICA translation by ATG-to-TAG substitutions did not perturb translation or function of its anORF product, collybistin. SPICA-null mice displayed abnormal maternal reproductive performance and enhanced anxiety-like behavior, characteristic of ARHGEF9-associated neurological disorders. This study demonstrates that mammalian uORFs can be independent genetic units, revising the prevailing dogma of the monocistronic gene in mammals, and even eukaryotes.

molecular biology

Genome-centric metagenomics revealed the spatial distribution and the diverse metabolic functions of lignocellulose degrading uncultured bacteria

The mechanisms by which specific anaerobic microorganisms remain firmly attached to lignocellulosic material allowing them to efficiently decompose the organic matter are far to be elucidated. To circumvent this issue, the microbiomes collected from anaerobic digesters treating pig manure and meadow grass were fractionated to separate the planktonic microbes from those adhered to lignocellulosic substrate. Assembly of shotgun reads followed by binning process recovered 151 population genomes, 80 out of which were completely new and were not previously deposited in any database. Genome coverage allowed the identification of microbial spatial distribution into the engineered ecosystem. Moreover, a composite bioinformatic analysis using multiple databases for functional annotation revealed that uncultured members of Bacteroidetes and Firmicutes follow diverse metabolic strategies for polysaccharide degradation. The structure of cellulosome in Firmicutes can vary depending on the number and functional roles of carbohydrate-binding modules. On contrary, members of Bacteroidetes are able to adhere and degrade lignocellulose due to the presence of multiple carbohydrate-binding family 6 modules in beta-xylosidase and endoglucanase proteins or S-layer homology modules in unknown proteins. This study combines the concept of variability in spatial distribution with genome-centric metagenomics allowing a functional and taxonomical exploration of the biogas microbiome.\n\nImportanceThis work contributes new knowledge about lignocellulose degradation in engineered ecosystems. Specifically, the combination of the spatial distribution of uncultured microbes with genome-centric metagenomics provides novel insights into the metabolic properties of planktonic and firmly attached to plant biomass bacteria. Moreover, the knowledge obtained in this study enabled us to understand the diverse metabolic strategies for polysaccharide degradation in different species of Bacteroidetes and Clostridiales. Even though structural elements of cellulosome were restricted to Clostridiales, our study identified in Bacteroidetes a putative mechanism for biomass decomposition based on a gene cluster responsible for cellulose degradation, disaccharide cleavage to glucose and transport to cytoplasm.

microbiology

Laboratory Validation of a Clinical Metagenomic Sequencing Assay for Pathogen Detection in Cerebrospinal Fluid

Metagenomic next-generation sequencing (mNGS) for pan-pathogen detection has been successfully tested in proof-of-concept case studies in patients with acute illness of unknown etiology, but to date has been largely confined to research settings. Here we developed and validated an mNGS assay for diagnosis of infectious causes of meningitis and encephalitis from cerebrospinal fluid (CSF) in a licensed clinical laboratory. A clinical bioinformatics pipeline, SURPI+, was developed to rapidly analyze mNGS data, automatically report detected pathogens, and provide a graphical user interface for evaluating and interpreting results. We established quality metrics, threshold values, and limits of detection of between 0.16 - 313 genomic copies or colony forming units per milliliter for each representative organism type. Gross hemolysis and excess host nucleic acid reduced assay sensitivity; however, a spiked phage used as an internal control was a reliable indicator of sensitivity loss. Diagnostic test accuracy was evaluated by blinded mNGS testing of 95 patient samples, revealing 73% sensitivity and 99% specificity compared to original clinical test results, with 81% positive percent agreement and 99% negative percent agreement after discrepancy analysis. Subsequent mNGS challenge testing of 20 positive CSF samples prospectively collected from a cohort of pediatric patients hospitalized with meningitis, myelitis, and/or encephalitis showed 92% sensitivity and 96% specificity relative to conventional microbiological testing of CSF in identifying the causative pathogen. These results demonstrate the analytic performance of a laboratory-validated mNGS assay for pan-pathogen detection, to be used clinically for diagnosis of neurological infections from CSF.

genomics

AgriSeqDB: an online RNA-Seq database for functional studies in agriculturally relevant plant species

BackgroundThe genome-wide expression profile of genes in different tissues/cell types and developmental stages is a vital component of many functional genomic studies. Transcriptome data obtained by RNA-sequencing (RNA-Seq) is often deposited in public databases that are made available via data portals. Data visualization is one of the first steps in assessment and hypothesis generation. However, these databases do not typically include visualization tools and establishing one is not trivial for users who are not computational experts. This, as well as the various formats in which data is commonly deposited, makes the processes of data access, sharing and utility more difficult. Our goal was to provide a simple and user-friendly repository that meets these needs for datasets from major agricultural crops.\n\nDescriptionAgriSeqDB (https://expression.latrobe.edu.au/agriseqdb), is a database for viewing, analysing and interpreting developmental and tissue/cell-specific transcriptome data from several species, including major agricultural crops such as wheat, rice, maize, barley and tomato. The disparate manner in which public transcriptome data is often warehoused and the challenge of visualizing raw data are both major hurdles to data reuse. The popular eFP browser does an excellent job of presenting transcriptome data in an easily interpretable view, but previous implementation has been mostly on a case-by-case basis. Here we present an integrated visualisation database of transcriptome datasets from six species that did not previously have public-facing visualisations. We combine the eFP browser, for gene-by-gene investigation, with the Degust browser, which enables visualisation of all transcripts across multiple samples. The two visualisation interfaces launch from the same point, enabling users to easily switch between analysis modes. The tools allow users, even those without bioinformatics expertise, to mine into datasets and understand the behaviour of transcripts of interest across samples and time. We have also incorporated an additional graphic download option to simplify incorporation into presentations or publications.\n\nConclusionPowered by eFP and Degust browsers, AgriSeqDB is a quick and easy-to-use platform for data analysis and visualization in five crops and Arabidopsis. Furthermore, it provides a tool that makes it easy for researchers to share their datasets, promoting research collaborations and dataset reuse.

genomics

GoFish: A Streamlined Environmental DNA Presence/Absence Assay for Marine Vertebrates

Here we describe GoFish, a streamlined environmental DNA (eDNA) presence/absence assay. The assay amplifies a 12S segment with broad-range vertebrate primers, followed by nested PCR with M13-tailed, species-specific primers. Sanger sequencing confirms positives detected by gel electrophoresis. We first obtained 12S sequences from 77 fish specimens representing 36 northwestern Atlantic taxa not well documented in GenBank. Using the newly obtained and published 12S records, we designed GoFish assays for 11 bony fish species common in the lower Hudson River estuary and tested seasonal abundance and habitat preference at two sites. Additional assays detected nine cartilaginous fish species and a marine mammal, bottlenose dolphin, in southern New York Bight. GoFish sensitivity was equivalent to Illumina MiSeq metabarcoding. Unlike quantitative PCR (qPCR), GoFish does not require tissues of target and related species for assay development and a basic thermal cycler is sufficient. Unlike Illumina metabarcoding, indexing and batching samples are unnecessary and advanced bioinformatics expertise is not needed. The assay can be carried out from water collection to result in three days. The main limitations so far are species with shared target sequences and inconsistent amplification of rarer eDNAs. We think this approach will be a useful addition to current eDNA methods when analyzing presence/absence of known species, when turnaround time is important, and in educational settings.

ecology

Identification and quantification of Lyme pathogen strains by deep sequencing of outer surface protein C (ospC) amplicons

Mixed infection of a single tick or host by Lyme disease spirochetes is common and a unique challenge for diagnosis, treatment, and surveillance of Lyme disease. Here we describe a novel protocol for differentiating Lyme strains based on deep sequencing of the hypervariable outer-surface protein C locus (ospC). Improving upon the traditional DNA-DNA hybridization method, the next-generation sequencing-based protocol is high-throughput, quantitative, and able to detect new pathogen strains. We applied the method to over one hundred infected Ixodes scapularis ticks collected from New York State, USA in 2015 and 2016. Analysis of strain distributions within individual ticks suggests an overabundance of multiple infections by five or more strains, inhibitory interactions among co-infecting strains, and presence of a new strain closely related to Borreliella bissettiae. A supporting bioinformatics pipeline has been developed. With the newly designed pair of universal ospC primers targeting intergenic sequences conserved among all known Lyme pathogens, the protocol could be used for culture-free identification and quantification of Lyme pathogens in wildlife and clinical specimens across the globe.

microbiology

Applications for deep learning in ecology

A lot of hype has recently been generated around deep learning, a group of artificial intelligence approaches able to break accuracy records in pattern recognition. Over the course of just a few years, deep learning revolutionized several research fields such as bioinformatics or medicine. Yet such a surge of tools and knowledge is still in its infancy in ecology despite the ever-growing size and the complexity of ecological datasets. Here we performed a literature review of deep learning implementations in ecology to identify its benefits in most ecological disciplines, even in applied ecology, up to decision makers and conservationists alike. We also provide guidelines on useful resources and recommendations for ecologists to start adding deep learning to their toolkit. At a time when automatic monitoring of populations and ecosystems generates a vast amount of data that cannot be processed by humans anymore, deep learning could become a necessity in ecology.

ecology

Removal of alleles by genome editing -- RAGE against the deleterious load

BackgroundIn this paper, we simulate deleterious load in an animal breeding program, and compare the efficiency of genome editing and selection for decreasing load. Deleterious variants can be identified by bioinformatics screening methods that use sequence conservation and biological prior information about protein function. Once deleterious variants have been identified, how can they be used in breeding?\n\nResultsWe simulated a closed animal breeding population subject to both natural selection against deleterious load and artificial selection for a quantitative trait representing the breeding goal. Deleterious load was polygenic and due to either codominant or recessive variants. We compared strategies for removal of deleterious alleles by genome editing (RAGE) to selection against carriers. Each strategy varied in how animals and variants were prioritized for editing or selection.\n\nConclusionsGenome editing of deleterious alleles reduces deleterious load, but requires simultaneous editing of multiple deleterious variants in the same sire to be effective when deleterious variants are recessive. In the short term, selection against carriers is a possible alternative to genome editing when variants are recessive. The dominance of deleterious variants affects both the efficiency of genome editing and selection against carriers, and which variant prioritization strategy is the most efficient. Our results suggest that in the future, there is the potential to use RAGE against deleterious load in animal breeding.

genetics

Identification of pathogens in culture-negative infective endocarditis with metagenomic analysis

Pathogens identification is critical for the proper diagnosis and precise treatment of infective endocarditis. Although blood and valve cultures are the gold standard for IE pathogens detection, many cases are culture-negative, especially in patients who had received long-term antibiotic treatment, and precise diagnosis has therefore become a major challenge in the clinic. Metagenomic sequencing can provide both information on the pathogenic strain and the antibiotic susceptibility profile of patient samples without culturing, offering a powerful method to deal with culture-negative cases. In this work, we assessed the feasibility of a metagenomic approach to detect the causative pathogens in resected valves from IE patients.\n\nUsing our in-house developed bioinformatics pipeline, we analyzed the sequencing results generated from both next-generation sequencing and Oxford Nanopore Technologies MinION nanopore sequencing for the direct identification of pathogens from the resected valves of seven clinically culture-negative IE patients according to the modified Duke criteria. Moreover, we were able to simultaneously characterize respective antimicrobial resistance features. This provides clinicians with valuable information to diagnose and treat IE patients after valve replacement surgery.

microbiology

Ataxia Telangiectasia triggers deficits in Reelin pathway

Autosomal recessive Ataxia Telangiectasia (A-T) is characterized by radiosensitivity, immunodeficiency and cerebellar neurodegeneration. A-T is caused by inactivating mutations in the Ataxia-Telangiectasia-Mutated (ATM) gene, a serine-threonine protein kinase involved in DNA-damage response and excitatory neurotransmission. The selective vulnerability of cerebellar Purkinje neurons (PN) to A-T is not well understood.\n\nEmploying global proteomic profiling of cerebrospinal fluid from patients at ages around 15 years we detected reduced Calbindin, Reelin, Cerebellin-1, Cerebellin-3, Protocadherin Fat 2, Sempahorin 7A and increased Apolipoprotein -B, -H, -J peptides. Bioinformatic enrichment was observed for pathways of chemical response, locomotion, calcium binding and complement immunity. This seemed important, since secretion of Reelin from glutamatergic afferent axons is crucial for PN radial migration and spine homeostasis. Reelin expression is downregulated by irradiation and its deficiency is a known cause of ataxia. Validation efforts in 2-month-old Atm-/- mice before onset of motor deficits confirmed transcript reductions for Reelin receptors Apoer2/Vldlr with increases for their ligands Apoe/Apoh and cholesterol 24-hydroxylase Cyp46a1. Concomitant dysregulations were found for Vglut2/Sema7a as climbing fiber markers, glutamate receptors like Grin2b and calcium homeostasis factors (Atp2b2, Calb1, Itpr1), while factors involved in DNA damage, oxidative stress, neuroinflammation and cell adhesion were normal at this stage.\n\nThese findings show that deficient levels of Reelin signaling factors reflect the neurodegeneration in A-T in a sensitive and specific way. As an extracellular factor, Reelin may be accessible for neuroprotective interventions.

neuroscience

Chromomycin A2 potently inhibits glucose-stimulated insulin secretion from pancreatic beta cells.

Enhancers or inhibitors of insulin secretion could become therapeutics as well as lead to the identification of requisite {beta}-cell regulatory pathways and increase our understanding of pancreatic islet function. Toward this goal, we previously used an insulin-linked luciferase that is co-secreted with insulin in MIN6 {beta}-cells to perform a high-throughput natural product screen for chronic effects on glucose-stimulated insulin secretion. Using multiple phenotypic analyses, we identified that one of the top natural product hits, chromomycin A2 (CMA2), potently inhibited insulin secretion through at least three mechanisms: disruption of Wnt signaling, interfering with {beta}-cell gene expression, and suppression of triggering calcium (Ca2+) influx. Chronic treatment with CMA2 largely ablated glucose-stimulated insulin secretion even post-washout, but did not inhibit glucose-stimulated generation of ATP or Ca2+ influx. However, by using the KATP channel-opener diazoxide, we uncovered defects in depolarization-induced Ca2+ influx which may contribute to the suppressed secretory response. Glucose-responsive ERK1/2 and S6 phosphorylation were also disrupted by chronic CMA2 treatment. The FUSION bioinformatic database indicated that the phenotypic effects of CMA2 clustered with a number of Wnt/GSK3 pathway-related genes. Consistently, CMA2 decreased GSK3 phosphorylation and suppressed activation of a {beta}-catenin activity reporter. CMA2 and a related compound mithramycin are described to have DNA-interaction properties, possibly abrogating transcription factor binding to critical {beta}-cell gene promoters. We observed that CMA2, but not mithramycin, suppressed expression of PDX1 and UCN3. However, neither expression of INSI/II nor insulin content was affected by chronic CMA2. The mechanisms of CMA2-induced insulin secretion defects may involve components both proximal and distal to Ca2+ influx. Therefore, CMA2 is an example of a chemical that can simultaneously disrupt {beta}-cell function through both non-cytotoxic and cytotoxic mechanisms. Future applications of CMA2 and similar aureolic acid analogs for disease therapies should consider the potential impacts on pancreatic islet function.

cell biology

TuxNet: A simple interface to process RNA sequencing data and infer gene regulatory networks

Predicting gene regulatory networks (GRNs) from gene expression profiles has become a common approach for identifying important biological regulators. Despite the increase in the use of inference methods, existing computational approaches do not integrate RNA-sequencing data analysis, are often not automated, and are restricted to users with bioinformatics and programming backgrounds. To address these limitations, we have developed TuxNet, an integrated user-friendly platform, which, with just a few selections, allows to process raw RNA-sequencing data (using the Tuxedo pipeline) and infer GRNs from these processed data. TuxNet is implemented as a graphical user interface and, using expression data from any organism with an existing reference genome, can mine the regulations among genes either by applying a dynamic Bayesian network inference algorithm, GENIST, or a regression tree-based pipeline that uses spatiotemporal data, RTP-STAR. To illustrate the use of TuxNet while getting insight into the regulatory cascade downstream of the Arabidopsis root stem cell regulator PERIANTHIA (PAN), we obtained time course gene expression data of a PAN inducible line and inferred a GRN using GENIST. Using RTP-STAR, we then inferred the network of a PAN secondary downstream gene, ATHB13, for which we obtained wildtype and mutant expression profiles. Our case studies feature the versatility of TuxNet to infer networks using different types of gene expression data (i.e time course and steady-state data) as well as how inference networks are used to identify important regulators.\n\nSUMMARYTuxNet offers a simple interface for non-computational biologists to infer GRNs from raw RNA-seq data.

plant biology

Developing A Programmable, Self-Assembling Squash Leaf Curl China Virus (SLCCNV) Capsid Proteins Into \"Nano-Cargo\"-Like Architecture: A Next-Generation \"Nanotool\" For Biomedical Applications

A new era has begun in which pathogens have become useful scaffolds for nanotechnology applications. In this research/study, an attempt has been made to generate an empty cargo-like architecture from a high-profile plant pathogen of Squash leaf curl China virus (SLCCNV). In this approach, SLCCNV coat protein monomers are obtained efficiently by using a yeast Pichia pastoris expression system. Further, dialysis of purified SLCCNV-CP monomers against various pH strengthenened (5-10) disassembly and assembly buffers produced a self-assembled \"Nanocargo\"-like architecture, which also exhibited an ability to encapsulate the magnetic nanoparticles at in vitro. Bioinformatics tools were also utilized to predict the possible self-assembly kinetics and bioconjugation sites as well. The biocompatibility of \"SLCNNV-CP-Nanocargo\" particles was also evaluated by in vitro cancer cells, which eventually proved the particles to be versatile material for the next generation \"nanotool\" capable of housing various therapeutic or imaging agents.

systems biology

Connectivity analyses of bioenergetic changes in schizophrenia: Identification of novel treatments

We utilized a cell-level approach to examine glycolytic pathways in the DLPFC of subjects with schizophrenia (n=16) and control (n=16) subjects and found decreased mRNA expression of glycolytic enzymes in pyramidal neurons, but not astrocytes. To replicate these novel bioenergetic findings, we probed independent datasets for bioenergetic targets and found similar abnormalities. Next, we used a novel strategy to build a schizophrenia bioenergetic profile by a tailored application of the Library of Integrated Network-Based Cellular Signatures data portal (iLINCS) and investigated connected cellular pathways, kinases, and transcription factors using Enrichr. Finally, with the goal of identifying drugs capable of \"reversing\" the bioenergetic schizophrenia signature, we performed a connectivity analysis with iLINCS and identified peroxisome proliferator-activated receptor (PPAR) agonists as promising therapeutic targets. We administered a PPAR agonist to the GluN1 knockdown model of schizophrenia and found it improved long-term memory. Taken together, our findings suggest that tailored bioinformatics approaches, coupled with the LINCS library of transcriptional signatures of chemical and genetic perturbagens may be employed to identify novel treatment strategies for schizophrenia and related diseases.

neuroscience

Adomaviruses: an emerging virus family provides insights into DNA virus evolution

Adenoviruses, papillomaviruses, and polyomaviruses are collectively known as small DNA tumor viruses. Although it has long been recognized that small DNA tumor virus oncoproteins and capsid proteins show a variety of structural and functional similarities, it is unclear whether these similarities reflect descent from a common ancestor, convergent evolution, horizontal gene transfer among virus lineages, or acquisition of genes from host cells. Here, we report the discovery of a dozen new members of an emerging virus family, the Adomaviridae, that unite a papillomavirus/polyomavirus-like replicase gene with an adenovirus-like virion maturational protease. Adomaviruses were initially discovered in a lethal disease outbreak among endangered Japanese eels. New adomavirus genomes were found in additional commercially important fish species, such as tilapia, as well as in reptiles. The search for adomavirus sequences also revealed an additional candidate virus family, which we refer to as xenomaviruses, in mollusk datasets. Analysis of native adomavirus virions and expression of recombinant proteins showed that the virion structural proteins of adomaviruses are homologous to those of both adenoviruses and another emerging animal virus family called adintoviruses. The results pave the way toward development of vaccines against adomaviruses and suggest a framework that ties small DNA tumor viruses into a shared evolutionary history. Author SummaryIn contrast to cellular organisms, viruses do not encode any universally conserved genes. Even within a given family of viruses, the amino acid sequences encoded by homologous genes can diverge to the point of unrecognizability. Although members of an emerging virus family, the Adomaviridae, encode replicative DNA helicase proteins that are recognizably similar to those of polyomaviruses and papillomaviruses, the functions of other adomavirus genes have been difficult to identify. Using a combination of laboratory and bioinformatic approaches, we identify the adomavirus virion structural proteins. The results link adomavirus virion protein operons to those of other midsize non-enveloped DNA viruses, including adenoviruses and adintoviruses.

microbiology