Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Computing structure-based lipid accessibility of membrane proteins with mp_lipid_acc in RosettaMP

BackgroundMembrane proteins are vastly underrepresented in structural databases, which has led to a lack of computational tools and the corresponding inappropriate use of tools designed for soluble proteins. For membrane proteins, lipid accessibility is an essential property. Even though programs are available for sequence-based prediction of lipid accessibility and structure-based identification of solvent-accessible surface area, the latter does not distinguish between water accessible and lipid accessible residues in membrane proteins.\n\nResultsHere we present mp_lipid_acc, the first method to identify lipid accessible residues from the protein structure, implemented in the RosettaMP framework and available as a webserver. Our method uses protein structures transformed in membrane coordinates, for instance from PDBTM or OPM databases, and a defined membrane thickness to classify lipid accessibility of residues. mp_lipid_acc is applicable to both -helical and {beta}-barrel membrane proteins of diverse architectures with or without water-filled pores and uses a concave hull algorithm for classification. We further provide a manually curated benchmark dataset, on which our method achieves prediction accuracies of 90%.\n\nConclusionWe present a novel tool to classify lipid accessibility from the protein structure, which is applicable to proteins of diverse architectures and achieves prediction accuracies of 90% on a manually curated database. mp_lipid_acc is part of the Rosetta software suite, available at www.rosettacommons.org. The webserver is available at http://rosie.graylab.jhu.edu/mp_lipid_acc/submit and the benchmark dataset is available at http://tinyurl.com/mp-lipid-acc-dataset.\n\nSupplementary informationSupplementary information is available at BMC Bioinformatics.

biophysics

Genome sequencing and analysis of the first spontaneous Nanosilver resistant bacterium Proteus mirabilis strain SCDR1

BackgroundP. mirabilis is a common uropathogenic bacterium that can cause major complications in patients with long-standing indwelling catheters or patients with urinary tract anomalies. In addition, P. mirabilis is a common cause of chronic osteomyelitis in Diabetic foot ulcer (DFU) patients. We isolated P. mirabilis SCDR1 from a Diabetic ulcer patient. We examined P. mirabilis SCDR1 levels of resistance against Nano-silver colloids, the commercial Nano-silver and silver containing bandages and commonly used antibiotics. We utilized next generation sequencing techniques (NGS), bioinformatics, phylogenetic analysis and pathogenomics in the characterization of the infectious pathogen.\n\nResultsP. mirabilis SCDR1 is a multi-drug resistant isolate that also showed high levels of resistance against Nano-silver colloids, Nano-silver chitosan composite and the commercially available Nano-silver and silver bandages. The P. mirabilis -SCDR1 genome size is 3,815,621 bp. with G+C content of 38.44%. P. mirabilis-SCDR1 genome contains a total of 3,533 genes, 3,414 coding DNA sequence genes, 11, 10, 18 rRNAs (5S, 16S, and 23S), and 76 tRNAs. Our isolate contains all the required pathogenicity and virulence factors to establish a successful infection. P. mirabilis SCDR1 isolate is a potential virulent pathogen that despite its original isolation site, wound, it can establish kidney infection and its associated complications. P. mirabilis SCDR1 contains several mechanisms for antibiotics and metals resistance including, biofilm formation, swarming mobility, efflux systems, and enzymatic detoxification.\n\nConclusionP. mirabilis SCDR1 is the first reported spontaneous Nanosilver resistant bacterial strain. P. mirabilis SCDR1 possesses several mechanisms that may lead to the observed Nanosilver resistance.

microbiology

Splice Expression Variation Analysis (SEVA) for Differential Gene Isoform Usage in Cancer

MotivationCurrent bioinformatics methods to detect changes in gene isoform usage in distinct phenotypes compare the relative expected isoform usage in phenotypes. These statistics model differences in isoform usage in normal tissues, which have stable regulation of gene splicing. Pathological conditions, such as cancer, can have broken regulation of splicing that increases the heterogeneity of the expression of splice variants. Inferring events with such differential heterogeneity in gene isoform usage requires new statistical approaches.\n\nResultsWe introduce Splice Expression Variability Analysis (SEVA) to model increased heterogeneity of splice variant usage between conditions (e.g., tumor and normal samples). SEVA uses a rank-based multivariate statistic that compares the variability of junction expression profiles within one condition to the variability within another. Simulated data show that SEVA is unique in modeling heterogeneity of gene isoform usage, and benchmark SEVAs performance against EBSeq, DiffSplice, and rMATS that model differential isoform usage instead of heterogeneity. We confirm the accuracy of SEVAin identifying known splice variants in head and neck cancer and perform cross-study validation of novel splice variants. A novel comparison of splice variant heterogeneity between subtypes of head and neck cancer demonstrated unanticipated similarity between the heterogeneity of gene isoform usage in HPV-positive and HPV-negative subtypes and anticipated increased heterogeneity among HPV-negative samples with mutations in genes that regulate the splice variant machinery.\n\nConclusionThese results show that SEVA accurately models differential heterogeneity of gene isoform usage from RNA-seq data.\n\nAvailabilitySEVA is implemented in the R/Bioconductor package GSReg.\n\nContactbahman@jhu.edu, favorov@sensi.org, ejfertig@jhmi.edu

genomics

Metagenomic Characteristics of Bacterial Response to Petroleum Hydrocarbon Contamination in Diverse Environments as Revealed by Functional Taxonomic Strategies

Microbial remediation of oil polluted habitats remains one of the foremost methods for restoration of petroleum hydrocarbon contaminated environments. The development of effective bioremediation strategies however, require an extensive understanding of the resident microbiome of these habitats. Recent developments such as high-throughput sequencing has greatly facilitated the advancement of microbial ecological studies in oil polluted habitats. However, effective interpretation of biological characteristics from these large datasets remains a considerable challenge. In this study, we have implemented recently developed bioinformatic tools for analyzing 65 publicly available 16S rRNA datasets from 12 diverse hydrocarbon polluted habitats to decipher metagenomic characteristics of bacterial communities of the same. We have comprehensively described phylogenetic and functional compositions of these habitats and additionally inferred a multitude of metagenomic features including 255 taxa and 414 functional modules which can be used as biomarkers for effective distinction between the 12 oil polluted sites. We have identified essential metabolic signatures and also showed that significantly over-represented taxa often contribute to either or both, hydrocarbon degradation and additional important functions. Our findings reveal significant differences between hydrocarbon contaminated sites and establishes the importance of endemic factors in addition to petroleum hydrocarbons as driving factors for sculpting hydrocarbon contaminated bacteriomes.

microbiology

Dual RNA sequencing (dRNA-Seq) of bacteria and their host cells

Bacterial pathogens subvert host cells by manipulating cellular pathways for survival and replication; in turn, host cells respond to the invading pathogen through cascading changes in gene expression. Deciphering these complex temporal and spatial dynamics to identify novel bacterial virulence factors or host response pathways is crucial for improved diagnostics and therapeutics. Dual RNA sequencing (dRNA-Seq) has recently been developed to simultaneously capture host and bacterial transcriptomes from an infected cell. This approach builds on the high sensitivity and resolution of RNA-Seq technology and is applicable to any bacteria that interact with eukaryotic cells, encompassing parasitic, commensal or mutualistic lifestyles. We pioneered dRNA-Seq to simultaneously capture prokaryotic and eukaryotic expression profiles of cells infected with bacteria, using in vitro Chlamydia-infected epithelial cells as proof of principle. Here we provide a detailed laboratory and bioinformatics protocol for dRNA-seq that is readily adaptable to any host-bacteria system of interest.

genomics

Multiplex PCR method for MinION and Illumina sequencing of Zika and other virus genomes directly from clinical samples

Genome sequencing has become a powerful tool for studying emerging infectious diseases; however, genome sequencing directly from clinical samples without isolation remains challenging for viruses such as Zika, where metagenomic sequencing methods may generate insufficient numbers of viral reads. Here we present a protocol for generating coding-sequence complete genomes comprising an online primer design tool, a novel multiplex PCR enrichment protocol, optimised library preparation methods for the portable MinION sequencer (Oxford Nanopore Technologies) and the Illumina range of instruments, and a bioinformatics pipeline for generating consensus sequences. The MinION protocol does not require an internet connection for analysis, making it suitable for field applications with limited connectivity. Our method relies on multiplex PCR for targeted enrichment of viral genomes from samples containing as few as 50 genome copies per reaction. Viral consensus sequences can be achieved starting with clinical samples in 1-2 days following a simple laboratory workflow. This method has been successfully used by several groups studying Zika virus evolution and is facilitating an understanding of the spread of the virus in the Americas.

genomics

Targeted capture of complete coding regions across divergent species

Despite continued advances in sequencing technologies, there is a need for methods that can efficiently sequence large numbers of genes from diverse species. One approach to accomplish this is targeted capture (hybrid enrichment). While these methods are well established for genome resequencing projects, cross-species capture strategies are still being developed and generally focus on the capture of conserved regions, rather than complete coding regions from specific genes of interest. The resulting data is thus useful for phylogenetic studies, but the wealth of comparative data that could be used for evolutionary and functional studies is lost. Here we design and implement a targeted capture method that enables recovery of complete coding regions across broad taxonomic scales. Capture probes were designed from multiple reference species and extensively tiled in order to facilitate cross-species capture. Using novel bioinformatics pipelines we were able to recover nearly all of the targeted genes with high completeness from species that were up to 200 myr divergent. Increased probe diversity and tiling for a subset of genes had a large positive effect on both recovery and completeness. The resulting data produced an accurate species tree, but importantly this same data can also be applied to studies of molecular evolution and function that will allow researchers to ask larger questions in broader phylogenetic contexts. Our method demonstrates the utility of cross-species approaches for the capture of full length coding sequences, and will substantially improve the ability for researchers to conduct large-scale comparative studies of molecular evolution and function.

evolutionary biology

Evaluating hybridization capture with RAD probes as a tool for museum genomics with historical bird specimens

Laboratory techniques for high-throughput sequencing have enhanced our ability to generate DNA sequence data from millions of natural history specimens collected prior to the molecular era, but remain poorly tested at shallower evolutionary time scales. Hybridization capture using restriction site associated DNA probes (hyRAD) is a recently developed method for population genomics with museum specimens (Suchan et al. 2016). The hyRAD method employs fragments produced in a restriction site associated double digestion as the basis for probes that capture orthologous loci in samples of interest. While promising in that it does not require a reference genome, hyRAD has yet to be applied across study systems in independent laboratories. Here we provide an independent assessment of the effectiveness of hyRAD on both fresh avian tissue and dried tissue from museum specimens up to 140 years old and investigate how variable quantities of input DNA affects sequencing, assembly, and population genetic inference. We present a modified bench protocol and bioinformatics pipeline, including three steps for detection and removal of microbial and mitochondrial DNA contaminants. We confirm that hyRAD is an effective tool for sampling thousands of orthologous SNPs from historic museum specimens to describe phylogeographic patterns. We find that modern DNA performs significantly better than historical DNA better during sequencing, but that assembly performance is largely equivalent. We also find that the quantity of input DNA predicts %GC content of assembled contiguous sequences, suggesting PCR bias. We caution against sampling schemes that include taxonomic or geographic autocorrelation across modern and historic samples.

genomics

In-Depth Resistome Analysis by Targeted Metagenomics

We developed ResCap, a targeted sequence capture platform based on SeqCapEZ technology, to analyse resistomes and other genes related to antimicrobial resistance (heavy metals, biocides and plasmids). ResCap includes probes for 8,667 canonical resistance genes (7,963 antibiotic resistance genes and 704 genes conferring resistance to metals or biocides), plus 2,517 relaxase genes (plasmid markers). Besides, it includes 78.600 genes homologous to the previous ones (47,806 for antibiotics and 30,794 for biocide or metals). ResCap enriched 279-fold the targeted sequences detected by metagenomic shotgun sequencing and improves their identification. Novel bioinformatic approaches allow quantifying \"gene abundance\" and \"gene diversity\". ResCap, the first targeted sequence capture specifically developed to analyse resistomes, enhances the sensitivity and specificity of available metagenomic methods to analyse antibiotic resistance in complex populations, enables the analysis of other genes related to antimicrobial resistance and opens the possibility to accurately study other complex microbial systems.

microbiology

Integrative genomics study of microglial transcriptome reveals effect of DLG4 (PSD95) on white matter in preterm infants.

Preterm birth places newborn infants in an adverse environment that leads to brain injury linked to neuroinflammation. To characterise this pathology, we present a translational bioinformatics investigation, with integration of human and mouse molecular and neuroimaging datasets to provide a deeper understanding of the role of microglia in preterm white matter damage. We examined preterm neuroinflammation in a mouse model of encephalopathy of prematurity induced by IL1B exposure, carrying out a gene network analysis of the cell-specific transcriptomic response to injury, which we extended to analysis of protein-protein interactions, transcription factors, and human brain gene expression, including translation to preterm infants by means of imaging-genetics approaches in the brain. We identified the endogenous synthesis of DLG4 (PSD95) protein by microglia in mouse and human, modulated by inflammation and development. Systemic genetic variation in DLG4 was associated with structural features in the preterm infant brain, suggesting that genetic variation in DLG4 may also impact white matter development and inter-individual susceptibility to injury.\n\nPreterm birth accounts for 11% of all births 1, and is the leading global cause of deaths under 5 years of age 2. Over 30% of survivors experience motor and/or cognitive problems from birth 3, 4, which last into adulthood 5. These problems include a 3-8 fold increased risk of symptoms and disorders associated with anxiety, inattention and social and communication problems compared to term-born infants 6. Prematurity is associated with a 4-12 fold increase in the prevalence of Autism Spectrum Disorders (ASD) compared to the general population 7, as well as a risk ratio of 7.4 for bipolar affective disorder among infants born below 32 weeks of gestation 8.\n\nThe characteristic brain injury observed in contemporary cohorts of preterm born infants includes changes to the grey and white matter tissues, that specifically include oligodendrocyte maturation arrest, hypomyelination and cortical changes visualised as decreases in fractional anisotropy 9-13. Exposure of the fetus and postnatal infant to systemic inflammation is an important contributing factor to brain injury in preterm born infants 12, 14, 15, and the persistence of inflammation is associated with poorer neurological outcome 16. Sources of systemic inflammation include maternal/fetal infections such as chorioamnionitis (which it is estimated affects a large number of women at a sub-clinical level), with the effect of systemic inflammation in the brain being mediated predominantly by the microglial response 17.\n\nMicroglia are unique yolk-sac derived resident phagocytes of the brain 18, 19, found preferentially within the developing white matter as a matter of normal developmental migration 12. Microglial products associated with white matter injury include pro-inflammatory cytokines, such as interleukin-1{beta} (IL1B) and tumour necrosis factor (TNF-)20, which can lead to a sub-clinical inflammatory situation associated with unfavourable outcomes 21. In addition to being key effector cells in brain inflammation, they are critical for normal brain development in processes such as axonal growth and synapse formation 22, 23. The role of microglia in neuroinflammation is dynamic and complex, reflected in their mutable phenotypes including both pro-inflammatory and restorative functions 24. Despite their important neurobiological role, the time course and nature of the microglial responses in preterm birth are currently largely unknown, and the interplay of inflammatory and developmental processes is also unclear. We, and others, believe that a better understanding of the molecular mechanisms underlying microglial function could harness their beneficial effects and mitigate the brain injury of prematurity and other states of brain inflammation25, 26\n\nA clinically relevant experimental mouse model of IL1B-induced systemic inflammation has been developed to study the changes occurring in the preterm human brain 27, 28. This model recapitulates the hallmarks of encephalopathy of prematurity including oligodendrocyte maturation delay with consequent dysmyelination, associated magnetic resonance imaging (MRI) phenotypes and behavioural deficits. Here, we take advantage of this model system to characterise the molecular underpinnings of the microglial response to IL1B-driven systemic inflammation and investigate its role in concurrent development.\n\nIn preterm infants MRI is used extensively to provide in-vivo correlates of white and grey matter pathology, allowing clinical assessment and prognostication. Diffusion MRI (d-MRI) measures the displacement of water molecules in the brain, and provides insight into the underlying tissue structure. Various d-MRI measures of white matter have been associated with developmental outcome in children born preterm 29-32, with up to 60% of inter-individual variability in structural and functional features attributable to genetic factors 33, 34. White matter abnormalities are linked to associated grey matter changes at both the imaging and cellular level 10, 35, 36, with functional and structural consequences lasting into adulthood 37, 38. Tract Based Statistics (TBSS) allows quantitative whole-brain white matter analysis of d-MRI data at the voxel level while avoiding problems due to contamination by signals arising from grey matter 39. This permits voxel-wise statistical testing and inferences to be made about group differences or associations with greater statistical power. TBSS has been shown to be an effective tool for studying white matter development and injury in the preterm brain 40, providing a macroscopic in vivo quantitative measure of white matter integrity that is associated with cognitive, fine motor, and gross motor outcome 11, 41, 42.\n\nIn this work we take a translational systems biology approach to investigate the role of microglia in preterm neuroinflammation and brain injury. We integrate microglial cell-type specific data from a mouse model of perinatal neuroinflammatory brain injury with experimental ex vivo and in vitro validation, translation to the human brain across the lifespan including analysis of human microglia, and assessment of the impact of genetic variation on structure of the preterm brain. We add to the understanding of the neurobiology of prematurity by: a) revealing the endogenous expression of DLG4 (PSD95) by microglia in early development, which is modulated by developmental stage and inflammation; and b) finding an association between systemic genetic variability in DLG4 and white matter structure in the preterm neonatal brain.

genomics

Asian lineage of Zika virus RNA pseudoknot may induce ribosomal frameshift and produce a new neuroinvasive protein ZIKV-NS1’

Zika virus (ZIKV) is a threat to humanity, and understanding its neuroinvasiveness is a major challenge. Microcephaly observed in neonates in Brazil is associated with ZIKV that belongs to the Asian lineage. What distinguishes the neuroinvasiveness between the RNA lineages from Asia and Africa is still unknown. Here we identify an aspect that may explain the different behavior between the two lineages. The distinction between the two groups is the occurrence of an alternative protein NS1 (ZIKV-NS1), which happens through a pseudoknot in the virus RNA that induces a ribosomal frameshift. Presence of NS1 protein is also observed in other Flavivirus that are neuroinvasive, and when NS1 production issuppressed, neuroinvasiveness is reduced.1 This evidence gives grounds to suggest that the ZIKV-NS1 occurring in the Asian lineage is responsible for neuro-tropism, which causes the neuro-pathologies associated with ZIKV infection, of which microcephaly is the most dev astating. The existence of ZIKV-NS1, which only exists in the Asian lineage, was inferred through bioinformatic methods, and it has yet to be experimentally observed. If its occurrence is confirmed, it will be a potential target in fighting the neuro-diseases associated with ZIKV.

epidemiology

Clusterization in head and neck squamous carcinomas based on lncRNA expression: molecular and clinical correlates

BackgroundLong non-coding RNAs (lncRNAs) have emerged as key players in a remarkably variety of biological processes and pathologic conditions, including cancer. Next-generation sequencing technologies and bioinformatics procedures predict the existence of tens of thousands of lncRNAs, from which we know the functions of only a handful of them, and very little is known in cancer types such as head and neck squamous cell carcinomas (HNSCCs).\n\nResultsHere, we use RNA-seq expression data from The Cancer Genome Atlas (TCGA) and various statistic and software tools in order to get insight about the lncRNome in HNSCC. Based on lncRNAs expression across 426 samples, we discover five distinct tumor clusters that we compare with reported clusters based on various genomic/genetic features. Results demonstrate significant associations between lncRNA-based clustering and DNA-methylation, TP53 mutation, and human papillomavirus infection. Using \"guilt by association\" procedures, we infer the possible biological functions of representative lncRNAs of each cluster. Furthermore, we found that lncRNA clustering is correlated with some important clinical and pathologic features, including patient survival after treatment, tumor grade or sub-anatomical location.\n\nConclusionsWe present a landscape of lncRNAs in HNSCC, and provide associations with important genotypic and phenotypic features that may help to understand the disease.

genomics

Heterogeneous chromatin mobility derived from chromatin states is a determinant of genome organisation in S. cerevisiae

Spatial organisation of the genome is essential for regulating gene activity, yet the mechanisms that shape this three-dimensional organisation in eukaryotes are far from understood. Here, we combine bioinformatic determination of chromatin states during normal growth and heat shock, and computational polymer modelling of genome structure, with quantitative microscopy and Hi-C to demonstrate that differential mobility of yeast chromosome segments leads to spatial self-organisation of the genome. We observe that more than forty percent of chromatin-associated proteins display a poised and heterogeneous distribution along the chromosome, creating a heteropolymer. This distribution changes upon heat shock in a concerted, state-specific manner. Simulating yeast chromosomes as heteropolymers, in which the mobility of each segment depends on its cumulative protein occupancy, results in functionally relevant structures, which match our experimental data. This thermodynamically driven self-organisation achieves spatial clustering of poised genes and mechanistically contributes to the directed relocalisation of active genes to the nuclear periphery upon heat shock.\n\nOne Sentence SummaryUnequal protein occupancy and chromosome segment mobility drive 3D organisation of the genome.

systems biology

Re-evaluating inheritance in genome evolution: widespread transfer of LINEs between species

Transposable elements (TEs) are mobile DNA sequences, colloquially known as jumping genes because of their ability to replicate to new genomic locations. Given a vector of transfer (e.g. tick or virus), TEs can jump further: between organisms or species in a process known as horizontal transfer (HT). Here we propose that LINE-1 (L1) and Bovine-B (BovB), the two most abundant TE families in mammals, were initially introduced as foreign DNA via ancient HT events. Using a 503-genome dataset, we identify multiple ancient L1 HT events in eukaryotes and provide evidence that L1s infiltrated the mammalian lineage after the monotreme-therian split. We also extend the BovB paradigm by increasing the number of estimated transfer events compared to previous studies, finding new potential blood-sucking parasite vectors and occurrences in new lineages (e.g. bats, frog). Given that these TEs make up nearly half of the genome sequence in todays mammals, our results provide the first evidence that HT can have drastic and long-term effects on the new host genomes. This revolutionizes our perception of genome evolution to consider external factors, such as the natural introduction of foreign DNA. With the advancement of genome sequencing technologies and bioinformatics tools, we anticipate our study to be the first of many large-scale phylogenomic analyses exploring the role of HT in genome evolution.\n\nSignificance statementLINE-1 (L1) elements occupy about half of most mammalian genomes (1), and they are believed to be strictly vertically inherited (2). Mutagenic L1 insertions are thought to account for approximately 1 of every 1000 random, disease-causing insertions in humans (4-7). Our research indicates that the very presence of L1s in humans, and other therian mammals, is due to an ancient transfer event - which has drastic implications for our perception of genome evolution. Using a machina analyses over 503 genomes, we trace the origins of L1 and BovB retrotransposons across the tree of life, and provide evidence of their long-term impact on eukaryotic evolution.

evolutionary biology

Validation and Implementation of CLIA-Compliant Whole Genome Sequencing (WGS) in Public Health Laboratory

BackgroundPublic health microbiology laboratories (PHL) are at the cusp of unprecedented improvements in pathogen identification, antibiotic resistance detection, and outbreak investigation by using whole genome sequencing (WGS). However, considerable challenges remain due to the lack of common standards.\n\nObjectives1) Establish the performance specifications of WGS applications used in PHL to conform with CLIA (Clinical Laboratory Improvements Act) guidelines for laboratory developed tests (LDT), 2) Develop quality assurance (QA) and quality control (QC) measures, 3) Establish reporting language for end users with or without WGS expertise, 4) Create a validation set of microorganisms to be used for future validations of WGS platforms and multi-laboratory comparisons and, 5) Create modular templates for the validation of different sequencing platforms.\n\nMethodsMiSeq Sequencer and Illumina chemistry (Illumina, Inc.) were used to generate genomes for 34 bacterial isolates with genome sizes from 1.8 to 4.7 Mb and wide range of GC content (32.1%-66.1%). A customized CLCbio Genomics Workbench - shell script bioinformatics pipeline was used for the data analysis.\n\nResultsWe developed a validation panel comprising ten Enterobacteriaceae isolates, five gram-positive cocci, five gram-negative non-fermenting species, nine Mycobacterium tuberculosis, and five miscellaneous bacteria; the set represented typical workflow in the PHL. The accuracy of MiSeq platform for individual base calling was >99.9% with similar results shown for reproducibility/repeatability of genome-wide base calling. The accuracy of phylogenetic analysis was 100%. The specificity and sensitivity inferred from MLST and genotyping tests were 100%. A test report format was developed for the end users with and without WGS knowledge.\n\nConclusionWGS was validated for routine use in PHL according to CLIA guidelines for LDTs. The validation panel, sequencing analytics, and raw sequences will be available for future multi-laboratory comparisons of WGS in PHL. Additionally, the WGS performance specifications and modular validation template are likely to be adaptable for the validation of other platforms and reagents kits.

microbiology

Genomic analysis of P elements in natural populations of Drosophila melanogaster

The Drosophila melanogaster P transposable element provides one of the best cases of horizontal transfer of a mobile DNA sequence in eukaryotes. Invasion of natural populations by the P element has led to a syndrome of phenotypes known as P-M hybrid dysgenesis that emerges when strains differing in their P element composition mate and produce offspring. Despite extensive research on many aspects of P element biology, many questions remain about the genomic basis of variation in P-M dysgenesis phenotypes in natural populations. Here we compare gonadal dysgenesis phenotypes and genomic P element predictions for isofemale strains obtained from three worldwide populations of D. melanogaster to illuminate the molecular basis of natural variation in cytotype status. We show that the number of predicted P element insertions in genome sequences from isofemale strains is highly correlated across different bioinformatics methods, but the absolute number of insertions per strain is sensitive to method and filtering strategies. Regardless of method used, we find that the number of euchromatic P element insertions predicted per strain varies significantly across populations, with strains from a North American population having fewer P element insertions than strains from populations sampled in Europe or Africa. Despite these geographic differences, numbers of euchromatic P element insertions are not strongly correlated with the degree of gonadal dysgenesis exhibited by an isofemale strain. Thus, variation in P element insertion numbers across different populations does not necessarily lead to corresponding geographic differences in gonadal dysgenesis phenotypes. Additionally, we show that pool-seq samples can uncover population differences in the number of P element insertions observed from isofemale lines, but that efforts to rigorously detect differences in the number of P elements across populations using pool-seq data must properly control for read depth per strain. Our work supports the view that euchromatic P element copy number is not sufficient to explain variation in gonadal dysgenesis across strains of D. melanogaster, and informs future efforts to decode the genomic basis of geographic and temporal differences in P element induced phenotypes.

genomics

Unmet Needs for Analyzing Biological Big Data: A Survey of 704 NSF Principal Investigators

In a 2016 survey of 704 National Science Foundation (NSF) Biological Sciences Directorate principal investigators (BIO PIs), nearly 90% indicated they are currently or will soon be analyzing large data sets. BIO PIs considered a range of computational needs important to their work--including high performance computing (HPC), bioinformatics support, multi-step workflows, updated analysis software, and the ability to store, share, and publish data. Previous studies in the United States and Canada emphasized infrastructure needs. However, BIO PIs said the most pressing unmet needs are training in data integration, data management, and scaling analyses for HPC--acknowledging that data science skills will be required to build a deeper understanding of life. This portends a growing data knowledge gap in biology and challenges institutions and funding agencies to redouble their support for computational training in biology.

scientific communication and education

A unifying mechanism for the biogenesis of prokaryotic membrane proteins co-operatively integrated by the Sec and Tat pathways

The vast majority of polytopic membrane proteins are inserted into the cytoplasmic membrane of prokaryotes by the general secretory (Sec) pathway. However, a subset of monotopic proteins that contain non-covalently-bound redox cofactors depend on the twin-arginine translocase (Tat) machinery for membrane integration. Recently actinobacterial Rieske iron-sulfur cluster-containing proteins were identified as an unusual class of membrane proteins that require both the Sec and Tat pathways for the insertion of their three transmembrane domains (TMDs). The Sec pathway inserts the first two TMDs of these proteins co-translationally, but releases the polypeptide prior to the integration of TMD3 to allow folding of the cofactor-containing domain and its translocation by Tat. Here we have investigated features of the Streptomyces coelicolor Rieske polypeptide that modulate its interaction with the Sec and Tat machineries. Mutagenesis of a highly conserved loop region between Sec-dependent TMD2 and Tat-dependent TMD3 shows that it plays no significant role in coordinating the activities of the two translocases, but that a minimum loop length of approximately eight amino acids is required for the Tat machinery to recognise TMD3. Instead we show that a combination of relatively low hydrophobicity of TMD3, coupled with the presence of C-terminal positively-charged amino acids, results in abortive insertion of TMD3 by the Sec pathway and its release at the cytoplasmic side of the membrane. Bioinformatic analysis identified two further families of polytopic membrane proteins that share features of dual Sec-Tat-targeted membrane proteins. A predicted heme-molybdenum cofactor-containing protein with five TMDs, and a polyferredoxin also with five predicted TMDs, are encoded across bacterial and archaeal genomes. We demonstrate that membrane insertion of representatives of each of these newly-identified protein families is dependent on more than one protein translocase, with the Tat machinery recognising TMD5. Importantly, the combination of low hydrophobicity of the final TMD and the presence of multiple C-terminal positive charges that serve as critical Sec-release features for the actinobacterial Rieske protein also dictate Sec release in these further protein families. Therefore we conclude that a simple unifying mechanism governs the assembly of dual targeted membrane proteins.

microbiology