Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

ExomeSlicer: a resource for the development and validation of exome-based clinical panels

Exome-based panels (exome slices) are becoming the preferred diagnostic strategy in clinical laboratories, especially for genetically heterogeneous disorders. The advantages of this approach include enabling frequent updates to gene content without the need for re-designing, reflexing to exome analysis bioinformatically without requiring additional sequencing, and streamlining laboratory operation by using established exome kits and protocols. Despite their increasing use, there are currently no guidelines or appropriate resources to support their clinical implementation. Here, we highlight principles and important considerations for the clinical development and validation of exome-based panels, guided by clinical data from a diagnostic epilepsy panel using this approach. We also present a novel, publically accessible web-based resource, ExomeSlicer, and demonstrate its clinical utility in predicting gene-specific and exome-wide technically challenging regions that are not amenable to Next Generation Sequencing (NGS), and that might significantly lead to increased post hoc Sanger fill in burden. Using this tool, we also characterize > 2000 low complexity, GC-rich and/or high homology, regions across the exome that can be a source of false positive or false negative variant calls thus potentially leading to misdiagnoses in tested patients.

genomics

Conserved and ubiquitous expression of piRNAs and PIWI genes in mollusks antedates the origin of somatic PIWI/piRNA expression to the root of bilaterians

PIWI proteins and a specific class of small non-coding RNAs, termed Piwi interacting RNAs (piRNAs), suppress transposon activity in animals on the transcriptional and post-transcriptional level, thus protecting genomes from detrimental insertion mutagenesis. While in vertebrates the PIWI/piRNA system appears to be restricted to the germline, somatic expression of piRNAs directed against transposons is widespread in arthropods, likely representing the ancestral state for this phylum. Here, we show that somatic expression of PIWI genes and piRNAs directed against transposons is conserved in mollusks, suggesting that somatic PIWI/piRNA expression was already realized in an early bilaterian ancestor. We further describe lineage specific adaptations regarding transposon composition of piRNA clusters and show that different piRNA clusters are dynamically expressed during oyster development. Finally, bioinformatics analyses suggest that different populations of piRNAs participate in the ping-pong amplification loop in a tissue specific manner.

evolutionary biology

A conserved mechanism for regulation of endo-lysosomal pH by histone deacetylases

The pH of the endo-lysosomal system is tightly regulated by a balance of proton pump and leak mechanisms that are critical for storage, recycling, turnover and signaling functions in the cell. Dysregulation of endo-lysosomal pH has been linked to aging, amyloidogenesis, synaptic dysfunction, and various neurodegenerative disorders including Alzheimers disease. Therefore, understanding mechanisms that regulate luminal pH may be key to identifying new targets for treatment of these disorders. Meta-analysis of yeast microarray databases revealed that nutrient limiting conditions upregulated transcription of the endosomal Na+/H+ exchanger Nhx1 by inhibition of the histone deacetylase (HDAC) Rpd3, resulting in vacuolar alkalinization. Consistent with these findings, Rpd3 inhibition by the HDAC inhibitor and antifungal drug trichostatin A induced Nhx1 expression and vacuolar alkalinization. Bioinformatics analysis of Drosophila and mouse databases revealed that caloric control of Nhx1 orthologs DmNHE3 and NHE6 respectively, was also mediated by histone deacetylases. We show that NHE6 is a target of cAMP-response element-binding (CREB) protein, providing a molecular mechanism for nutrient and HDAC dependent regulation of endosomal pH. Control of NHE6 expression by pharmacological targeting of the CREB pathway can be used to regulate endosomal pH and restore defective amyloid A{beta} clearance in an ApoE4 astrocyte model of Alzheimers disease. These observations from yeast, fly, mouse and cell culture models reveal an evolutionarily conserved mechanism for regulation of endosomal NHE expression by histone deacetylases and offer new therapeutic strategies for modulation of endo-lysosomal pH in fungal infection and human disease.

molecular biology

Control of artefactual variation in reported inter-sample relatedness during clinical use of a Mycobacterium tuberculosis sequencing pipeline

Contact tracing requires reliable identification of closely related bacterial isolates. When we noticed the reporting of artefactual variation between M. tuberculosis isolates during routine next generation sequencing of Mycobacterium spp, we investigated its basis in 2,018 consecutive M. tuberculosis isolates. In the routine process used, clinical samples were decontaminated and inoculated into broth cultures; from positive broth cultures DNA was extracted, sequenced, reads mapped, and consensus sequences determined. We investigated the process of consensus sequence determination, which selects the most common nucleotide at each position. Having determined the high-quality read depth and depth of minor variants across 8,006 M. tuberculosis genomic regions, we quantified the relationship between the minor variant depth and the amount of non-Mycobacterial bacterial DNA, which originates from commensal microbes killed during sample decontamination. In the presence of non-Mycobacterial bacterial DNA, we found significant increases in minor variant frequencies of more than 1.5 fold in 242 regions covering 5.1% of the M. tuberculosis genome. Included within these were four high variation regions strongly influenced by the amount of non-Mycobacterial bacterial DNA. Excluding these four regions from pairwise distance comparisons reduced biologically implausible variation from 5.2% to 0% in an independent validation set derived from 226 individuals. Thus, we have demonstrated an approach identifying critical genomic regions contributing to clinically relevant artefactual variation in bacterial similarity searches. The approach described monitors the outputs of the complex multi-step laboratory and bioinformatics process, allows periodic process adjustments, and will have application to quality control of routine bacterial genomics.

microbiology

Analysis of proteins in computational models of synaptic plasticity

The desire to explain how synaptic plasticity arises from interactions between ions, proteins and other signalling molecules has propelled the development of biophysical models of molecular pathways in hippocampal, striatal and cerebellar synapses. The experimental data underpinning such models is typically obtained from low-throughput, hypothesis-driven experiments. We used high-throughput proteomic data and bioinformatics datasets to assess the coverage of biophysical models.\n\nTo determine which molecules have been modelled, we surveyed biophysical models of synaptic plasticity, identifying which proteins are involved in each model. We were able to map 4.2% of previously reported synaptic proteins to entities in biophysical models. Linking the modelled protein list to Gene Ontology terms shows that modelled proteins are focused on functions such as calmodulin binding, cellular responses to glucagon stimulus, G-alpha signalling and DARPP-32 events.\n\nWe cross-linked the set of modelled proteins with sets of genes associated with common neurological diseases. We find some examples of disease-associated molecules that are well represented in models, such as voltage-dependent calcium channel family (CACNA1C), dopamine D1 receptor, and glutamate ionotropic NMDA type 2A and 2B receptors. Many other disease-associated genes have not been included in models of synaptic plasticity, for example catechol-O-methyltransferase (COMT) and MAO A. By incorporating pathway enrichment results, we identify LAMTOR, a gene uniquely associated with Schizophrenia, which is closely linked to the MAPK pathway found in some models.\n\nOur analysis provides a map of how molecular pathways underpinning neurological diseases relate to synaptic biophysical models that can in turn be used to explore how these molecular events might bridge scales into cellular processes and beyond. The map illustrates disease areas where biophysical models have good coverage as well as domain gaps that require significant further research.\n\nAuthor summaryThe 100 billion neurons in the human brain are connected by a billion trillion structures called synapses. Each synapse contains hundreds of different proteins. Some proteins sense the activity of the neurons connecting the synapse. Depending on what they sense, the proteins in the synapse are rearranged and new proteins are synthesised. This changes how strongly the synapse influences its target neuron, and underlies learning and memory. Scientists build computational models to reason about the complex interactions between proteins. Here we list the proteins that have been included in computational models to date. For good reasons, models do not always specify proteins precisely, so to make the list we had to translate the names used for proteins in models to gene names, which are used to identify proteins. Our translation could be used to label computational models in the future. We found that the list of modelled proteins contains only 4.2% of proteins associated with synapses, suggesting more proteins should be added to models. We used lists of genes associated with neurological diseases to suggest proteins to include in future models.

neuroscience

Identification of Variable and Joining germline genes and alleles for Rhesus macaque from B-cell receptor repertoires

The Rhesus macaque is a valuable preclinical animal model to estimate vaccine effectiveness, and is also important for understanding antibody maturation and B-cell repertoire evolution responding to vaccination; however, incomplete mapping of rhesus immunoglobulin germline genes hinders the research efforts. To address this deficiency, we sequenced B-cell receptor (BCR) repertoires of 75 India Rhesus macaques. Using a bioinformatic method that has been validated with BCR repertoire analysis of three human donors, we were able to infer rhesus Variable (V) and Joint(J) germline alleles, identifying a total of 122 V and 20 J germline alleles. Importantly, 91 V and 13 J alleles were novel, and 40 V and 13 J genes were found at a novel genome region that has not been previously recorded. The novelty of these newly identified alleles was supported by two observations. Firstly, 50 V and 5 J novel alleles were observed in whole genome sequencing data of 10 Rhesus macaques. Secondly, using alignment reference including the novel alleles, the mutation rate of rearranged repertoires was significant declined in 9 other irrelevant samples, and all our identified novel V and J alleles were 100% identity mapped by rearranged repertoire data. These newly identified novel alleles, along with previous reported alleles, provide an important reference for future investigations of rhesus immune repertoire evolution, in response to vaccination or infection. In addition, the method outlined in our study offered an example to future efforts in identifying novel immunoglobulin alleles.

immunology

Exploratory re-encoding of Yellow Fever Virus genome: analysis of in vitro and in vivo replicative phenotypes

Virus attenuation by genome re-encoding is a pioneering approach for generating effective live-attenuated vaccine candidates. Its core principle is to introduce a large number of synonymous substitutions into the viral genome to produce stable attenuation of the targeted virus. Introduction of large numbers of mutations has also been shown to maintain stability of the attenuated phenotype by lowering the risk of reversion and recombination of re-encoded genomes. Identifying mutations with low fitness cost is pivotal as this increases the number that can be introduced and generates more stable and attenuated viruses. Here, we sought to identify mutations with low deleterious impact on the in vivo replication and virulence of yellow fever virus (YFV). Following comparative bioinformatic analyses of flaviviral genomes, we categorized synonymous transition mutations according to their impact on CpG/UpA composition and secondary RNA structures. We then designed 17 re-encoded viruses with 100-400 synonymous mutations in the NS2A-to-NS4B coding region of YFV Asibi and Ap7M (hamster-adapted) genomes. Each virus contained a panel of synonymous mutations designed according to the above categorisation criteria. The replication and fitness characteristics of parent and re-encoded viruses were compared in vitro using cell culture competition experiments. In vivo laboratory hamster models were also used to compare relative virulence and immunogenicity characteristics. Most of the re-encoded strains showed no decrease in replicative fitness in vitro. However, they showed reduced virulence and, in some instances, decreased replicative fitness in vivo. Importantly, the most attenuated of the re-encoded strains induced robust, protective immunity in hamsters following challenge with Ap7M, a virulent virus. Overall, the introduction of transitions with no or a marginal increase in the number of CpG/UpA dinucleotides had the mildest impact on YFV replication and virulence in vivo. Thus, this strategy can be incorporated in procedures for the finely tuned creation of substantially re-encoded viral genomes.

microbiology

KLF6 activates complementary gene modules relevant to axon growth and promotes corticospinal tract regeneration after spinal injury

Members of the KLF family of transcription factors can exert both positive and negative effects on axon regeneration in the central nervous system, but the underlying mechanisms are unclear. KLF6 and -7 share nearly identical DNA binding domains and stand out as the only known growth-promoting family members. Here we confirm that similar to KLF7, expression of KLF6 declines during postnatal cortical development and that forced re-expression of KLF6 in corticospinal tract neurons of adult female mice enhances axon regeneration after cervical spinal injury. Unlike KLF7, however, these effects were achieved with wildtype KLF6, as opposed constitutively active mutants, thus simplifying the interpretation of mechanistic studies. To clarify the molecular basis of growth promotion, RNA sequencing identified 454 genes whose expression changed upon forced KLF6 expression in cortical neurons. Network analysis of these genes revealed sub-networks of downregulated genes that were highly enriched for synaptic functions, and sub-networks of upregulated genes with functions relevant to axon extension including cytoskeleton remodeling, lipid synthesis and transport, and bioenergetics. The promoter regions of KLF6-sensitive genes showed enrichment for the binding sequence of STAT3, a previously identified regeneration-associated gene. Notably, co-expression of constitutively active STAT3 along with KLF6 in cortical neurons produced synergistic increases in neurite length. Finally, genome-wide ATAC-seq footprinting detected frequent co-binding by the two factors in pro-growth gene networks, indicating co-occupancy as an underlying mechanism for the observed synergy. These findings advance understanding of KLF-stimulated axon growth and indicate functional synergy of KLF6 transcriptional effects with those of STAT3.\n\nSIGNIFICANCE STATEMENTThe failure of axon regeneration in the CNS limits recovery from damage and disease. These findings show the transcription factor KLF6 to be a potent promoter of axon growth after spinal injury, and more importantly clarify the underlying transcriptional changes. In addition, bioinformatics analysis predicted a functional interaction between KLF6 and a second transcription factor, STAT3, and genome-wide footprinting confirmed frequent co-occupancy. Co-expression of the two factors yielded synergistic elevation of neurite growth in primary neurons. These data point the way toward novel transcriptional interventions to promote CNS regeneration.

neuroscience

Divergent PAM Specificity of a Highly-Similar SpCas9 Ortholog

RNA-guided DNA endonucleases of the CRISPR-Cas system are widely used for genome engineering and thus have numerous applications in a wide variety of fields. The range of sequences that CRISPR endonucleases can recognize, however, is constrained by the need for a specific protospacer adjacent motif (PAM) flanking the target site. In this study, we demonstrate the natural PAM plasticity of a highly-similar, yet previously uncharacterized, Cas9 from Streptococcus canis (ScCas9) through rational manipulation of distinguishing motif insertions. To this end, we report a divergent affinity to 5-NNGT-3 PAM sequences and demonstrate the editing capabilities of the ortholog in both bacterial and human cells. Finally, we build an automated bioinformatics pipeline, the Search for PAMs by ALignment Of Targets (SPAMALOT), which further explores the microbial PAM diversity of otherwise-overlooked Streptococcus Cas9 orthologs. Our results establish that ScCas9 can be utilized both as an alternative genome editing tool and as a functional platform to discover novel Streptococcus PAM specificities.

microbiology

Developmental chromatin restriction of pro-growth gene networks acts as an epigenetic barrier to axon regeneration in cortical neurons

Axon regeneration in the central nervous system is prevented in part by a developmental decline in the intrinsic regenerative ability of maturing neurons. This loss of axon growth ability likely reflects widespread changes in gene expression, but the mechanisms that drive this shift remain unclear. Chromatin accessibility has emerged as a key regulatory mechanism in other cellular contexts, raising the possibility that chromatin structure may contribute to the age-dependent loss of regenerative potential. Here we establish an integrated bioinformatic pipeline that combines analysis of developmentally dynamic gene networks with transcription factor regulation and genome-wide maps of chromatin accessibility. When applied to the developing cortex, this pipeline detected overall closure of chromatin in sub-networks of genes associated with axon growth. We next analyzed mature CNS neurons that were supplied with various pro-regenerative transcription factors. Unlike prior results with SOX11 and KLF7, here we found that neither JUN nor an activated form of STAT3 promoted substantial corticospinal tract regeneration. Correspondingly, chromatin accessibility in JUN or STAT3 target genes was substantially lower than in predicted targets of SOX11 and KLF7. Finally, we used the pipeline to predict pioneer factors that could potentially relieve chromatin constraints at growth-associated loci. Overall this integrated analysis substantiates the hypothesis that dynamic chromatin accessibility contributes to the developmental decline in axon growth ability and influences the efficacy of pro-regenerative interventions in the adult, while also pointing toward selected pioneer factors as high-priority candidates for future combinatorial experiments.

neuroscience

Template switching causes artificial junction formation and false identification of circular RNAs

Hundreds of thousands of putative circular RNAs have been identified through deep sequencing and bioinformatic analyses. However, the circularity of these putative RNA circles has not been experimentally validated due to limited methodologies currently available. We reported here that the template-switching capability of commonly used reverse transcriptases (e.g., SuperScript II) leads to the formation of artificial junction sequences, and consequently misclassification of large linear RNAs as RNA circles. Use of reverse transcriptases without terminal transferase activity (e.g., MonsterScript) for cDNA synthesis is critical for the identification of physiological circular RNAs. We also report two methods, MonsterScript junction PCR and high-resolution melting curve analyses, which can reliably distinguish circular RNAs from their linear forms and thus, can be used to discover and validate true circular RNAs.\n\nSignificance StatementThe vast majority of circular RNAs were identified through computational detection of junction sequences in the deep sequencing reads because these unique fusion sequences represent back-splicing events. We found that artificial junction sequences could be formed through template switching (TS) when MMLV-derived reverse transcriptases, e.g., SuperScript II, are used to synthesize cDNAs. Thus, many of the reported circular RNAs may not be RNA circles, but rather experimental artifacts. Fake circular RNAs can be avoided by using reverse transcriptases without terminal transferase activity (e.g., MonsterScript) for cDNA synthesis. We developed two novel methods, MonsterScript junction PCR and high-resolution melting curve analyses, for distinguishing circular RNAs from their linear form.

molecular biology

Gene Coregulation and Coexpression in the Aryl Hydrocarbon Receptor-mediated Transcriptional Regulatory Network in the Mouse Liver

Tissue-specific network models of chemical-induced gene perturbation can improve our mechanistic understanding of the intracellular events leading to adverse health effects resulting from chemical exposure. The aryl hydrocarbon receptor (AHR) is a ligand-inducible transcription factor (TF) that activates a battery of genes and produces a variety of species-specific adverse effects in response to the potent and persistent environmental contaminant 2,3,7,8-tetrachlorodibenzo-p-dioxin (TCDD). Here we assemble a global map of the AHR gene regulatory network in TCDD-treated mouse liver from a combination of previously published gene expression and genome-wide TF binding data sets. Using Kohonen selforganizing maps and subspace clustering, we show that genes co-regulated by common upstream TFs in the AHR network exhibit a pattern of co-expression. Specifically, directly-bound, indirectly-bound and non-genomic AHR target genes exhibit distinct patterns of gene expression, with the directly bound targets generally associated with highest median expression. Further, among the directly bound AHR target genes, the expression level increases with the number of AHR binding sites in the proximal promoter regions. Finally, we show that co-regulated genes in the AHR network activate distinct groups of downstream biological processes, with AHR-bound target genes enriched for metabolic processes and enrichment of immune responses among AHR-unbound target genes, likely reflecting infiltration of immune cells into the mouse liver upon TCDD treatment. This work describes an approach to the reconstruction and analysis of transcriptional regulatory cascades underlying cellular stress response using bioinformatic and statistical tools.

genomics

Long-read sequencing reveals the splicing profile of the calcium channel gene CACNA1C in human brain

RNA splicing is a key mechanism linking genetic variation with psychiatric disorders. Splicing profiles are particularly diverse in brain and difficult to accurately identify and quantify. We developed a new approach to address this challenge, combining long-range PCR and nanopore sequencing with a novel bioinformatics pipeline. We identify the full-length coding transcripts of CACNA1C in human brain. CACNA1C is a psychiatric risk gene that encodes the voltage-gated calcium channel CaV1.2. We show that CACNA1Cs transcript profile is substantially more complex than appreciated, identifying 38 novel exons and 241 novel transcripts. Importantly, many of the novel variants are abundant, and predicted to encode channels with altered function. The splicing profile varies between brain regions, especially in cerebellum. We demonstrate that human transcript diversity (and thereby protein isoform diversity) remains under-characterised, and provide a feasible and cost-effective methodology to address this. A detailed understanding of isoform diversity will be essential for the translation of psychiatric genomic findings into pathophysiological insights and novel psychopharmacological targets.

neuroscience

Genome-wide study identifies 611 loci associated with risk tolerance and risky behaviors

Humans vary substantially in their willingness to take risks. In a combined sample of over one million individuals, we conducted genome-wide association studies (GWAS) of general risk tolerance, adventurousness, and risky behaviors in the driving, drinking, smoking, and sexual domains. We identified 611 approximately independent genetic loci associated with at least one of our phenotypes, including 124 with general risk tolerance. We report evidence of substantial shared genetic influences across general risk tolerance and risky behaviors: 72 of the 124 general risk tolerance loci contain a lead SNP for at least one of our other GWAS, and general risk tolerance is moderately to strongly genetically correlated ([Formula] to 0.50) with a range of risky behaviors. Bioinformatics analyses imply that genes near general-risk-tolerance-associated SNPs are highly expressed in brain tissues and point to a role for glutamatergic and GABAergic neurotransmission. We find no evidence of enrichment for genes previously hypothesized to relate to risk tolerance.

genetics

NeoDTI: Neural integration of neighborinformation from a heterogeneous network fordiscovering new drug-target interactions

MotivationAccurately predicting drug-target interactions (DTIs) in silico can guide the drug discovery process and thus facilitate drug development. Computational approaches for DTI prediction that adopt the systems biology perspective generally exploit the rationale that the properties of drugs and targets can be characterized by their functional roles in biological networks.\n\nResultsInspired by recent advance of information passing and aggregation techniques that generalize the convolution neural networks (CNNs) to mine large-scale graph data and greatly improve the performance of many network-related prediction tasks, we develop a new nonlinear end-to-end learning model, called NeoDTI, that integrates diverse information from heterogeneous network data and automatically learns topology-preserving representations of drugs and targets to facilitate DTI prediction. The substantial prediction performance improvement over other state-of-the-art DTI prediction methods as well as several novel predicted DTIs with evidence supports from previous studies have demonstrated the superior predictive power of NeoDTI. In addition, NeoDTI is robust against a wide range of choices of hyperparameters and is ready to integrate more drug and target related information (e.g., compound-protein binding affinity data). All these results suggest that NeoDTI can offer a powerful and robust tool for drug development and drug repositioning.\n\nAvailability and implementationThe source code and data used in NeoDTI are available at: https://github.com/FangpingWan/NeoDTI.\n\nContactzengjy321@tsinghua.edu.cn\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

systems biology

Functional characterization of sensory neuron membrane proteins (SNMPs)

Sensory neuron membrane proteins (SNMPs) play a critical role in the insect olfactory system but there is a deficit of functional studies beyond Drosophila. Here, we provide functional characterisation of insect SNMPs through the use of bioinformatics, genome curation, transcriptome data analysis, phylogeny, expression profiling, and RNAi gene knockdown techniques. We curated 81 genes from 35 insect species and identified a novel lepidopteran SNMP gene family, SNMP3. Phylogenetic analysis shows that lepidopteran SNMP3, but not the previously annotated lepidopteran SNMP2, is the true homologue of the dipteran SNMP2. Digital expression, microarray and qPCR analyses show that the lepidopteran SNMP1 is specifically expressed in adult antennae. SNMP2 is widely expressed in multiple tissues while SNMP3 is specifically expressed in the larval midgut. Microarray analysis suggest SNMP3 may be involved in the silkworm immunity response to virus and bacterial infections. We functionally characterised SNMP1 in the silkworm using RNAi and behavioural assays. Our results suggested that Bombyx mori SNMP1 is a functional orthologue of the Drosophila melanogaster SNMP1 and plays a critical role in pheromone detection. Split-ubiquitin yeast hybridization study shows that BmorSNMP1 has a protein-protein interaction with the BmorOR1 pheromone receptor, and the BmorOrco co-receptor. Concluding, we propose a novel molecular model in which BmorOrco, BmorSNMP1 and BmorOR1 form a heteromer in the detection of the silkworm sex pheromone bombykol.

molecular biology

Structural model of Cyc2, the primary electron acceptor of Acidithiobacillus ferrooxidans respiratory chain, as a modular cytochrome - β-barrel fusion protein, and mechanistic proposals based on this model

Acidithiobacillus ferrooxidans oxidizes Fe(II) to Fe(III) to feed electrons into its respiratory chain. The primary electron acceptor of this complex system is Cyc2, an outer membrane protein of unknown structure. This work proposes a feasible model of Cyc2s global structure, based on homology modeling, residue-residue coevolution data, bioinformatics predictions and limited knowledge about Cyc2s function. The proposal is that the sequence segment spanning residues ~30 to ~90 folds as a cytochrome-like domain that contains a heme group which would presumably bind and oxidize external Fe(II), whereas the remaining segment from residue ~90 until the end adopts a {beta}-barrel fold similar to that of most outer membrane proteins. Such model differs strongly from a published model, but is backed up by more data and is more compatible with the known topology of outer membrane proteins and with Cyc2s function of internalizing reducing equivalents. The small size of the cytochrome-like domain would allow it to reside inside, and/or slide through, the {beta}-barrel domain, thus communicating in a controlled fashion the extracellular medium with the periplasm to import electrons through the outer membrane. All the models discussed are provided as PyMOL session files in the Supporting Information and can be visualized online at http://lucianoabriata.altervista.org/modelshome.html

biophysics

Primate MHC class I from Genomes

The major histocompatibility complex (MHC) molecule plays a central role in the adaptive immunity of jawed vertebrates. Allelic variations have been studied extensively in some primate species, however a comprehensive description of the number of genes remains incomplete. Here, a bioinformatics program was developed to identify three MHC Class I exons (EX2, EX3 and EX4) from Whole Genome Sequencing (WGS) datasets. With this algorithm, MHC Class I exons sequences were extracted from 30 WGS datasets of primates, representatives of Apes, Old World and New World monkeys and prosimians. There is a high variability in the number of genes between species. From human WGS, six viable genes (HLA-A, -B, -C, -E, -F, and -G) and four pseudogene sequences (HLA-H, -J, -L, -V) are obtained. These genes serve to identify the phylogenetic clades of MHC-I in primates. The results indicate that human clades of HLA-A -B and -C were generated shortly after the separation of Old World monkeys. The clades pertaining to HLA-E, -H and -F are found in all primate families, except in Prosimians. In the clades defined by HLA-G, -L and -J, there are sequences from Old world monkeys. Specific clades are found in the four primate families. The evolution of these genes is consistent with birth and death processes having a high turnover rates.

immunology