Search bioRxivSearch

Biology subjects

Khan, A.

Publications and source records attributed to Khan, A..

16 recordsLinked to original sources

Cohort Profile: East London Genes & Health (ELGH), a community based population genomics and health study in people of British-Bangladeshi and -Pakistani heritage.

Cohort profile in a nutshellO_LIEast London Genes & Health (ELGH) is a large scale, community genomics and health study (to date >34,000 volunteers; target 100,000 volunteers). C_LIO_LIELGH was set up in 2015 to gain deeper understanding of health and disease, and underlying genetic influences, in British-Bangladeshi and British-Pakistani people living in east London. C_LIO_LIELGH prioritises studies in areas important to, and identified by, the community it represents. Current priorities include cardiometabolic diseases and mental illness, these being of notably high prevalence and severity. However studies in any scientific area are possible, subject to community advisory group and ethical approval. C_LIO_LIELGH combines health data science (using linked UK National Health Service (NHS) electronic health record data) with exome sequencing and SNP array genotyping to elucidate the genetic influence on health and disease, including the contribution from high rates of parental relatedness on rare genetic variation and homozygosity (autozygosity), in two understudied ethnic groups. Linkage to longitudinal health record data enables both retrospective and prospective analyses. C_LIO_LIThrough Stage 2 studies, ELGH offers researchers the opportunity to undertake recall-by-genotype and/or recall-by-phenotype studies on volunteers. Sub-cohort, trial-within-cohort, and other study designs are possible. C_LIO_LIELGH is a fully collaborative, open access resource, open to academic and life sciences industry scientific research partners. C_LI

genomics

Modeling RNA-binding protein specificity in vivo by precisely registering protein-RNA crosslink sites

RNA-binding proteins (RBPs) regulate post-transcriptional gene expression by recognizing short and degenerate sequence elements in their target transcripts. Despite the expanding list of RBPs with in vivo binding sites mapped genomewide using crosslinking and immunoprecipitation (CLIP), defining precise RBP binding specificity remains challenging. We previously demonstrated that the exact protein-RNA crosslink sites can be mapped using CLIP data at single-nucleotide resolution and observed that crosslinking frequently occurs at specific positions in RBP motifs. Here we have developed a computational method, named mCross, to jointly model RBP binding specificity while precisely registering the crosslinking position in motif sites. We applied mCross to 112 RBPs using ENCODE eCLIP data and validated the reliability of the resulting motifs by genome-wide analysis of allelic binding sites also detected by CLIP. We found that the prototypical SR protein SRSF1 recognizes GGA clusters to regulate splicing in a much larger repertoire of transcripts than previously appreciated.

molecular biology

Evolutionary history of Alzheimer Disease causing protein family Presenilins with pathological implications.

Presenilin proteins are type II transmembrane proteins. They make the catalytic component of Gamma secretase, a multiportion transmembrane protease. Amyloid protein, Notch and beta catenin are among more than 90 substrates of Presenilins. Mutations in Presenilins lead to defects in proteolytic cleavage of its substrate resulting in some of the most devastating pathological conditions including Alzheimer disease (AD), developmental disorders and cancer. In addition to catalytic roles, Presenilin protein is also shown to be involved in many non-catalytic roles i.e. calcium homeostasis, regulation of autophagy and protein trafficking etc. These proteolytic proteins are highly conserved, present in almost all the major eukaryotic groups. Studies on wide variety of organisms ranging from human to unicellular dictyostelium have shown the important catalytic and non-catalytic roles of Presenilins. In the current research project, we aimed to elucidate the phylogenetic history of Presenilins. We showed that Presenilins are the most ancient of the Gamma secretase proteins and might have their origin in last common eukaryotic ancestor (LCEA). We also demonstrated that these proteins have been evolving under strong purifying selection. Through evolutionary trace analysis, we showed that Presenilin protein sites which undergoes mutations in Familial Alzheimer Disease are highly conserved in metazoans. Finally, we discussed the evolutionary, physiological and pathological implication of our findings and proposed that evolutionary profile of Presenilins supports the loss of function hypothesis of AD pathogenesis.

evolutionary biology

A map of direct TF-DNA interactions in the human genome

Chromatin immunoprecipitation followed by sequencing (ChIP-seq) is the most popular assay to identify genomic regions, called ChIP-seq peaks, that are bound in vivo by transcription factors (TFs). These regions are derived from direct TF-DNA interactions, indirect binding of the TF to the DNA (through a co-binding partner), nonspecific binding to the DNA, and noise/bias/artifacts. Delineating the bona fide direct TF-DNA interactions within the ChIP-seq peaks remains challenging. We developed a dedicated software, ChIP-eat, that combines computational TF binding models and ChIP-seq peaks to automatically predict direct TF-DNA interactions. Our work culminated with predicted interactions covering >4% of the human genome, obtained by uniformly processing 1,983 ChIP-seq peak data sets from the ReMap database for 232 unique TFs. The predictions were a posteriori assessed using protein binding microarray and ChIP-exo data, and were predominantly found in high quality ChIP-seq peaks. The set of predicted direct TF-DNA interactions suggested that high-occupancy target regions are likely not derived from direct binding of the TFs to the DNA. Our predictions derived co-binding TFs supported by protein-protein interaction data and defined cis-regulatory modules enriched for disease- and trait-associated SNPs. Finally, we provide this collection of direct TF-DNA interactions and cis-regulatory modules in the human genome through the UniBind web-interface (http://unibind.uio.no).

bioinformatics

POPSICLE: A Software Suite to Study Population Structure and Ancestral Determinants of Phenotypes using Whole Genome Sequencing Data

The advent of new sequencing technologies has provided access to genome-wide markers which may be evaluated for their association with phenotypes. Recent studies have leveraged these technologies and sequenced hundreds and sometimes thousands of strains to improve the accuracy of genotype-phenotype predictions. Sequencing of thousands of strains is not practical for many research groups which argues for the formulation of new strategies to improve predictability using lower sample sizes and more cost-effective methods. We introduce here a novel computational algorithm called POPSICLE that leverages the local genetic variations to infer blocks of shared ancestries to construct complex evolutionary relationships. These evolutionary relationships are subsequently visualized using chromosome painting, as admixtures and as clades to acquire general as well as specific ancestral relationships within a population. In addition, POPSICLE evaluates the ancestral blocks for their association with phenotypes thereby bridging two powerful methodologies from population genetics and genome-wide association studies. In comparison to existing tools, POPSICLE offers substantial improvements in terms of accuracy, speed and automation. We evaluated POPSICLEs ability to find genetic determinants of Artemisinin resistance within P. falciparum using 57 randomly selected strains, out of 1,612 that were used in the original study. POPSICLE found Kelch, a gene implicated in the original study, to be significant (p-value 0) towards resistance to Artemisinin. We further extended this analysis to find shared ancestries among closely related P. falciparum, P. reichenowi and P. gaboni species from the Laverania subgenus of Plasmodium. POPSICLE was able to accurately infer the population structure of the Laverania subgenus and detected 4 strains from a chimpanzee in Koulamoutou with significant shared ancestries with P. falciparum and P. gaboni. We simulated 4 datasets to asses if these shared ancestries indicated a hybrid or mixed infections involving P. falciparum and P. gaboni. The analysis based on the simulated data and genome-wide heterozygosity profiles of the strains indicate these are most likely mixed infections although the possibility of hybrids cannot be ruled out. POPSICLE is a java-based utility that requires no installation and can be downloaded freely from https://popsicle-admixture.sourceforge.io/\n\nAuthor SummaryThe associations between genotypes and phenotypes have traditionally been performed using markers such as single nucleotide polymorphisms. Often, these markers are independently evaluated for their association with phenotypes. A genomic region is deemed significant if multiple markers with significance colocalize. However, multiple markers that are in linkage disequilibrium can sometimes work synergistically and contribute to phenotypic variations. These synergistic associations across markers and across subpopulations have traditionally been captured by population genetic approaches that determine local ancestries. We sought to bridge these two powerful but independent methodologies to improve genotype-phenotype predictions. We developed a new software called POPSICLE that employs an innovative approach to determine local ancestries and evaluates them for their association with phenotypes. Validity of POPSICLE in determining the genes that are responsible for Plasmodium Falciparums resistance to Artemisinin and in determining the population structure of Laverania subgenus of Plasmodium are discussed.

bioinformatics

Super-enhancers are transcriptionally more active and cell-type-specific than stretch enhancers

BackgroundSuper-enhancers and stretch enhancers represent classes of transcriptional enhancers that have been shown to control the expression of cell identity genes and carry disease- and trait-associated variants. Specifically, super-enhancers are clusters of enhancers defined based on the binding occupancy of master transcription factors (TFs), chromatin regulators, or chromatin marks, while stretch enhancers are large chromatin-defined regulatory regions of at least 3,000 base pairs. Several studies have characterized these regulatory regions in numerous cell types and tissues to decipher their functional importance. However, the differences and similarities between these regulatory regions have not been fully assessed.\n\nResultsWe integrated genomic, epigenomic, and transcriptomic data from ten human cell types to perform a comparative analysis of super and stretch enhancers with respect to their chromatin profiles, cell-type-specificity, and ability to control gene expression. We found that stretch enhancers are more abundant, more distal to transcription start sites, cover twice as much the genome and are significantly less conserved than super-enhancers. In contrast, super-enhancers are significantly more enriched for active chromatin marks and cohesin complex and transcriptionally active than stretch enhancers. Importantly, a vast majority of superenhancers (85%) overlap with only a small subset of stretch enhancers (13%), which are enriched for cell-type-specific biological functions, and control cell identity genes.\n\nConclusionsThese results suggest that super-enhancers are transcriptionally more active and cell-type-specific than stretch enhancers, and importantly, most of the stretch enhancers that are distinct from superenhancers do not show an association with cell identity genes, are less active, and more likely to be poised enhancers.

genomics

VirTect: a computational method for detecting virus species from RNA-Seq and its application in head and neck squamous cell carcinoma

Next generation sequencing (NGS) provides an opportunity to detect viral species from RNA-seq data on human tissues, but existing computational approaches do not perform optimally on clinical samples. We developed a bioinformatics method called VirTect for detecting viruses in neoplastic human tissues using RNA-seq data. Here, we used VirTect to analyze RNA-seq data from 363 HNSCC (head and neck squamous cell carcinoma) patients and identified 22 HPV-induced HNSCCs. These predictions were validated by manual review of pathology reports on histopathologic specimens. Compared to two existing prediction methods, VirusFinder and VirusSeq, VirTect demonstrated superior performance with many fewer false positives and false negatives. The majority of HPV carcinogenesis studies thus far have been performed on cervical cancer and generalized to HNSCC. Our results suggest that HPV-induced HNSCC involves unique mechanisms of carcinogenesis, so understanding these molecular mechanisms will have a significant impact on therapeutic approaches and outcomes. In summary, VirTect can be an effective solution for the detection of viruses with NGS data, and can facilitate the clinicopathologic characterization of various types of cancers with broad applications for oncology.\n\nSignificance StatementWe developed a new bioinformatics tool, and reported the new inside of HPV carcinogenesis mechanism in HPV-induced head and neck squamous cell carcinoma (HNSCC). This novel bioin-formatics tool and the new knowledge of HPV-induced HNSCC will facilitate the development of target therapies for treating HNSCC.

bioinformatics

Role of the novel endoribonuclease SLFN14 and its disease causing mutations in ribosomal degradation

Platelets are anucleate and mostly ribosome-free cells within the bloodstream, derived from megakaryocytes within bone marrow and crucial for cessation of bleeding at sites of injury. Inherited thrombocytopenias are a group of disorders characterized by alow platelet count and are frequently associated with excessive bleeding. SLFN14 is one of the most recently discovered genes linked to inherited thrombocytopenia where several heterozygous missense mutations in SLFN14 were identified to cause defective megakaryocyte maturation and platelet dysfunction. Yet, SLFN14 was recently described as a ribosome-associated protein resulting in rRNA and ribosome-bound mRNA degradation in rabbit reticulocytes. To unveil the cellular function of SLFN14 and the link between SLFN14 and thrombocytopenia, we examined SLFN14 (WT/mutants) in in vitro models. Here, we show that all SLFN14 variants co-localize with ribosomes and mediate rRNA endonucleolytic degradation and ribosome clearance. Compare dto SLFN14 WT, expression of mutants is dramatically reduced as a result of post-translational degradation due to partial misfolding of the protein. Moreover, all SLFN14 variants tend to form oligomers. These findings could explain the dominant negative effect of heterozygous mutation on SLFN14 expression in patients platelets. Overall we suggest that SLFN14 could be involved in ribosome degradation during platelet formation and maturation.

cell biology

Control of septin filament flexibility and bundling

Septins self-assemble into heteromeric rods and filaments to act as scaffolds and modulate membrane properties. How cells tune the biophysical properties of septin filaments to control filament flexibility and length, and in turn the size, shape, and position of higher-order septin structures is not well understood. We examined how rod composition and nucleotide availability influence physical properties of septins such as annealing, fragmentation, bundling and bending. We found that septin complexes have symmetric termini, even when both Shs1 and Cdc11 are coexpressed. The relative proportion of Cdc11/Shs1 septin complexes controls the biophysical properties of filaments and influences the rate of annealing, fragmentation, and filament flexibility. Additionally, the presence and exchange of guanine nucleotide also alters filament length and bundling. An Shs1 mutant that is predicted to alter nucleotide hydrolysis has altered filament length and dynamics in cells and impacts cell morphogenesis. These data show that modulating filament properties through rod composition and nucleotide binding can control formation of septin assemblies that have distinct physical properties and functions.

cell biology

Intragenic differential expression in archaea transcriptomes revealed by computational analysis of tiling microarrays

Recent advances, in high-throughput technologies allows whole transcriptome analysis, providing a complete and panoramic view of intragenic differential expression in eukaryotes. However, intragenic differential expression in prokaryotes still mystery and incompletely understood. In this study, we investigated and collected the evidence for intragenic differential expression in several archaeal transcriptomes such as, Halobacterium salinarum NRC-1, Pyrococcus furiosus, Methanococcus maripaludis, and Sulfolobus solfataricus, based on computational methods; specifically, by well-known self-organizing map (SOM) for cluster analysis, which transforms high dimensional data into low dimensional. We found 104 (3.86%) of genes in Halobacterium salinarum NRC-1, 59 (2.56%) of genes in Pyrococcus furiosus, 43 (2.41%) of genes Methanococcus maripaludis and 13 (0.42%) of genes in Sulfolobus solfataricus have two or more clusters, i.e., showed the intragenic differential expression at different conditions.

bioinformatics

JASPAR RESTful API: accessing JASPAR data from any programming language

JASPAR is a widely used open-access database of curated, non-redundant transcription factor binding profiles. Currently, data from JASPAR can be retrieved as flat files or by using programming language-specific interfaces. Here, we present a programming language-independent application programming interface (API) to access JASPAR data using the Representational State Transfer (REST) architecture. The REST API enables programmatic access to JASPAR by most programming languages and returns data in seven widely used formats. Further, it provides an endpoint to infer the TF binding profile(s) likely bound by a given DNA binding domain protein sequence. Additionaly, it provides an interactive browsable interface for bioinformatics tool developers. The REST API is implemented in Python using the Django REST Framework. It is accessible at http://jaspar.genereg.net/api/ and the source code is freeiy available at https://bitbucket.org/CBGR/jaspar under GPL v3 iicense.

bioinformatics

Assembly Of Whole-Chromosome Pseudomolecules For Polyploid Plant Genomes Using Outcrossed Mapping Populations

The assembly of whole-chromosome pseudomolecules for plant genomes remains challenging due to polyploidy and high repeat content. We developed an approach for constructing complete pseudomolecules for polyploid species using genotyping-by-sequencing data from outcrossing mapping populations coupled with high coverage whole genome sequence data of a reference genome. Our approach combines de novo assembly with linkage mapping to arrange scaffolds into pseudomolecules. We show that the method is able to reconstruct simulated chromosomes for both diploid and tetraploid genomes. Comparisons to three existing genetic mapping tools show that our method outperforms the other methods in accuracy on both grouping and ordering, and is robust to the presence of substantial amounts of missing data and genotyping errors. We applied our method to three real datasets including a diploid Ipomoea trifida and two tetraploid potato mapping populations. The linkage maps show significant concordance with the reference chromosomes. We resolved seven assembly errors for the published Ipomoea trifida genome assembly as well as anchored an unplaced scaffold in the published potato genome.

bioinformatics

Intervene: a tool for intersection and visualization of multiple gene or genomic region sets

BackgroundA common task for scientists relies on comparing lists of genes or genomic regions derived from high-throughput sequencing experiments. While several tools exist to intersect and visualize sets of genes, similar tools dedicated to the visualization of genomic region sets are currently limited.\n\nResultsTo address this gap, we have developed the Intervene tool, which provides an easy and automated interface for the effective intersection and visualization of genomic region or list sets, thus facilitating their analysis and interpretation. Intervene contains three modules: venn to generate Venn diagrams of up to six sets, upset to generate UpSet plots of multiple sets, and pairwise to compute and visualize intersections of multiple sets as clustered heat maps. Intervene, and its interactive web ShinyApp companion, generate publication-quality figures for the interpretation of genomic region and list sets.\n\nConclusionsIntervene and its web application companion provide an easy command line, and an interactive web interface to compute intersections of multiple genomic and list sets. They also have the capacity to plot intersections using easy-to-interpret visual approaches. Intervene is developed and designed to meet the needs of both computer scientists and biologists. The source code is freely available at https://bitbucket.org/CBGR/intervene, with the web application available at https://asntech.shinyapps.io/intervene.

bioinformatics

Analysis and prediction of super-enhancers using sequence and chromatin signatures

BackgroundSuper-enhancers are clusters of transcriptional enhancers densely occupied by the Mediators, transcription factors and chromatin regulators. They control the expression of cell identity genes and disease associated genes. Current studies demonstrated the possibility of multiple factors with important roles in super-enhancer formation; however, a systematic analysis to assess the relative importance of chromatin and sequence signatures of super-enhancers and their constituents remain unclear.\n\nResultsHere, we integrated diverse types of genomic and epigenomic datasets to identify key signatures of super-enhancers and their constituents and to investigate their relative importance. Through computational modelling, we found that Cdk8, Cdk9 and Smad3 as new key features of super-enhancers along with many known features such as H3K27ac. Comprehensive analysis of these features in embryonic stem cells and pro-B cells revealed their role in the super-enhancer formation and cellular identity. We also observed that super-enhancers are significantly GC-rich in contrast with typical enhancers. Further, we observed significant correlation among many cofactors at the constituents of super-enhancers.\n\nConclusionsOur analysis and ranking of super-enhancer signatures can serve as a resource to further characterize and understand the formation of super-enhancers. Our observations are consistent with a cooperative or synergistic model underlying the interaction of super-enhancers and their constituents with numerous factors.

genomics

Canonical and cross-reactive binding of NK cell inhibitory receptors to HLA-C allotypes is dictated by peptides bound to HLA-C

BackgroundHuman natural killer (NK) cell activity is regulated by a family of killer-cell Ig-like receptors (KIR) that bind human leucocyte antigen (HLA) class I. Combinations of KIR and HLA genotypes are associated with disease, including susceptibility to viral infection and disorders of pregnancy. KIR2DL1 binds HLA-C alleles of group C2 (Lys80) and KIR2DL2 and KIR2DL3 bind HLA-C alleles of group C1 (Asn80). However, this model does not capture allelic diversity in HLA-C or the impact of HLA-bound peptides. The goal of this study was to determine the extent to which the endogenous HLA-C peptide repertoire can influence the specific binding of inhibitory KIR to HLA-C allotypes.\n\nResultsThe impact of HLA-C bound peptide on inhibitory KIR binding was investigated taking advantage of the fact that HLA-C*05:01 (HLA-C group 2, C2) and HLA-C*08:02 (HLA-C group 1, C1) have identical sequences apart from the key KIR specificity determining epitope at residues 77 and 80. Endogenous peptides were eluted from HLA-C*05:01 and used to test the peptide dependence of KIR2DL1 and KIR2DL2/3 binding to HLA-C*05:01 and HLA-C*08:02 and subsequent impact on NK cell function. Specific binding of KIR2DL1 to the C2 allotype occurred with the majority of peptides tested. In contrast, KIR2DL2/3 binding to the C1 allotype occurred with only a subset of peptides. Cross-reactive binding of KIR2DL2/3 with the C2 allotype was restricted to even fewer peptides. Unexpectedly, two peptides promoted binding of the C2 allotype-specific KIR2DL1 to the C1 allotype. We showed that presentation of endogenous peptides, or predicted HIV Gag peptides, by HLA-C can promote KIR cross-reactive binding.\n\nConclusionsKIR2DL2/3 binding to C1 is more peptide selective than that of KIR2DL1 binding to C2, which provides an explanation for why KIR2DL3-C1 interactions appear weaker than KIR2DL1-C2. In addition, cross-reactive binding of KIR is characterized by even higher peptide selectivity. We demonstrate a hierarchy of functional peptide selectivity of KIR-HLA-C interactions with relevance to NK cell biology and human disease associations. This selective peptide sequence-driven binding of KIR provides a potential mechanism for pathogen as well as self-peptide to modulate NK cell activation through altering levels of inhibition.

immunology

The molecular mechanism of the type IVa pilus motors

Type IVa pili are protein filaments essential for virulence in many bacterial pathogens; they extend and retract from the surface of bacterial cells to pull the bacteria forward with unprecedented force. They are used for attachment, swarming and twitching motility, biofilm formation, up-regulation of other virulence factors, and natural competence. The pilus is assembled by the motor subcomplex which consists of the inner membrane protein PilC and the cytoplasmic ATPase PilB. How PilB catalyzes this process is unknown, due in part to the lack of high-resolution structural information. Phylogenetic analysis of PilB-like ATPases, including GspE, PilT, BfpD, FlaI, and archaeal GspE2 revealed highly conserved residues essential for function in this family of ATPases. Here we report the structure of the core ATPase domains of Geobacter metalloreducens PilB bound to ADP and the non-hydrolysable ATP analogue, AMPPNP, at 3.4 and 2.3[A], respectively. Importantly, these structures were determined in non-saturating nucleotide conditions, revealing important differences in nucleotide binding between chains. Analysis of these differences revealed the sequential turnover of nucleotide by the chains, and the corresponding domain movements. Our data indicate a clockwise rotation of movement in PilB, which would support the assembly of a right-handed helical pilus. Conversely, our analysis suggests a counterclockwise rotation in PilT that would enable right-handed pilus disassembly. The proposed model provides insight into how this family of ATPases can power pilus extension and retraction with extraordinary forces.

microbiology