Search bioRxivSearch

Biology subjects

Sinha, A.

Publications and source records attributed to Sinha, A..

6 recordsLinked to original sources

Complete Genome Sequence of the Wolbachia wAlbB Endosymbiont of Aedes albopictus

Wolbachia, an alpha-proteobacterium closely related to Rickettsia is a maternally transmitted, intracellular symbiont of arthropods and nematodes. Aedes albopictus mosquitoes are naturally infected with Wolbachia strains wAlbA and wAlbB. Cell line Aa23 established from Ae. albopictus embryos retains only wAlbB and is a key model to study host-endosymbiont interactions. We have assembled the complete circular genome of wAlbB from the Aa23 cell line using long-read PacBio sequencing at 500X median coverage. The assembled circular chromosome is 1.48 megabases in size, an increase of more than 300 kb over the published draft wAlbB genome. The annotation of the genome identified 1,205 protein coding genes, 34 tRNA, 3 rRNA, 1 tmRNA and 3 other ncRNA loci. The long reads enabled sequencing over complex repeat regions which are difficult to resolve with short-read sequencing. Thirteen percent of the genome is comprised of IS elements distributed throughout the genome, some of which cause pseudogenization. Prophage WO genes encoding some essential components of phage particle assembly are missing, while the remainder are scattered around the genome. Orthology analysis identified a core proteome of 536 orthogroups across all completed Wolbachia genomes. The majority of proteins could be annotated using Pfam and eggNOG analyses, including ankyrins and components of the T4SS. KEGG analysis revealed the absence of 5 genes in wAlbB which are present in other Wolbachia. The availability of a complete circular chromosome from wAlbB will enable further biochemical, molecular and genetic analyses on this strain and related Wolbachia.\n\nData depositionRaw data from PacBio sequencing have been deposited in the NCBI SRA database under BioProject accession number PRJNA454708, as runs SRR7784284, SRR7784285, SRR7784286, SRR7784287. The paired-end reads from Illumina library used for indel correction are available from NCBI SRA database as accession SRR7623731. The assembled genome and annotations have been submitted to the NCBI GenBank database with accession number CP031221.

genomics

Integrative analysis of single cell expression data reveals distinct regulatory states in bidirectional promoters

BackgroundBidirectional promoters (BPs) are prevalent in eukaryotic genomes. However, it is poorly understood how the cell integrates different epigenomic information, such as transcription factor (TF) binding and chromatin marks, to drive gene expression at BPs. Single cell sequencing technologies are revolutionizing the field of genome biology. Therefore, this study focuses on the integration of single cell RNA-seq data with bulk ChIP-seq and other epigenetics data, for which single cell technologies are not yet established, in the context of BPs.\n\nResultsWe performed integrative analyses of novel human single cell RNA-seq (scRNA-seq) data with bulk ChIP-seq and other epigenetics data. scRNA-seq data revealed distinct transcription states of BPs that were previously not recognized. We find associations between these transcription states to distinct patterns in structural gene features, DNA accessibility, histone modification, DNA methylation and TF binding profiles.\n\nConclusionsOur results suggest that a complex interplay of all of these elements is required to achieve BP-specific transcriptional output in this specialized promoter configuration. Further, our study implies that novel statistical methods can be developed to deconvolute masked subpopulations of cells measured with different bulk epigenomic assays using scRNA-seq data.

genomics

Barcoded oligonucleotides ligated on RNA amplified for multiplex and parallel in-situ analyses

We present Barcoded Oligonucleotides Ligated On RNA Amplified for Multiplexed and parallel In-Situ analysis (BOLORAMIS), a reverse-transcription (RT)-free method for spatially-resolved, targeted, in-situ RNA identification of single or multiple targets. For this proof of concept, we have profiled 154 distinct coding and small non-coding transcripts ranging in sizes 18 nucleotides in length and upwards, from over 200, 000 individual human induced pluripotent stem cells (iPSC) and demonstrated compatibility with multiplexed detection, enabled by fluorescent in-situ sequencing. We use BOLORAMIS data to identify differences in spatial localization and cell-to-cell expression heterogeneity. Our results demonstrate BOLORAMIS to be a generalizable toolset for targeted, in-situ detection of coding and small non-coding RNA for single or multiplexed applications.

cell biology

The Delirium and Population Health Informatics Cohort study protocol: ascertaining the determinants and outcomes from delirium in a whole population.

BackgroundDelirium affects 25% of older inpatients and is associated with long-term cognitive impairment and future dementia. However, no population studies have systematically ascertained cognitive function before, cognitive deficits during, and cognitive impairment after delirium. Therefore, there is a need to address the following question: does delirium, and its features (including severity, duration, and presumed aetiologies), predict long-term cognitive impairment, independent of cognitive impairment at baseline?\n\nMethodsThe Delirium and Population Health Informatics Cohort (DELPHIC) study is an observational population-based cohort study based in the London Borough of Camden. It is recruiting 2000 individuals aged [≥]70 years and prospectively following them for two years, including daily ascertainment of all inpatient episodes for delirium. Daily inpatient assessments include the Memorial Delirium Assessment Scale, the Observational Scale for Level of Arousal, and the Hierarchical Assessment of Balance and Mobility. Data on delirium aetiology is also collected. The primary outcome is the change in the modified Telephone Interview for Cognitive Status at two years.\n\nDiscussionDELPHIC is the first population sample to assess older persons before, during and after hospitalisation. The cumulative incidence of delirium in the general population aged [≥]70 will be described. DELPHIC offers the opportunity to quantify the impact of delirium on cognitive and functional outcomes. Overall, DELPHIC will provide a real-time public health observatory whereby information from primary, secondary, intermediate and social care can be integrated to understand how acute illness is linked to health and social care outcomes.

epidemiology

normR: Regime enrichment calling for ChIP-seq data

ChIP-seq probes genome-wide localization of DNA-associated proteins. To mitigate technical biases ChIP-seq read densities are normalized to read densities obtained by a control. Our statistical framework \"normR\" achieves a sensitive normalization by accounting for the effect of putative protein-bound regions on the overall read statistics. Here, we demonstrate normRs suitability in three studies: (i) calling enrichment for high (H3K4me3) and low (H3K36me3) signal-to-ratio data; (ii) identifying two previously undescribed H3K27me3 and H3K9me3 heterochromatic regimes of broad and peak enrichment; and (iii) calling differential H3K4me3 or H3K27me3-enrichment between HepG2 hepatocarcinoma cells and primary human Hepatocytes. normR is readily available on http://bioconductor.org/packages/normr

bioinformatics

Combining transcription factor binding affinities with open-chromatin data for accurate gene expression prediction

The binding and contribution of transcription factors (TF) to cell specific gene expression is often deduced from open-chromatin measurements to avoid costly TF ChIP-seq assays. Thus, it is important to develop computational methods for accurate TF binding prediction in open-chromatin regions (OCRs). Here, we report a novel segmentation-based method, TEPIC, to predict TF binding by combining sets of OCRs with position weight matrices. TEPIC can be applied to various open-chromatin data, e.g. DNaseI-seq and NOMe-seq. Additionally, Histone-Marks (HMs) can be used to identify candidate TF binding sites. TEPIC computes TF affinities and uses open-chromatin/HM signal intensity as quantitative measures of TF binding strength. Using machine learning, we find low affinity binding sites to improve our ability to explain gene expression variability compared to the standard presence/absence classification of binding sites. Further, we show that both footprints and peaks capture essential TF binding events and lead to a good prediction performance. In our application, gene-based scores computed by TEPIC with one open-chromatin assay nearly reach the quality of several TF ChIP-seq datasets. Finally, these scores correctly predict known transcriptional regulators as illustrated by the application to novel DNaseI-seq and NOMe-seq data for primary human hepatocytes and CD4+ T-cells, respectively.

bioinformatics