Search bioRxivSearch

Biology subjects

Fang, L.

Publications and source records attributed to Fang, L..

6 recordsLinked to original sources

Generation of human neural retina transcriptome atlas by single cell RNA sequencing

The retina is a highly specialized neural tissue that senses light and initiates image processing. Although the functional organisation of specific cells within the retina has been well-studied, the molecular profile of many cell types remains unclear in humans. To comprehensively profile cell types in the human retina, we performed single cell RNA-sequencing on 20,009 cells obtained post-mortem from three donors and compiled a reference transcriptome atlas. Using unsupervised clustering analysis, we identified 18 transcriptionally distinct clusters representing all known retinal cells: rod photoreceptors, cone photoreceptors, Muller glia cells, bipolar cells, amacrine cells, retinal ganglion cells, horizontal cells, retinal astrocytes and microglia. Notably, our data captured molecular profiles for healthy and early degenerating rod photoreceptors, and revealed a novel role of MALAT1 in putative rod degeneration. We also demonstrated the use of this retina transcriptome atlas to benchmark pluripotent stem cell-derived cone photoreceptors and an adult Muller glia cell line. This work provides an important reference with unprecedented insights into the transcriptional landscape of human retinal cells, which is fundamental to our understanding of retinal biology and disease.

systems biology

LinkedSV: Detection of mosaic structural variants from linked-read exome and genome sequencing data

Linked-read sequencing provides long-range information on short-read sequencing data by barcoding reads originating from the same DNA molecule, and can improve the detection and breakpoint identification for structural variants (SVs). We present LinkedSV for SV detection on linked-read sequencing data. LinkedSV considers barcode overlapping and enriched fragment endpoints as signals to detect large SVs, while it leverages read depth, paired-end signals and local assembly to detect small SVs. Benchmarking studies demonstrates that LinkedSV outperforms existing tools, especially on exome data and on somatic SVs with low variant allele frequencies. We demonstrate clinical cases where LinkedSV identifies disease causal SVs from linked-read exome sequencing data missed by conventional exome sequencing, and show examples where LinkedSV identifies SVs missed by high-coverage long-read sequencing. In summary, LinkedSV can detect SVs missed by conventional short-read and long-read sequencing approaches, and may resolve negative cases from clinical genome/exome sequencing studies.

bioinformatics

A simple cloning-free method to efficiently induce gene expression using CRISPR/Cas9

Gain-of-function studies often require the tedious cloning of transgene cDNA into vectors for overexpression beyond the physiological expression levels. The rapid development of CRISPR/Cas technology presents promising opportunities to address these issues. Here we report a simple, cloning-free method to induce gene expression at endogenous locus using CRISPR/Cas9 activators. Our strategy utilises synthesized sgRNA expression cassettes to direct a nuclease-null Cas9 complex fused with transcriptional activators (VP64, p65 and Rta) for site-specific induction of endogenous genes. This strategy allows rapid initiation of gain-of-function studies in the same day. Using this cloning-free approach, we tested two CRISPR activation systems, dSpCas9VPR and dSaCas9VPR, for induction of multiple genes in human and rat cells. Our results showed that both CRISPR activators allow efficient induction of six different neural development genes (CRX, RORB, RAX, OTX2, ASCL1 and NEUROD1) in human cells, whereas the rat cells exhibit a more variable and less efficient levels of gene induction, as observed in three different genes (Ascl1, Neurod1, Nrl). Altogether, this study provides a simple method to efficiently activate endogenous gene expression using CRISPR/Cas9 activators, which can be applies as a rapid workflow to initiate gain-of-function studies for a range of molecular and cell biology disciplines.

molecular biology

PIRD: Pan immune repertoire database

MotivationT and B cell receptors (TCRs and BCRs) play a pivotal role in the adaptive immune system by recognizing an enormous variety of external and internal antigens. Understanding these receptors is critical for exploring the process of immunoreaction and exploiting potential applications in immunotherapy and antibody drug design. Although a large number of samples have had their TCR and BCR repertoires sequenced using high-throughput sequencing in recent years, very few databases have been constructed to store these kinds of data. To resolve this issue, we developed a database.\n\nResultsWe developed a database, the Pan Immune Repertoire Database (PIRD), located in China National GeneBank (CNGBdb), to collect and store annotated TCR and BCR sequencing data, including from Homo sapiens and other species. In addition to data storage, PIRD also provides functions of data visualisation and interactive online analysis. Additionally, a manually curated database of TCRs and BCRs targeting known antigens (TBAdb) was also deposited in PIRD.\n\nAvailability and ImplementationPIRD can be freely accessed at https://db.cngb.org/pird.

immunology

Single-molecule optical mapping enables accurate molecular diagnosis of facioscapulohumeral muscular dystrophy (FSHD)

Facioscapulohumeral Muscular Dystrophy (FSHD) is a common adult muscular dystrophy in which the muscles of the face, shoulder blades and upper arms are among the most affected. FSHD is the only disease in which \"junk\" DNA is reactivated to cause disease, and the only known repeat array-related disease where fewer repeats cause disease. More than 95% of FSHD cases are associated with copy number loss of a 3.3kb tandem repeat (D4Z4 repeat) at the subtelomeric chromosomal region 4q35, of which the pathogenic allele contains less than 10 repeats and has a specific genomic configuration called 4qA. Currently, genetic diagnosis of FSHD requires pulsed-field gel electrophoresis followed by Southern blot, which is labor-intensive, semi-quantitative and requires long turnaround time. Here, we developed a novel approach for genetic diagnosis of FSHD, by leveraging Bionano Saphyr single-molecule optical mapping platform. Using a bioinformatics pipeline developed for this assay, we found that the method gives direct quantitative measurement of repeat numbers, can differentiate 4q35 and the highly paralogous 10q26 regions, can determine the 4qA/4qB allelic configuration, and can quantitate levels of post-zygotic mosaicism. We evaluated this approach on 5 patients (including two with post-zygotic mosaicism) and 2 patients (including one with post-zygotic mosaicism) from two separate cohorts, and had complete concordance with Southern blots, but with improved quantification of repeat numbers resolved between haplotypes. We concluded that single-molecule optical mapping is a viable approach for molecular diagnosis of FSHD and may be applied in clinical diagnostic settings once more validations are performed.

genomics

Evaluation on Efficient Detection of Structural Variants at Low Coverage by Long-Read Sequencing

BackgroundStructural variants (SVs) in human genomes are implicated in a variety of human diseases. Long-read sequencing delivers much longer read lengths than short-read sequencing and may greatly improve SV detection. However, due to the relatively high cost of long-read sequencing, it is unclear what coverage is needed and how to optimally use the aligners and SV callers.\n\nResultsIn this study, we developed NextSV, a meta-caller to perform SV calling from low coverage long-read sequencing data. NextSV integrates three aligners and three SV callers and generates two integrated call sets (sensitive/stringent) for different analysis purposes. We evaluated SV calling performance of NextSV under different PacBio coverages on two personal genomes, NA12878 and HX1. Our results showed that, compared with running any single SV caller, NextSV stringent call set had higher precision and balanced accuracy (F1 score) while NextSV sensitive call set had a higher recall. At 10X coverage, the recall of NextSV sensitive call set was 93.5% to 94.1% for deletions and 87.9% to 93.2% for insertions, indicating that ~10X coverage might be an optimal coverage to use in practice, considering the balance between the sequencing costs and the recall rates. We further evaluated the Mendelian errors on an Ashkenazi Jewish trio dataset.\n\nConclusionsOur results provide useful guidelines for SV detection from low coverage whole-genome PacBio data and we expect that NextSV will facilitate the analysis of SVs on long-read sequencing data.

genomics