Search bioRxivSearch

Biology subjects

Fan, F.

Publications and source records attributed to Fan, F..

4 recordsLinked to original sources

The Utilization of Pharmacophore-based 3D QSAR Modeling and Virtual Screening in Safety Profiling: a Case Study to Identify Antagonistic Activities against Adenosince Receptor, A2aR, using 1,897 known drugs

ABSTRACTSafety pharmacology screening against a wide range of unintended vital targets using in vitro assays is crucial to understand off-target interactions with drug candidates. With the increasing demand for in vitro assays, ligand-and structure-based virtual screening approaches have been evaluated for potential utilization in safety profiling. Although ligand based approaches have been actively applied in retrospective analysis or prospectively within well-defined chemical space during the early discovery stage (i.e., HTS screening and lead optimization), virtual screening is rarely implemented in later stage of drug discovery (i.e., safety). Here we present a case study to evaluate ligand-based 3D QSAR models built based on in vitro antagonistic activity data against adenosine receptor 2A (A2aR). The resulting models, obtained from 268 chemically diverse compounds, were used to test a set of 1,897 chemically distinct drugs, simulating the real-world challenge of safety screening when presented with novel chemistry and a limited training set. Due to the unique requirements of safety screening versus discovery screening, the limitations of 3D QSAR methods (i.e., chemotypes, dependence on large training set, and prone to false positives) are less critical than early discovery screen. We demonstrated that 3D QSAR modelling can be effectively applied in safety assessment prior to in vitro assays, even with chemotypes that are drastically different from training compounds. It is also worth noting that our model is able to adequately make the mechanistic distinction between agonists and antagonists, which is important to inform subsequent in vivo studies. Overall, we present an in-depth analysis of the appropriate utilization and interpretation of pharmacophore-based 3D QSAR models for safety screening.

pharmacology and toxicology

Single tube bead-based DNA co-barcoding for cost effective and accurate sequencing, haplotyping, and assembly

Obtaining accurate sequences from long DNA molecules is very important for genome assembly and other applications. Here we describe single tube long fragment read (stLFR), a technology that enables this a low cost. It is based on adding the same barcode sequence to sub-fragments of the original long DNA molecule (DNA co-barcoding). To achieve this efficiently, stLFR uses the surface of microbeads to create millions of miniaturized barcoding reactions in a single tube. Using a combinatorial process up to 3.6 billion unique barcode sequences were generated on beads, enabling practically non-redundant co-barcoding with 50 million barcodes per sample. Using stLFR, we demonstrate efficient unique co-barcoding of over 8 million 20-300 kb genomic DNA fragments. Analysis of the genome of the human genome NA12878 with stLFR demonstrated high quality variant calling and phasing into contigs up to N50 34 Mb. We also demonstrate detection of complex structural variants and complete diploid de novo assembly of NA12878. These analyses were all performed using single stLFR libraries and their construction did not significantly add to the time or cost of whole genome sequencing (WGS) library preparation. stLFR represents an easily automatable solution that enables high quality sequencing, phasing, SV detection, scaffolding, cost-effective diploid de novo genome assembly, and other long DNA sequencing applications.

genomics

The Origins and Consequences of Localized and Global Somatic Hypermutation

Cancer is a disease of the genome, but the dramatic inter-patient variability in mutation number is poorly understood. Tumours of the same type can differ by orders of magnitude in their mutation rate. To understand potential drivers and consequences of the underlying heterogeneity in mutation rate across tumours, we evaluated both local and global measures of mutation density: both single-stranded and double-stranded DNA breaks in 2,460 tumours of 38 cancer types. We find that SCNAs in thousands of genes are associated with elevated rates of point-mutations, while similarly point-mutation patterns in dozens of genes are associated with specific patterns of DNA double-stranded breaks. These candidate drivers of mutation density are enriched for known cancer drivers, and preferentially occur early in tumour evolution, appearing clonally in all cells of a tumour. To supplement this understanding of global mutation density, we developed and validated a tool called SeqKat to identify localized \"rainstorms\" of point-mutations (kataegis). We show that rates of kataegis differ by four orders of magnitude across tumour types, with malignant lymphomas showing the highest. Tumours with TP53 mutations were 2.6-times more likely to harbour a kataegic event than those without, and 239 SCNAs were associated with elevated rates of kataegis, including loss of the tumour-suppressor CDKN2A. We identify novel subtypes of kataegic events not associated with aberrant APOBEC activity, and find that these are localized to specific cellular regions, enriched for MYC-target genes. Kataegic events were associated with patient survival in some, but not all tumour types, highlighting a combination of global and tumour-type specific effects. Taken together, we reveal a landscape of genes driving localized and tumour-specific hyper-mutation, and reveal novel mutational processes at play in specific tumour types.

cancer biology

A Convenient Non-harm Cervical Spondylosis Intelligent Identity method based on Machine Learning

Cervical spondylosis(CS), a most common orthopedic diseases, is mainly identified by the doctors judgment from the clinical symptoms and cervical change provided by expensive instruments in hospital. Owing to the development of the surface electromyography(sEMG) technique and artificial intelligence, we proposed a convenient non-harm CS intelligent identify method EasiCNCSII, including the sEMG data acquisition and the CS identification. For the convenience and efficiency of data acquisition with the limited testable muscles provided by the sEMG technology, we proposed a data acquisition method based on the relationship between muscle activity pattern, the tendons theory and CS etiology. It is easily performed in less than 20 minutes, even outside the hospital. Faced with the challenge of high-dimension and the weak availability, the 3-tier model EasiAI is developed to intelligently identify CS. The common features and new features are extracted from raw sEMG data in first tier. The EasiRF is proposed in second tier to further reduce the data dimension and improve the performance. With the limited and weakly available data, the gradient boosted regression tree is developed in third tier to effectively identify CS. The EasiAI achieve the best performance with 91.02% in accuracy, 97.14% in sensitivity, and 81.43% in specificity compared with 4 common machine learning classification model, validating the EasiCNCSII effectiveness.

bioinformatics