Search bioRxiv⌕ Search

Biology subjects

Ming, W.

Publications and source records attributed to Ming, W..

3 recordsLinked to original sources

abCAN: a Practical and Novel Attention Network for Predicting Mutant Antibody Affinity

Accurate prediction of mutation effects on antibody-antigen interactions is critical for antibody engineering and drug design. In this study, we present abCAN, a practical and novel attention network designed to predict changes in binding affinity caused by mutations. abCAN requires only the pre-mutant antibody-antigen complex structure and mutation information to perform its predictions. abCAN introduces an innovative approach, Progressive Encoding, which progressively integrates structural, residue-level, and sequential information to construct the complex representation in a systematic manner, effectively capturing both the topological features of the structure and contextual features of the sequence. During which, extra weight to interface residues would also be applied through attention mechanisms. These learned representations are then transferred to a predictor that estimates changes in antibody-antigen binding affinity induced by mutations. On the benchmark dataset, abCAN achieved a root-mean-square error (RMSE) of 1.195 (kcal/mol-1) and a Pearson correlation coefficient (PCC) of 0.841, setting a new state-of-the-art (SOTA) benchmark for prediction accuracy in the field of antibody affinity prediction.

bioinformatics↗

Investigation of contributions from cortical and subcortical brain structures for speech decoding

Language impairments often arise from severe neurological disorders, prompting the development of neural prosthetics based on electrophysiological signals for the restoration of comprehensible language information. Previous decoding efforts have focused mainly on signals from the cerebral cortex, neglecting the potential contributions of subcortical brain structures to speech decoding in brain-computer interfaces (BCIs). This study aims to explore the role of subcortical structures for speech decoding by utilizing stereotactic electroencephalography (sEEG). Two native Mandarin Chinese speakers, who underwent sEEG implantation for pharmaco-resistant epilepsy, participated in this study. sEEG contacts were primarily located in the superior temporal gyrus, middle temporal gyrus, inferior temporal gyrus, thalamus, hippocampus, insular gyrus, amygdala, and parahippocampal gyrus. The participants were asked to read Chinese text, which included 407 Chinese characters (covering all Chinese syllables), displayed on a screen after receiving prompts. 1-30, 30-70 and 70-150 Hz frequency band powers of sEEG signals were used as key features. A deep learning model based on long short-term memory (LSTM) was developed to evaluate the contribution of different brain structures during encoding of speech. Prediction of speech characteristics of consonants (articulatory place and manner) and tone within single words based on the selected features and electrode contact locations was made. Cortical signals were generally better at articulatory place prediction (86.5% accuracy, chance level = 12.5%), while cortical and subcortical signals predicted articulatory manner at similar level (51.5% vs 51.7% accuracy, respectively, chance level = 14.3%). Subcortical signals generated better prediction for tone (around 58.3% accuracy, chance level = 25%). Superior temporal gyrus remains highly relevant during speech decoding for both consonants and tone. Prediction reached the highest level when cortical and subcortical inputs were combined, especially for tone prediction. Our findings indicate that both cortical and subcortical structures can play crucial roles for speech decoding, each contributing to different aspects of speech.

neuroscience↗

Detecting Full-Length EccDNA with FLED and long-reads sequencing

Reconstructing the full-length sequence of extrachromosomal circular DNA (eccDNA) from short sequencing reads has proved challenging given the similarity of eccDNAs and their corresponding linear DNAs. Previous sequencing methods were unable to achieve high-throughput detection of full-length eccDNAs. Here we describe a new strategy that combined rolling circle amplification (RCA) and nanopore long-reads sequencing technology to generate full-length eccDNAs. We further developed a novel algorithm, called Full-Length eccDNA Detection (FLED), to reconstruct the sequence of eccDNAs. We used FLED to analyze seven human epithelial and cancer cell line samples and identified over 5,000 full-length eccDNAs per sample. The structures of identified eccDNAs were validated by both PCR and Sanger sequencing. Compared to other published nanopore-based eccDNA detectors, FLED exhibited higher sensitivity. In cancer cell lines, the genes overlapped with eccDNA regions were enriched in cancer-related pathways and cis-regulatory elements can be predicted in the up-stream or downstream of intact genes on eccDNA molecules, and the expressions of these cancer-related genes were dysregulated in tumor cell lines, indicating the regulatory potency of eccDNAs in biological processes. Our method takes advantage of nanopore long reads and enables unbiased reconstruction of full-length eccDNA sequences. FLED is imple-mented using Python3 which is freely available on GitHub (https://github.com/FuyuLi/FLED).

genomics↗