Search bioRxiv⌕ Search

Biology subjects

Lian, D.

Publications and source records attributed to Lian, D..

6 recordsLinked to original sources

A long-read human pangenome initiative for comprehensive interpretation of nuclear-embedded mitochondrial DNA

Nuclear-embedded mitochondrial DNA segments (NUMT) preserve a record of ongoing mitochondrial-to-nuclear DNA transfer during evolution, with important implications for disease mechanisms and genome organization. Here, we developed a Pangenome Graph-based NUMT Detection (PG-NUMT) approach to improve NUMT detection sensitivity by 2.52-fold compared to short-read based approaches and comprehensively resolved seven concatenated mega-NUMTs with a maximum length of 127.7 kbp. We generated a high-resolution human NUMT map comprising 774 fixed, 280 polymorphic, and 123 pericentromeric NUMTs, alongside 74 superpopulation-stratified NUMTs. Both fixed and polymorphic NUMTs are hypermethylated and enriched within segmental duplications, whereas fixed NUMTs preferentially localize to intergenic regions and avoid transposable elements. Interestingly, NUMTs derived from the 3-end of the mtDNA D-loop are less frequently fixed in human genomes and exhibit potential cis-regulatory activity in the nuclear genome. These observations suggest that selective pressures shape the genomic features of fixed NUMTs. Seven NUMTs associated with gene expression or alternative splicing were further identified, suggesting potential modulatory functions of common NUMTs. Using 20 complete non-human primate genomes, we identified variable NUMT insertion rates across primate lineages (1.5-18.7 insertions per million years), with particularly high rates in the Pan lineage. Notably, we uncovered two NUMT-derived variable number tandem repeats (VNTRs), establishing NUMTs as a novel source of VNTRs. In summary, the integrated analysis enhances our understanding of NUMT genomic architecture, population dynamics, and evolutionary implications in human and non-human primates, establishing NUMTs as dynamic genomic components of biomedical relevance.

genomics↗

Joint Segmental Duplication Co-option Drives Human-specific Transcriptional Readthrough and Expression Fine-tuning of NPEPPS-TBC1D3

Segmental duplication (SD) is a major driver of functional changes in evolution and disease. Many genes embedded within SDs, such as NPEPPS and TBC1D3, display substantial copy number variation (CNV) across individuals. Yet, the precise identification of the functional copies and their transcriptional outputs remains largely unstudied. Focusing on NPEPPS and TBC1D3, we illustrate human-specific expression fine-tuning mechanisms associated with readthrough transcripts. We identified a human-specific NPEPPS-TBC1D3 digenic genomic structure that originated from a joint SD pair and became fixed across populations. Experiments demonstrate that this structure generates NPEPPS-TBC1D3 readthrough transcripts, which are the predominant isoforms of TBC1D3 expression in various cell types, fine-tuning its protein level. Furthermore, a human-specific hypomethylation signal within an upstream CpG island of NPEPPS precisely pinpoints the expressed TBC1D3 paralog. Moreover, we reveal transcriptional readthrough events are [~]3-fold enriched for joint-SD-associated transcriptional readthrough (JSDTR) and identify 109 JSDTR gene pairs, including neurodevelopmentally important pairs and clinically interesting SERF1A/B-SMN1/2. Taken together, our findings comprehensively describe an example of how a joint SD event shaped evolution and suggest that JSDTR is a broad mechanism for the emergence of new functions.

genetics↗

The Polymorphism of Globin and Virus RdRP and Protein Space

The divergence of protein sequences, structures, and functions reveals Natures exploration of evolutionary possibilities within a given structural fold. Globins are ubiquitous across bacteria, fungi, plants, and animals; RNA-dependent RNA polymerases (RdRPs) are responsible for the replication of RNA viruses. These proteins provide excellent models for the study of protein evolution. In this study we analyzed the polymorhpisms of globins and viruses RdRPs with structure and function. We found that the protein polymorphisms are positively correlated with residue solvent accessibility (RSA). From our analysis of protein space of globins and virus RdRPs, we proposed that purifying selection is the predominant force of protein evolution. We also discuss the theory of last universal common ancestor (LUCA), proposed that it may be put to test.

bioinformatics↗

Polymorphisms, Solvent Accessibility, and Evolutionary Conservation of Influenza A Virus PB1 Protein

Protein polymorphisms, reflecting amino acid and nucleotide sequence divergence, provide insights into protein evolution. Here, we analyze sequence variation and structural features of influenza A virus (IAV) PB1 protein, an RNA-dependent RNA polymerase critical for viral replication. Our findings demonstrate that residue solvent accessibility strongly predicts polymorphism likelihood, with exposed sites exhibiting higher variability. Despite extensive polymorphism, we observe pervasive purifying selection across PB1, maintaining functional and structural constraints. These results highlight how protein architecture shapes evolutionary dynamics in viral proteins.

bioinformatics↗

The Polymorphisms, Solvent Accessibility and Conservatism of Hepatitis C Virus Nonstructural 5B Protein

ABSTRACTThe polymorphisms of protein or protein family, that is, the divergences of amino acid and nucleotide sequences have provided much useful information on the divergent evolution of proteins. In this paper, we analyzed the polymorphisms of enzyme NS5B of HCV for which sequence variation among most isolates have been characterized and protein structures of the catalytic domain form of this enzyme are also known. For this protein, we found that solvent accessibility of residues in the protein structure is a strong predictor of whether or not an amino acid will be polymorphic and the residue variability. Apart from polymorphism, we found conservatism at every level among site is universal for this protein. We also found that purifying selection at different levels was strong in the forming of the polymorphisms and conservatism of this protein.

bioinformatics↗

Prediction of Carbon Emissions in Guizhou Province-Based on Different Neural Network Models

Global warming caused by greenhouse gas emissions has become a major challenge facing people all over the world. The study of regional human activities and their impacts on carbon emissions is of great significance to achieve the ambitious goal of carbon neutrality and sustainable economic development. Guizhou Province is a typical karst area in China, and its energy consumption is mainly based on fossil fuels.Therefore, it is necessary to predict and analyze its carbon emissions. In this paper, BP neural network and extreme learning machine (ELM) model, which have the advantage of nonlinear processing, will be used to predict the carbon emissions of Guizhou Province from 2020 to 2040. Based on the energy consumption data of Guizhou Province, the carbon emissions of Guizhou Province are calculated by using the conversion method and the inventory compilation method. The data show that the carbon emissions of Guizhou Province show an "S" growth trend; In this paper, 12 influencing factors of carbon emissions are selected, and five influencing factors with larger correlation are screened out by using grey correlation analysis method, and the prediction model of carbon emissions in Guizhou Province is established and simulated, and the prediction performance of BP neural network, ELM and WOA-ELM are compared respectively. Compared with ELM model and BP neural network model, the prediction accuracy of WOA-ELM model is higher; Finally, three development scenarios of carbon emissions are set by scenario analysis, which are baseline scenario, high-speed scenario and low-carbon scenario. On this basis, the size and time of peak carbon emissions in Liaoning Province from 2020 to 2040 are predicted based on WOA-ELM model. The results show that the peak value of carbon dioxide in the low carbon scenario is up to 0.98 million tons 31294 in 2033, the peak value of carbon emissions in the high speed scenario is up to 0.37 million tons 30251 in 2036, and the peak value of carbon emissions in the baseline scenario is up to 0.61 million tons 26243 in 2038. Based on the peak time and prediction results of carbon emissions under the three scenarios, the main factors contributing to the reduction of carbon emissions in Guizhou Province are analyzed, and important data basis is provided for energy conservation and emission reduction in Guizhou Province.

ecology↗