Search bioRxivSearch

Biology subjects

Lu, P.

Publications and source records attributed to Lu, P..

6 recordsLinked to original sources

In-cell structural analysis reveals a distinctive chloroplast ribosome in Chlamydomonas reinhardtii

Chloroplast ribosomes synthesize plastid-encoded components of photosynthetic machinery, yet their structure and organization remain poorly understood. We combined cryo-focused ion beam milling, cryo-electron tomography and subtomogram averaging to determine native chloroplast ribosomes in Chlamydomonas reinhardtii. The 4.4-4.9 [A] structure revealed a large arch-like extension on the small subunit (SSU). Comparisons with bacterial and plant chloroplast ribosomes, supported by proteomics, AlphaFold3 predictions and a recent atomic model, indicate that the arch is formed by insertions and extensions in SSU proteins. Classification resolved active, thylakoid-associated ribosomes with density adjacent to the nascent peptide exit and an arch-moved state enriched among thylakoid-associated particles, with coordinated displacement of the arch and beak. Phylogenetic analysis revealed an evolutionary mosaic: the uS3c insertion is broadly distributed across Chlorophyceae, whereas the uS2c insertion, uS5c and PSRP7 are concentrated in Chlamydomonadales, with PSRP7 also in Sphaeropleales. Nuclear-encoded components were recruited stepwise onto a plastid-encoded scaffold, with all four under comparable purifying selection. These findings link a lineage-specific SSU extension to ribosome dynamics, thylakoid association and evolution, highlighting the value of in-cell structural analysis.

plant biology

Spatiotemporal Gene Coexpression and Regulation in Mouse Cardiomyocytes of Early Cardiac Morphogenesis

Cardiac looping is an early morphogenic process critical for the formation of four-chambered mammalian hearts. To study the roles of signaling pathways, transcription factors (TFs) and genetic networks in the process, we constructed gene co-expression networks and identified gene modules highly activated in individual cardiomyocytes (CMs) at multiple anatomical regions and developmental stages. Function analyses of the module genes uncovered major pathways important for spatiotemporal CM differentiation. Interestingly, about half of the pathways were highly active in cardiomyocytes at outflow tract (OFT) and atrioventricular canal (AVC), including many well-known signaling pathways for cardiac development and several newly identified ones. Most of the OFT-AVC pathways were predicted to be regulated by 6 6 transcription factors (TFs) actively expressed at the OFT-AVC locations, with the prediction supported by motif enrichment analysis of the TF targets, including 10 TFs that have not been previously associated with cardiac development, e.g., Etv5, Rbpms, and Baz2b. Finally, our study showed that the OFT-AVC TF targets were significantly enriched with genes associated with mouse heart developmental abnormalities and human congenital heart defects.

genetics

Reproducible evaluation of classification methods in Alzheimer’s disease: framework and application to MRI and PET data

A large number of papers have introduced novel machine learning and feature extraction methods for automatic classification of Alzheimers disease (AD). However, while the vast majority of these works use the public dataset ADNI for evaluation, they are difficult to reproduce because different key components of the validation are often not readily available. These components include selected participants and input data, image preprocessing and cross-validation procedures. The performance of the different approaches is also difficult to compare objectively. In particular, it is often difficult to assess which part of the method (e.g. preprocessing, feature extraction or classification algorithms) provides a real improvement, if any. In the present paper, we propose a framework for reproducible and objective classification experiments in AD using three publicly available datasets (ADNI, AIBL and OASIS). The framework comprises: i) automatic conversion of the three datasets into a standard format (BIDS); ii) a modular set of preprocessing pipelines, feature extraction and classification methods, together with an evaluation framework, that provide a baseline for benchmarking the different components. We demonstrate the use of the framework for a large-scale evaluation on 1960 participants using T1 MRI and FDG PET data. In this evaluation, we assess the influence of different modalities, preprocessing, feature types (regional or voxel-based features), classifiers, training set sizes and datasets. Performances were in line with the state-of-the-art. FDG PET outperformed T1 MRI for all classification tasks. No difference in performance was found for the use of different atlases, image smoothing, partial volume correction of FDG PET images, or feature type. Linear SVM and L2-logistic regression resulted in similar performance and both outperformed random forests. The classification performance increased along with the number of subjects used for training. Classifiers trained on ADNI generalized well to AIBL and OASIS, performing better than the classifiers trained and tested on each of these datasets independently. All the code of the framework and the experiments is publicly available.

neuroscience

Latent-Based Imputation of Laboratory Measures from Electronic Health Records: Case for Complex Diseases

Imputation is a key step in Electronic Health Records-mining as it can significantly affect the conclusions derived from the downstream analysis. There are three main categories that explain the missingness in clinical settings-incompleteness, inconsistency, and inaccuracy-and these can capture a variety of situations: the patient did not seek treatment, the health care provider did not enter the information, etc. We used EHR data from patients diagnosed with Inflammatory Bowel Disease from Geisinger Health System to design a novel imputation that focuses on a complex phenotype. Our approach is based on latent-based analysis integrated with clustering to group patients based on their comorbidities before imputation. IBD is a chronic illness of unclear etiology and without a complete cure. We have taken advantage of the complexity of IBD to pre-process the EHR data of 10,498 IBD patients and show that imputation can be improved using shared latent comorbidities. The R code and sample simulated input data will be available at a future time.

bioinformatics

Gene-expression profiling of single cells from archival tissue with laser-capture microdissection and Smart-3SEQ

RNA sequencing (RNA-seq) is a sensitive and accurate method for quantifying gene expression. Small samples or those whose RNA is degraded, such as formalin-fixed, paraffin-embedded (FFPE) tissue, remain challenging to study with nonspecialized RNA-seq protocols. Here we present a new method, Smart-3SEQ, that accurately quantifies transcript abundance even with small amounts of total RNA and effectively characterizes small samples extracted by laser-capture microdissection (LCM) from FFPE tissue. We also obtain distinct biological profiles from FFPE single cells, which have been impossible to study with previous RNA-seq protocols, and we use these data to identify possible new macrophage phenotypes associated with the tumor microenvironment. We propose Smart-3SEQ as a highly cost-effective method to enable large gene-expression profiling experiments unconstrained by sample size and tissue availability. In particular, Smart-3SEQs compatibility with FFPE tissue unlocks an enormous number of archived clinical samples, and combined with LCM it allows unprecedented studies of small cell populations and single cells isolated by their in situ context.

genomics

Molecular Mapping Of YrTZ2, A Stripe Rust Resistance Gene In Wild Emmer Accession TZ-2 And Its Comparative Analyses With Aegilops tauschii

Wheat stripe rust, caused by Puccinia striiformis f. sp. tritici (Pst), is a devastating disease that can cause severe yield losses. Identification and utilization of stripe rust resistance genes are essential for effective breeding against the disease. Wild emmer accession TZ-2, originally collected from Mount Hermon, Israel, confers near-immunity resistance against several prevailing Pst races in China. A set of 200 F6:7 recombinant inbred lines (RILs) derived from a cross between susceptible durum wheat cultivar Langdon and TZ-2 was used for stripe rust evaluation. Genetic analysis indicated that the stripe rust resistance of TZ-2 to Pst race CYR34 was controlled by a single dominant gene, temporarily designated YrTZ2. Through bulked segregant analysis (BSA) and SSR mapping, YrTZ2 was located on chromosome arm 1BS and flanked by SSR markers Xwmc230 and Xgwm413 with genetic distance of 0.8 cM (distal) and 0.3 cM (proximal), respectively. By applying wheat 90K iSelect SNP genotyping assay, 11 polymorphic loci (consist of 250 SNP markers) closely linked with YrTZ2 were identified. YrTZ2 was further delimited into a 0.8 cM genetic interval between SNP marker IWB19368 and SSR marker Xgwm413, and co-segregated with SNP marker IWB28744 (attached with 28 SNP markers). Comparative genomics analyses revealed high level of collinearity between the YrTZ2 genomic region and the orthologous region of Aegilops tauschii 1DS. The genomic region between loci IWB19368 and IWB31649 harboring YrTZ2 is orthologous to a 24.5 Mb genomic region between AT1D0112 and AT1D0150, spanning 15 contigs on chromosome 1DS. The genetic and comparative maps of YrTZ2 provide framework for map-based cloning and marker-assisted selection (MAS) of YrTZ2.

plant biology