Search bioRxivSearch

Biology subjects

Gong, Y.

Publications and source records attributed to Gong, Y..

5 recordsLinked to original sources

BioIMA: a one-click desktop tool for standardized extraction of phenotypic traits from biological images

Standardized extraction of quantitative phenotypes from images is increasingly important across plant biology, from ecological and evolutionary studies to genetics, breeding, and functional genomics. However, as large image datasets are increasingly used for trait analysis, many biologically relevant traits, including size, shape, color, and spatial patterning, are still measured manually or using fragmented semi-automated workflows. These limitations reduce throughput, reproducibility, and accessibility, especially for researchers without computational expertise. Here, we present BioIMA, an open-source desktop tool for rapid and standardized phenotyping from biological images. BioIMA integrates foundation model-based segmentation with automated trait computation, allowing users to extract quantitative measurements from images through an intuitive graphical interface and without model training. To validate its performance, we quantified a set of knot morphological traits in two Populus species, as these measurements are typically time-consuming to perform manually. Automatic measurements showed strong agreement with manual ImageJ-based measurements (R2 > 0.95), while reducing per-image processing time by approximately 75% (from ~15 s to ~4 s). BioIMA was further applied to diverse plant datasets, including Helianthus and Rhododendron images with varying morphologies and background conditions. Although developed for plant phenotyping, BioIMA may also be extended to other biological samples where region-based size, shape, or color traits are of interest. By combining accessibility and standardization in a lightweight local application, BioIMA provides a practical community resource for image-based phenotyping in ecological and evolutionary studies.

bioinformatics

Investigation of novel functions of three genes in oriental river prawn, Macrobrachium nipponense: Molecular Cloning, Expression, and In situ Hybridization Analysis

Three genes were predicted to be potentially involved in the male sexual development in M. nipponense, including the Gem-associated protein 2-like isoform X1 (GEM), Ferritin peptide, and DNA polymerase zeta catalytic subunit (Rev3). In this study, we aimed to investigate their novel functions in depth. The full-length cDNA sequence of Mn-GEM was 1,018 bp, encoding 258 amino acids. The partial Mn-Rev3 cDNA sequence was 6,832 bp, encoding 1,203 amino acids. Tissue distribution indicated that all of these three genes have higher expression level in testis and androgenic gland, implying their novel functions in male sexual development. In situ hybridization analysis further confirmed the novel roles of these three genes. Rev3 promote the testis development during the whole reproductive cycle, while GEM and ferritin only promote the activation of testis development. Besides, these three genes play essential roles in funicular structure development surrounding the androgenic gland cells, which promote and support the formation of androgenic gland cells. The expression in hepatopancreas cells also suggested their role in immune system in M. nipponense. This study advances our understanding of male sexual development in M. nipponense, as well as providing the basis for further studies of male sexual differentiation and development in crustaceans.

zoology

Molecular Detection of H.pylori Antibiotic-Resistant Genes and Bioinformatics Predictive Analysis

To explore the mutation characteristics of H.pylori resistance-related genes to antibiotics of clarithromycin, levofloxacin and metronidazole. 23S rRNA, gyrA, gyrB, rdxA and frxA genes were amplified and sequenced, respectively. Their structural alteration after mutation was predicted using bioinformatics software. In the clarithromycin-resistant strains, the mutation rate in site A2143G was 74.2% (n=23). The mutations in sites C1883T, C2131T and T2179G might cause structural alteration. In the levofloxacin-resistant strains, the mutation rates in 87 (N to K/I) and 91 (D to N/Y/G) of gyrA were 28.6% (n=16) and 12.5% (n =7), respectively. Meanwhile, one of the mutation strains in site 91 was accompanied by D99N variation. Additionally, a D143E mutation was found in one drug-resistant strain. Some changes of tertiary structure occurred after these mutations. The mutation types of RdxA protein consisted of protein truncation caused by premature stop codons (n=26, 33.3%), frameshift mutations (n=8, 10.3%), FMN-binding sites (n=16, 20.5%) and the others (n=11, 14.1%). Predictive analysis showed that mutations in the first three groups and the A118S of the last group could lead to structural alteration. Our study suggested the clarithromycin-resistant sites of H.pylori were mainly located in A2143G of 23S rRNA. C1883T, C2131T and T2179G might also be related to resistance. Levofloxacin resistance was mainly based on the amino acid changes in 87 and 91 sites of gyrA. The new sites D99N and D143E might also be associated with resistance. Metronidazole resistance was related to RdxA protein truncation, frameshift, and FMN binding. The new site A118S might also be linked to drug resistance.

microbiology

Robust Estimation Of Hi-C Contact Matrices Using Fused Lasso Reveals Preferential Insulation Of Super-Enhancers By Strong TAD Boundaries

The metazoan genome is compartmentalized in megabase-scale areas of highly interacting chromatin known as topologically associating domains (TADs), typically identified by computational analyses of Hi-C sequencing data. TADs are demarcated by boundaries that are largely conserved across cell types and even across species, although, increasing evidence suggests that the seemingly invariant TAD boundaries may exhibit plasticity and their insulating strength can vary. However, a genome-wide characterization of TAD boundary strength in mammals is still lacking. A systematic classification and characterization of TAD boundaries may generate new insights into their function. In this study, we first use fused two-dimensional lasso as a machine learning method to improve Hi-C contact matrix reproducibility, and, subsequently, we categorize TAD boundaries based on their insulation score. We demonstrate that higher TAD boundary insulation scores are associated with elevated CTCF levels and that they may differ across cell types. Intriguingly, we observe that super-enhancer elements are preferentially insulated by strong boundaries, i.e. boundaries of higher insulation score. Furthermore, we perform a pan-cancer analysis to demonstrate that strong TAD boundaries and super-enhancer elements are frequently co-duplicated in cancer patients. Taken together, our findings suggest that super-enhancers insulated by strong TAD boundaries may be exploited, as a functional unit, by cancer cells to promote oncogenesis.

bioinformatics

lncRNA-screen: an interactive platform for computationally screening long non-coding RNAs in large genomics datasets

Long non-coding RNAs (lncRNAs) have emerged as a class of factors that are important for regulating development and cancer. Computational prediction of lncRNAs from ultra-deep RNA sequencing has been successful in identifying candidate lncRNAs. However, the complexity of handling and integrating different types of genomics data poses significant challenges to experimental laboratories that lack extensive genomics expertise. To address this issue, we have developed lncRNA-screen, a comprehensive pipeline for computationally screening putative lncRNA transcripts over large multimodal datasets. The main objective of this work is to facilitate the computational discovery of lncRNA candidates to be further examined by functional experiments. lncRNA-screen provides a fully automated easy-to-run pipeline which performs data download, RNA-seq alignment, assembly, quality assessment, transcript filtration, novel lncRNA identification, coding potential estimation, expression level quantification, histone mark enrichment profile integration, differential expression analysis, annotation with other type of segmented data (CNVs, SNPs, Hi-C, etc.) and visualization. Importantly, lncRNA-screen generates an interactive report summarizing all interesting lncRNA features including genome browser snapshots and lncRNA-mRNA interactions based on Hi-C data. In summary, our pipeline provides a comprehensive solution for lncRNA discovery and an intuitive interactive report for identifying promising lncRNA candidates. lncRNA-screen is available as free open-source software on GitHub.

bioinformatics