Search bioRxiv⌕ Search

Biology subjects

Carrion, J. T.

Publications and source records attributed to Carrion, J. T..

2 recordsLinked to original sources

CryoJAM: Automating Protein Homolog Fitting in Medium Resolution Cryo-EM Density Maps

Obtaining atomic structures of large protein complexes from medium-resolution cryogenic electron-microscopy (cryo-EM) density maps is a critical bottleneck in the cryo-EM workflow. CryoJAM aims to automate this process by using a 3D Convolutional Neural Network model within a U-Net architecture. This model is trained on a novel loss function that leverages Fourier-Shell Correlation (FSC), as a proxy for quality of fit, and Root Mean Squared Error (RMSE) to help optimize fits within real space. Capitalizing on the gold-standard status of FSC in cryo-EM, this method introduces an innovative implementation of FSC into cryo-EM model fitting software, enhancing the precision and efficiency of structural analysis. After 25 epochs, CryoJAM successfully reduced the RMSE in 21 out of 26 of the test cases, effectively fitting homologous protein structures into medium-resolution cryo-EM densities.

biophysics↗

A data-fusion approach to identifying developmental dyslexia from multi-omics datasets

This exploratory study tested and validated the use of data fusion and machine learning techniques to probe high-throughput omics and clinical data with a goal of exploring the etiology of developmental dyslexia. Developmental dyslexia is the leading learning disability in school aged children affecting roughly 5-10% of the US population. The complex biological and neurological phenotype of this life altering disability complicates its diagnosis. Phenome, exome, and metabolome data was collected allowing us to fully explore this system from a behavioral, cellular, and molecular point of view. This study provides a proof of concept showing that data fusion and ensemble learning techniques can outperform traditional machine learning techniques when provided small and complex multi-omics and clinical datasets. Heterogenous stacking classifiers consisting of single-omic experts/models achieved an accuracy of 86%, F1 score of 0.89, and AUC value of 0.83. Ensemble methods also provided a ranked list of important features that suggests exome single nucleotide polymorphisms found in the thalamus and cerebellum could be potential biomarkers for developmental dyslexia and heavily influenced the classification of DD within our machine learning models.

bioinformatics↗