Search bioRxiv⌕ Search

Biology subjects

BIANCO, S.

Publications and source records attributed to BIANCO, S..

2 recordsLinked to original sources

The projection basis determines the information ceiling for perturbation prediction

Recent benchmarks show that deep-learning models for perturbation prediction do not outperform simple baselines operating in principal-component (PCA) space. We explain this with an information-theoretic ceiling: for any orthonormal projection basis{Phi} , the squared correlation between prediction and truth is bounded by the variance the basis explains (r2 [≤] VE), so no model complexity can recover signal discarded at projection. On chemical perturbations (sciPlex3, LINCS L1000), the eigenbasis of a gene association network captures only 10-12% of drug-response variance and yields chance-level predictions, while PCA captures 90-99%. Graph wavelets built on the same network recover [~]88%, localising the drug signal in high-frequency modes that the standard low-pass eigenbasis discards. On CRISPRa genetic perturbations the ranking inverts: the network basis outperforms PCA across all dimensions tested. Controls on topology, null networks and data leakage confirm the effect is structural. The right basis depends on the perturbation modality: PCA captures the variance that drives chemical responses, the network basis captures the cascade structure that drives genetic ones, and bases that access the networks full graph spectrum (such as graph wavelets) recover both from the same topology.

molecular biology↗

Annotation-free Learning of Plankton for Classification and Anomaly Detection

The acquisition of increasingly large plankton digital image datasets requires automatic methods of recognition and classification. As data size and collection speed increases, manual annotation and database representation are often bottlenecks for utilization of machine learning algorithms for taxonomic classification of plankton species in field studies. In this paper we present a novel set of algorithms to perform accurate detection and classification of plankton species with minimal supervision. Our algorithms approach the performance of existing supervised machine learning algorithms when tested on a plankton dataset generated from a custom-built lensless digital device. Similar results are obtained on a larger image dataset obtained from the Woods Hole Oceanographic Institution. Our algorithms are designed to provide a new way to monitor the environment with a class of rapid online intelligent detectors. Author SummaryPlankton are at the bottom of the aquatic food chain and marine phytoplankton are estimated to be responsible for over 50% of all global primary production [1] and play a fundamental role in climate regulation. Thus, changes in plankton ecology may have a profound impact on global climate, as well as deep social and economic consequences. It seems therefore paramount to collect and analyze real time plankton data to understand the relationship between the health of plankton and the health of the environment they live in. In this paper, we present a novel set of algorithms to perform accurate detection and classification of plankton species with minimal supervision. The proposed pipeline is designed to provide a new way to monitor the environment with a class of rapid online intelligent detectors.

bioinformatics↗