Search bioRxiv⌕ Search

Biology subjects

Marcu, D.

Publications and source records attributed to Marcu, D..

2 recordsLinked to original sources

Within-ejaculate haploid selection reduces disease biomarkers in human sperm

The germline is widely regarded as a checkpoint against the inheritance of damaged genomes, yet the mechanisms that could enact such filtering remain poorly resolved. Of the millions of sperm in a human ejaculate, only one fertilises the egg, creating a strong opportunity for selection among gametes. Combining within-ejaculate selection on sperm quality with whole-genome sequencing and proteomics in healthy donors, we find that this selection is biased against molecular signatures of age-related disease. The most reproducibly diverging genes are tumour suppressors, and genes diverging under sperm longevity-based selection are enriched for senescence-associated genes involved in genome maintenance and oxidative stress response. High-quality sperm are further depleted of inflammation- and cancer-associated proteins. This genomic signature is conserved in zebrafish, in which longer-lived sperm sire longer-lived offspring. We propose that within-ejaculate selection acts as a pre-fertilisation filter against age-related disease alleles, with implications for offspring lifespan and healthspan.

molecular biology↗

NERO: A Biomedical Named-entity (Recognition) Ontology with a Large, Annotated Corpus Reveals Meaningful Associations Through Text Embedding

Machine reading is essential for unlocking valuable knowledge contained in the millions of existing biomedical documents. Over the last two decades 1,2, the most dramatic advances in machine-reading have followed in the wake of critical corpus development3. Large, well-annotated corpora have been associated with punctuated advances in machine reading methodology and automated knowledge extraction systems in the same way that ImageNet 4 was fundamental for developing machine vision techniques. This study contributes six components to an advanced, named-entity analysis tool for biomedicine: (a) a new, Named-Entity Recognition Ontology (NERO) developed specifically for describing entities in biomedical texts, which accounts for diverse levels of ambiguity, bridging the scientific sublanguages of molecular biology, genetics, biochemistry, and medicine; (b) detailed guidelines for human experts annotating hundreds of named-entity classes; (c) pictographs for all named entities, to simplify the burden of annotation for curators; (d) an original, annotated corpus comprising 35,865 sentences, which encapsulate 190,679 named entities and 43,438 events connecting two or more entities; (e) validated, off-the-shelf, named-entity recognition automated extraction, and; (f) embedding models that demonstrate the promise of biomedical associations embedded within this corpus.

systems biology↗