Search bioRxiv⌕ Search

bioRxiv · 10.1101/2024.02.12.580001

Named Entity Recognition of Pharmacokinetic parameters in the scientific literature

Abstract

The development of accurate predictions for a new drugs absorption, distribution, metabolism, and excretion profiles in the early stages of drug development is crucial due to high candidate failure rates. The absence of comprehensive, standardised, and updated pharmacokinetic (PK) repositories limits pre-clinical predictions and often requires searching through the scientific literature for PK parameter estimates from similar compounds. While text mining offers promising advancements in automatic PK parameter extraction, accurate Named Entity Recognition (NER) of PK terms remains a bottleneck due to limited resources. This work addresses this gap by introducing novel corpora and language models specifically designed for effective NER of PK parameters. Leveraging active learning approaches, we developed an annotated corpus containing over 4,000 entity mentions found across the PK literature on PubMed. To identify the most effective model for PK NER, we fine-tuned and evaluated different NER architectures on our corpus. Fine-tuning BioBERT exhibited the best results, achieving a strict F1 score of 90.37% in recognising PK parameter mentions, significantly outperforming heuristic approaches and models trained on existing corpora. To accelerate the development of end-to-end PK information extraction pipelines and improve pre-clinical PK predictions, the PK NER models and the labelled corpus were released open source at https://github.com/PKPDAI/PKNER.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hernandez, F. G., Nguyen, Q., Smith, V. C., Cordero, J. A., Ballester, M. R., Duran, M., Sole, A., Chotsiri, P., Wattanakul, T., Mundin, G., Lilaonitkul, W., Standing, J. F., Kloprogge, F.. 2024-02-14. Named Entity Recognition of Pharmacokinetic parameters in the scientific literature. https://doi.org/10.1101/2024.02.12.580001

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Aquaporin-9 and aquaporin-10 but not aquaporin-3 confer susceptibility to dimethylarsinic acid genotoxicity in human cells

Human metabolism converts inorganic arsenic to the pentavalent methylated species MMA(V) and DMA(V), the forms most people excrete, and the forms long read as the end of a detoxification pathway. Whether a transporter sets how much of these metabolites reaches the genome has not been tested in a mammalian cell. We expressed human AQP3, AQP7, AQP9 or AQP10 in HEK293T and MRC5-SV40 cells and measured gamma-H2AX by flow cytometry across dose series of As(V), MMA(V) and DMA(V), pairing every aquaporin with a GFP-Tubulin control and an untransfected mock acquired in the same replicate. As(V) was inactive in HEK293T cells and only weakly active in MRC5-SV40 cells to 20 micromolar, and both methylated species damaged DNA only in the millimolar range, DMA(V) being the more potent of the two in both cell lines. Against that weak baseline, AQP9 and AQP10 raised DMA(V)-induced gamma-H2AX in HEK293T cells by roughly 17 percentage points over the matched control, more than doubling the damage the same exposure produced in control cells, whereas AQP3 and AQP7 changed it not at all. AQP9 alone remained active with MMA(V). The ranking held in MRC5-SV40 fibroblasts at one-sixth the size, and within single wells the damage rose with the amount of AQP9 a cell carried while the control was flat. Aquaglyceroporins therefore discriminate among arsenic species, and AQP9 and AQP10 turn a weakly genotoxic metabolite into a substantially more genotoxic one.

pharmacology and toxicology↗

Quantitative Systems Pharmacology Model for Trop-2 Targeting Antibody-Drug Conjugate in Triple-Negative Breast Cancer

TROP2-targeted antibody-drug conjugates (ADCs) have demonstrated promising clinical activity in triple-negative breast cancer (TNBC) as monotherapies; however, therapeutic benefit varies among patients. Combination strategies pairing TROP2-targeted ADCs with immune checkpoint inhibitors are also being investigated. Elucidating the mechanistic drivers of ADC monotherapy variability and enabling the rational development of combination regimens require computational frameworks that integrate ADC pharmacology with tumor-immune interactions. A quantitative systems pharmacology (QSP) model is presented that incorporates an ADC module into our established immuno-oncology model for TNBC. The module captures ADC and payload pharmacokinetics and pharmacodynamics. TNBC heterogeneity is represented by two tumor cell clones with high and low TROP2 expression, informed by prior characterizations, and differential sensitivity to the ADC payload is incorporated as an intrinsic property of each clone. Although generalizable, the model was applied to the TROP2-targeted ADC sacituzumab govitecan (SG, TRODELVY). A virtual patient cohort was generated using Latin hypercube sampling and calibrated against objective response rate (ORR) data from SG Phase I/II TNBC basket trial. The model predicted an ORR of 33.2% consistent with ASCENT study (NCT02574455). Simulations suggest TROP2-mediated delivery contributes modestly to SG efficacy with tumor exposure driven largely by systemically released SN-38 payload being sufficient to induce cytotoxicity. Tumor heterogeneity emerged as a key determinant of response with ORR increasing as the fraction of payload-sensitive clones increased. Overall, this QSP framework for TROP2-targeted ADCs accounts for TNBC heterogeneity and is extendable to other ADCs and targets enabling interrogation of ADC mechanisms of action in conjunction with tumor-immune interactions.

pharmacology and toxicology↗

Computer-Assisted Systematic Chemical-Space Mapping of a First-in-Class Peripherally Restricted α2AAR Agonist through Scaffold-Seeded Enumeration

CC10137 is a first-in-class peripherally restricted 2A-adrenergic receptor (2AAR) agonist with broad-spectrum analgesic efficacy and a favorable safety profile. Systematic exploration of the chemical space surrounding first-in-class leads is important for defining series boundaries and guiding continued optimization, but conventional analogue-by-analogue medicinal chemistry samples only a small fraction of the accessible structural space. Here, we used a scaffold-seeded enumeration strategy to expand the chemical space surrounding CC10137 from four SAR-informed seed compounds comprising CC10137 and three closely related structural variants. Application of predefined medicinal chemistry transformation rules in StarDrop generated a virtual library of 16,601,163 unique structures. Morgan fingerprint-based principal component analysis indicated that the library occupied a highly multidimensional structural space involving variation in scaffold substitution, peripheral functional groups, and side-chain composition. A retrospective comparison set of 43 compounds independently designed and experimentally characterized in the earlier CC10137 program represented only approximately 0.00026% of the 16.6-million-member library, yet all 43 were recovered as exact structural matches. Three compounds selected directly from the virtual library retained 2AAR binding affinity and agonist potency below 25 nM. Five representative compounds further showed significant anti-allodynic effects in the in vivo spared nerve injury model, with inhibition rates ranging from 39.3% to 55.7%. These findings support scaffold-seeded computational enumeration as a practical strategy for systematic chemical-space mapping around a first-in-class lead and for identifying additional pharmacologically active structural regions for further optimization.

pharmacology and toxicology↗