Search bioRxivSearch

Biology subjects

Barton, C.

Publications and source records attributed to Barton, C..

5 recordsLinked to original sources

Multicenter validation of a machine learning algorithm for 48 hour all-cause mortality prediction

PurposeThis study evaluates a machine-learning-based mortality prediction tool.\n\nMaterials and MethodsWe conducted a retrospective study with data drawn from three academic health centers. Inpatients of at least 18 years of age and with at least one observation of each vital sign were included. Predictions were made at 12, 24, and 48 hours before death. Models fit to training data from each institution were evaluated on hold-out test data from the same institution and data from the remaining institutions. Predictions were compared to those of qSOFA and MEWS using area under the receiver operating characteristic curve (AUROC).\n\nResultsFor training and testing on data from a single institution, machine learning predictions averaged AUROCs of 0.97, 0.96, and 0.95 across institutional test sets for 12-, 24-, and 48-hour predictions, respectively. When trained and tested on data from different hospitals, the algorithm achieved AUROC up to 0.95, 0.93, and 0.91, for 12-, 24-, and 48-hour predictions, respectively. MEWS and qSOFA had average 48-hour AUROCs of 0.86 and 0.82, respectively.\n\nConclusionThis algorithm may help identify patients in need of increased levels of clinical care.

bioinformatics

The contribution of mitochondrial metagenomics to largescale data mining and phylogenetic analysis of Coleoptera

A phylogenetic tree at the species level is still far off for highly diverse insect orders, including the Coleoptera, but the taxonomic breadth of public sequence databases is growing. In addition, new types of data may contribute to increasing taxon coverage, such as metagenomic shotgun sequencing for assembly of mitogenomes from bulk specimen samples. The current study explores the application of these techniques for large-scale efforts to build the tree of Coleoptera. We used shotgun data from 17 different ecological and taxonomic datasets (5 unpublished) to assemble a total of 1942 mitogenome contigs of >3000 bp. These sequences were combined into a single dataset together with all mitochondrial data available at GenBank, in addition to nuclear markers widely used in molecular phylogenetics. The resulting matrix of nearly 16000 species with two or more loci produced trees (RAxML) showing overall congruence with the Linnaean taxonomy at hierarchical levels from suborders to genera. We tested the role of full-length mitogenomes in stabilizing the tree from GenBank data, as mitogenomes might link terminals with non-overlapping gene representation. However, the mitogenome data were only partly useful in this respect, presumably because of the purely automated approach to assembly and gene delimitation, but improvements in future may be possible by using multiple assemblers and manual curation. In conclusion, the combination of data mining and metagenomic sequencing of bulk samples provided the largest phylogenetic tree of Coleoptera to date, which represents a summary of existing phylogenetic knowledge and a defensible tree of great utility, in particular for studies at the intra-familial level, despite some shortcomings for resolving basal nodes.

evolutionary biology

Multicenter validation of a sepsis prediction algorithm using only vital sign data in the emergency department, general ward and ICU

ObjectivesWe validate a machine learning-based sepsis prediction algorithm (InSight) for detection and prediction of three sepsis-related gold standards, using only six vital signs. We evaluate robustness to missing data, customization to site-specific data using transfer learning, and generalizability to new settings.\n\nDesignA machine learning algorithm with gradient tree boosting. Features for prediction were created from combinations of only six vital sign measurements and their changes over time.\n\nSettingA mixed-ward retrospective data set from the University of California, San Francisco (UCSF) Medical Center (San Francisco, CA) as the primary source, an intensive care unit data set from the Beth Israel Deaconess Medical Center (Boston, MA) as a transfer learning source, and four additional institutions datasets to evaluate generalizability.\n\nParticipants684,443 total encounters, with 90,353 encounters from June 2011 to March 2016 at UCSF.\n\nInterventionsnone\n\nPrimary and secondary outcome measuresArea under the receiver operating characteristic curve (AUROC) for detection and prediction of sepsis, severe sepsis, and septic shock.\n\nResultsFor detection of sepsis and severe sepsis, InSight achieves an area under the receiver operating characteristic (AUROC) curve of 0.92 (95% CI 0.90 - 0.93) and 0.87 (95% CI 0.86 - 0.88), respectively. Four hours before onset, InSight predicts septic shock with an AUROC of 0.96 (95% CI 0.94 -0.98), and severe sepsis with an AUROC of 0.85 (95% CI 0.79 - 0.91).\n\nConclusionsInSight outperforms existing sepsis scoring systems in identifying and predicting sepsis, severe sepsis, and septic shock. This is the first sepsis screening system to exceed an AUROC of 0.90 using only vital sign inputs. InSight is robust to missing data, can be customized to novel hospital data using a small fraction of site data, and retained strong discrimination across all institutions.\n\nStrengths and limitations of this studyO_LIMachine learning is applied to the detection and prediction of three separate sepsis standards in the emergency department, general ward and intensive care settings.\nC_LIO_LIOnly six commonly measured vital signs are used as input for the algorithm.\nC_LIO_LIThe algorithm is robust to randomly missing data.\nC_LIO_LITransfer learning successfully leverages large dataset information to a target dataset.\nC_LIO_LIRetrospective nature of the study does not predict clinician reaction to information.\nC_LI

bioinformatics

Pediatric Severe Sepsis Prediction Using Machine Learning

Early detection of pediatric severe sepsis is necessary in order to administer effective treatment. In this study, we assessed the efficacy of a machine-learning-based prediction algorithm applied to electronic healthcare record (EHR) data for the prediction of severe sepsis onset. The resulting prediction performance was compared with the Pediatric Logistic Organ Dysfunction score (PELOD-2) and pediatric Systemic Inflammatory Response Syndrome score (SIRS) using cross-validation and pairwise t-tests. EHR data were collected from a retrospective set of de-identified pediatric inpatient and emergency encounters drawn from the University of California San Francisco (UCSF) Medical Center, with encounter dates between June 2011 and March 2016. Patients (n = 11,127) were 2-17 years of age and 103 [0.93%] were labeled severely septic. In four-fold cross-validation evaluations, the machine learning algorithm achieved an AUROC of 0.912 for discrimination between severely septic and control pediatric patients at onset and AUROC of 0.727 four hours before onset. Under the same measure, the prediction algorithm also significantly outperformed PELOD-2 (p < 0.05) and SIRS (p < 0.05) in the prediction of severe sepsis four hours before onset. This machine learning algorithm has the potential to deliver high-performance severe sepsis detection and prediction for pediatric inpatients.

bioinformatics

Prediction of Acute Kidney Injury with a Machine Learning Algorithm using Electronic Health Record Data

BackgroundA major problem in treating acute kidney injury (AKI) is that clinical criteria for recognition are markers of established kidney damage or impaired function; treatment before such damage manifests is desirable. Clinicians could intervene during what may be a crucial stage for preventing permanent kidney injury if patients with incipient AKI and those at high risk of developing AKI could be identified.\n\nMethodsWe used a machine learning technique, boosted ensembles of decision trees, to train an AKI prediction tool on retrospective data from inpatients at Stanford Medical Center and intensive care unit patients at Beth Israel Deaconess Medical Center. We tested the algorithms ability to detect AKI at onset, and to predict AKI 12, 24, 48, and 72 hours before onset, and compared its 3-fold cross-validation performance to the SOFA score for AKI identification in terms of Area Under the Receiver Operating Characteristic (AUROC).\n\nResultsThe prediction algorithm achieves AUROC of 0.872 (95% CI 0.867, 0.878) for AKI onset detection, superior to the SOFA score AUROC of 0.815 (P < 0.01). At 72 hours before onset, the algorithm achieves AUROC of 0.728 (95% CI 0.719, 0.737), compared to the SOFA score AUROC of 0.720 (P < 0.01).\n\nConclusionsThe results of these experiments suggest that a machine-learning-based AKI prediction tool may offer important prognostic capabilities for determining which patients are likely to suffer AKI, potentially allowing clinicians to intervene before kidney damage manifests.

bioinformatics