Search bioRxivSearch

Biology subjects

Jackson, R.

Publications and source records attributed to Jackson, R..

7 recordsLinked to original sources

The effects of training population design on genomic prediction accuracy in wheat

Genomic selection offers several routes for increasing genetic gain or efficiency of plant breeding programs. In various species of livestock there is empirical evidence of increased rates of genetic gain from the use of genomic selection to target different aspects of the breeders equation. Accurate predictions of genomic breeding value are central to this and the design of training sets is in turn central to achieving sufficient levels of accuracy. In summary, small numbers of close relatives and very large numbers of distant relatives are expected to enable accurate predictions.\n\nTo quantify the effect of some of the properties of training sets on the accuracy of genomic selection in crops we performed an extensive field-based winter wheat trial. In summary, this trial involved the construction of 44 F2:4 bi- and triparental populations, from which 2992 lines were grown on four field locations and yield was measured. For each line, genotype data were generated for 25,000 segregating single nucleotide polymorphism markers. The overall heritability of yield was estimated to 0.65, and estimates within individual families ranged between 0.10 and 0.85. Within cross genomic prediction accuracies of yield BLUEs were 0.125 - 0.127 using two different cross-validation approaches, and generally increased with training set size. Using related crosses in training and validation sets generally resulted in higher prediction accuracies than using unrelated crosses. The results of this study emphasize the importance of the training set design in relation to the genetic material to which the resulting prediction model is to be applied.

genetics

Human Cervical Keratinocyte-Derived Monolayer and Organoid Cultures for Disease Modelling and Drug Screening

The successful isolation and propagation of patient-derived keratinocytes from cervical lesions constitute a more appropriate model of cervical disease than traditional cervical cancer-derived cell lines such as SiHa and CaSki. Our aim was to streamline the growth of patient-obtained, cervical keratinocytes into a reproducible process. We performed an observational case series study with 60 women referred to colposcopy for a diagnostic biopsy. Main outcome measures were how many samples could be passaged at least once, and where enough cells could be established, to precisely define their proliferation profile over time. Altering cell culture conditions over those reported by other groups markedly improved outcomes. We were also successful in making freeze backs which could be resuscitated for additional experiments. For best results, biopsy-intrinsic factors such as size and tissue digestion appear to be major variables. This seems to be the first systematic report with a well characterized and defined sample size, detailed protocol, carefully assessed cell yield and performance, and to successfully grow multi-layered, organoid cultures from cervical keratinocytes. This research is particularly impactful for constituting a sample repository-on-demand for appropriate disease modelling and drug screening under the umbrella of personalized health.

cancer biology

Genome-wide Meta-analysis of 158,000 Individuals of European Ancestry Identifies Three Loci Associated with Chronic Back Pain

OBJECTIVESTo conduct a genome-wide association study (GWAS) meta-analysis of chronic back pain (CBP).\n\nMETHODSAdults of European ancestry were included from 16 cohorts in Europe and North America. CBP cases were defined as those reporting back pain present for >3-6 months; non-cases were included as comparisons (\"controls\"). Each cohort conducted genotyping using commercially available arrays followed by imputation. GWAS used logistic regression models with additive genetic effects, adjusting for age, sex, study-specific covariates, and population substructure. The threshold for genome-wide significance in the fixed-effect inverse-variance weighted meta-analysis was p<5x10-8. Suggestive (p<5x10-7) and genome-wide significant (p<5x10-8) variants were carried forward for replication or further investigation in an independent sample.\n\nRESULTSThe discovery sample was comprised of 158,025 individuals, including 29,531 CBP cases. A genome-wide significant association was found for the intronic variant rs12310519 in SOX5 (OR 1.08, p=7.2x10-10). This was subsequently replicated in an independent sample of 283,752 subjects, including 50,915 cases (OR 1.06, p=5.3x10-11), and exceeded genome-wide significance in joint meta-analysis (0R=1.07, p=4.5x10-19). We found suggestive associations at three other loci in the discovery sample, two of which exceeded genome-wide significance in joint meta-analysis: an intergenic variant, rs7833174, located between CCDC26 and GSDMC (OR 1.05, p=4.4x10-13), and an intronic variant, rs4384683, in DCC (OR 0.97, p=2.4x10-10).\n\nDISCUSSIONIn this first reported meta-analysis of GWAS for CBP, we identified and replicated a genetic locus associated with CBP (SOX5). We also identified 2 other loci that reached genome-wide significance in a 2-stage joint meta-analysis (CCDC26/GSDMC and DCC).

genomics

SemEHR: A General-purpose Semantic Search System to Surface Semantic Data from Clinical Notes for Tailored Care, Trial Recruitment and Clinical Research

ObjectiveUnlocking the data contained within both structured and unstructured components of Electronic Health Records (EHRs) has the potential to provide a step change in data available forsecondary research use, generation of actionable medical insights, hospital management and trial recruitment. To achieve this, we implemented SemEHR - a semantic search and analytics, open source tool for EHRs.\n\nMethodsSemEHR implements a generic information extraction (IE) and retrieval infrastructure by identifying contextualised mentions of a wide range of biomedical concepts within EHRs. Natural Language Processing (NLP) annotations are further assembled at patient level and extended with EHR-specific knowledge to generate a timeline for each patient. The semantic data is serviced via ontology-based search and analytics interfaces.\n\nResultsSemEHR has been deployed to a number of UK hospitals including the Clinical Record Interactive Search (CRIS), an anonymised replica of the EHR of the UK South London and Maudsley (SLaM) NHS Foundation Trust, one of Europes largest providers of mental health services. In two CRIS-based studies, SemEHR achieved 93% (Hepatitis C case) and 99% (HIV case) F-Measure results in identifying true positive patients. At Kings College Hospital in London, as part of the CogStack programme (github.com/cogstack), SemEHR is being used to recruit patients into the UK Dept of Health 100k Genome Project (genomicsengland.co.uk). The validation study suggests that the tool can validate previously recruited cases and is very fast in searching phenotypes - time for recruitment criteria checking reduced from days to minutes. Validated on an open intensive care EHR data - MIMICIII, the vital signs extracted by SemEHR can achieve around 97% accuracy.\n\nConclusionResults from the multiple case studies demonstrate SemEHRs efficiency - weeks or months of work can be done within hours or minutes in some cases. SemEHR provides a more comprehensive view of a patient, bringing in more and unexpected insight compared to study-oriented bespoke information extraction systems.\n\nSemEHR is open source available at https://github.com/CogStack/SemEHR.

bioinformatics

Epithelial stratification shapes infection dynamics

Infections of stratified epithelia collectively represent a large burden on global health. Experimental models provide a means to understand how the cell dynamics themselves influence the outcomes of these infections. Mathematical approaches are needed to improve quantification and theoretical advancement of these complex systems. Here, we develop a general ecology-inspired model for stratified epithelial dynamics, which allows us to simulate infections and to estimate parameters that are difficult to measure with organotypic cell cultures. To explore how epithelial cell dynamics affect infection dynamics, we focus on two contrasting pathogens of the cervicovaginal epithelium: Chlamydia trachomatis and Human papillomaviruses. We find that key infection symptoms stem from differential interactions with the layers, while clearance and pathogen burden are bottom-up processes. Cell protective responses to infections (e.g. increased cell proliferation) generally lowered pathogen load but there were specific effects based on infection strategies. These generic responses by the epithelium, then, will have varying results depending on the pathogens infection strategy. Our modeling approach opens new perspectives for 3D tissue culture experimental systems of infections and, more generally, for developing and testing hypotheses related to infections of stratified epithelia.

systems biology

Pathogen-Host Analysis Tool (PHAT): an Integrative Platform to Analyze Pathogen-Host Relationships in Next-Generation Sequencing Data

SummaryThe Pathogen-Host Analysis Tool (PHAT) is an application for processing and analyzing next-generation sequencing (NGS) data as it relates to relationships between pathogen and host organisms. Unlike custom scripts and tedious pipeline programming, PHAT provides an integrative platform encompassing raw and aligned sequence and reference file input, quality control (QC) reporting, alignment and variant calling, linear and circular alignment viewing, and graphical and tabular output. This novel tool aims to be user-friendly for life scientists studying diverse pathogen-host relationships.\n\nAvailability and ImplementationThe project is publicly available on GitHub (https://github.com/chgibb/PHAT) and includes convenient installers, as well as portable and source versions, for both Windows and Linux (Debian and RedHat). Up-to-date documentation for PHAT, including user guides and development notes, can be found at https://chgibb.github.io/PHATDocs/. We encourage users and developers to provide feedback (error reporting, suggestions, and comments) using GitHub Issues.\n\nContactLead software developer: chris.gibb@outlook.com

bioinformatics

CogStack - Experiences Of Deploying IntegratedInformation Retrieval And Extraction Services In A Large National Health Service Foundation Trust Hospital

BackgroundTraditional health information systems are generally devised to support clinical data collection at the point of care. However, as the significance of the modern information economy expands in scope and permeates the healthcare domain, there is an increasing urgency for healthcare organisations to offer information systems that address the expectations of clinicians, researchers and the business intelligence community alike. Amongst other emergent requirements, the principal unmet need might be defined as the 3R principle (right data, right place, right time) to address deficiencies in organisational data flow while retaining the strict information governance policies that apply within the UK National Health Service (NHS). Here, we describe our work on creating and deploying a low cost structured and unstructured information retrieval and extraction architecture within Kings College Hospital, the management of governance concerns and the associated use cases and cost saving opportunities that such components present.\n\nResultsTo date, our CogStack architecture has processed over 300 million lines of clinical data, making it available for internal service improvement projects at Kings College London. On generated data designed to simulate real world clinical text, our de-identification algorithm achieved up to 94% precision and up to 96% recall.\n\nConclusionWe describe a toolkit which we feel is of huge value to the UK (and beyond) healthcare community. It is the only open source, easily deployable solution designed for the UK healthcare environment, in a landscape populated by expensive proprietary systems. Solutions such as these provide a crucial foundation for the genomic revolution in medicine.

bioinformatics