Search bioRxivSearch

Biology subjects

Le, M.

Publications and source records attributed to Le, M..

3 recordsLinked to original sources

Direct readout of neural stem cell transgenesis with an integration-coupled gene expression switch

SUMMARYStable genomic integration of exogenous transgenes is critical for neurodevelopmental and neural stem cell studies. Despite the emergence of tools driving genomic insertion at high rates with DNA vectors, transgenesis procedures remain fundamentally hindered by the impossibility to distinguish integrated transgenes from residual episomes. Here, we introduce a novel genetic switch termed iOn that triggers gene expression upon insertion in the host genome, enabling simple, rapid and faithful identification of integration events following transfection with naked plasmids accepting large cargoes. In vitro, iOn permits rapid drug-free stable transgenesis of mouse and human pluripotent stem cells with multiple vectors. In vivo, we demonstrate accurate cell lineage tracing, assessment of regulatory elements and mosaic analysis of gene function in somatic transgenesis experiments that reveal new aspects of neural progenitor potentialities and interactions. These results establish iOn as an efficient and widely applicable strategy to report transgenesis and accelerate genetic engineering in cultured systems and model organisms.

developmental biology

Identifying conservation priorities in a defaunated tropical biodiversity hotspot

AimUnsustainable hunting is leading to widespread defaunation across the tropics. To mitigate against this threat with limited conservation resources, stakeholders must make decisions on where to focus anti-poaching activities. Identifying priority areas in a robust way allows decision-makers to target areas of conservation importance, therefore maximizing the impact of conservation interventions.\n\nLocationAnnamite mountains, Vietnam and Laos.\n\nMethodsWe conducted systematic landscape-scale surveys across five study sites (four protected areas, one unprotected area) using camera-trapping and leech-derived environmental DNA. We analyzed detections within a Bayesian multi-species occupancy framework to evaluate species responses to environmental and anthropogenic influences. Species responses were then used to predict occurrence to unsampled regions. We used predicted species richness maps and occurrence of endemic species to identify areas of conservation importance for targeted conservation interventions.\n\nResultsAnalyses showed that habitat-based covariates were uninformative. Our final model therefore incorporated three anthropogenic covariates as well as elevation, which reflects both ecological and anthropogenic factors. Conservation-priority species tended to found in areas that are more remote now or have been less accessible in the past, and at higher elevations. Predicted species richness was low and broadly similar across the sites, but slightly higher in the more remote site. Occupancy of the three endemic species showed a similar trend.\n\nMain conclusionIdentifying spatial patterns of biodiversity in heavily-defaunated landscapes may require novel methodological and analytical approaches. Our results indicate to build robust prediction maps it is beneficial to sample over large spatial scales, use multiple detection methods to increase detections for rare species, include anthropogenic covariates that capture different aspects of hunting pressure, and analyze data within a Bayesian multi-species framework. Our models further suggest that more remote areas should be prioritized for anti-poaching efforts to prevent the loss of rare and endemic species.

ecology

Scalable probabilistic PCA for large-scale genetic variation data

Principal component analysis (PCA) is a key tool for understanding population structure and controlling for population stratification in genome-wide association studies (GWAS). With the advent of large-scale datasets of genetic variation, there is a need for methods that can compute principal components (PCs) with scalable computational and memory requirements. We present ProPCA, a highly scalable method based on a probabilistic generative model, which computes the top PCs on genetic variation data efficiently. We applied ProPCA to compute the top five PCs on genotype data from the UK Biobank, consisting of 488,363 individuals and 146,671 SNPs, in less than thirty minutes. Leveraging the population structure inferred by ProPCA within the White British individuals in the UK Biobank, we scanned for SNPs that are not well-explained by the PCs to identify several novel genome-wide signals of recent putative selection including missense mutations in RPGRIP1L and TLR4.\n\nAuthor SummaryPrincipal component analysis is a commonly used technique for understanding population structure and genetic variation. With the advent of large-scale datasets that contain the genetic information of hundreds of thousands of individuals, there is a need for methods that can compute principal components (PCs) with scalable computational and memory requirements. In this study, we present ProPCA, a highly scalable statistical method to compute genetic PCs efficiently. We systematically evaluate the accuracy and robustness of our method on large-scale simulated data and apply it to the UK Biobank. Leveraging the population structure inferred by ProPCA within the White British individuals in the UK Biobank, we identify several novel signals of putative recent selection.

bioinformatics