Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

Engineering targeted deletions in the mitochondrial genome

Summary ParagraphMitochondria are a network of critical intracellular organelles with diverse functions ranging from energy production to cell signaling. The mitochondrial genome (mtDNA) consists of 37 genes that support oxidative phosphorylation and are prone to dysfunction that can lead to currently untreatable diseases. Further characterization of mtDNA gene function and creation of more accurate models of human disease will require the ability to engineer precise genomic sequence modifications. To date, mtDNA has been inaccessible to direct modification using traditional genome engineering tools due to unique DNA repair contexts in mitochondria1. Here, we report a new DNA modification process using sequence-specific transcription activator-like effector (TALE) proteins to manipulate mtDNA in vivo and in vitro for reverse genetics applications. First, we show mtDNA deletions can be induced in Danio rerio (zebrafish) using site-directed mitoTALE-nickases (mito-nickases). Using this approach, the protein-encoding mtDNA gene nd4 was deleted in injected zebrafish embryos. Furthermore, this DNA engineering system recreated a large deletion spanning from nd5 to atp8, which is commonly found in human diseases like Kearns-Sayre syndrome (KSS) and Pearson syndrome. Enrichment of mtDNA-deleted genomes was achieved using targeted mitoTALE-nucleases (mitoTALENs) by co-delivering both mito-nickases and mitoTALENs into zebrafish embryos. This combined approach yielded deletions in over 90% of injected animals, which were maintained through adulthood in various tissues. Subsequently, we confirmed that large, targeted deletions could be induced with this approach in human cells. In addition, we show that, when provided with a single nick on the mtDNA light strand, the binding of a terminal TALE protein alone at the intended recombination site is sufficient for deletion induction. This \"block and nick\" approach yielded engineered mitochondrial molecules with single nucleotide precision using two different targeted deletion sites. This precise seeding method to engineer mtDNA variants is a critical step for the exploration of mtDNA function and for creating new cellular and animal models of mitochondrial disease.

molecular biology

Genome sequencing and assessment of plant growth-promoting properties of a Serratia marcescens strain isolated from vermicompost

Plant-bacteria associations have been extensively studied for their potential in increasing crop productivity in a sustainable manner. Serratia marcescens is a Gram-negative species found in a wide range of environments, including soil. Here we describe the genome sequencing and assessment of plant-growth promoting abilities of S. marcescens UENF-22GI (SMU), a strain isolated from mature cattle manure vermicompost. In vitro, SMU is able to solubilize P and Zn, to produce indole compounds (likely IAA), to colonize hyphae and counter the growth of two phytopathogenic fungi. Inoculation of maize with SMU remarkably increased seedling growth and biomass under greenhouse conditions. The SMU genome has 5 Mb, assembled in 17 scaffolds comprising 4,662 genes (4,528 are protein-coding). No plasmids were identified. SMU is phylogenetically placed within a clade comprised almost exclusively of environmental strains. We were able to find the genes and operons that are likely responsible for all the interesting plant-growth promoting features that were experimentally described. Genes involved other interesting properties that were not experimentally tested (e.g. tolerance against metal contamination) were also identified. The SMU genome harbors a horizontally-transferred genomic island involved in antibiotic production, antibiotic resistance, and anti-phage defense via a novel ADP-ribosyltransferase-like protein and possible modification of DNA by a deazapurine base, which likely contributes to the SMU competitiveness against other bacteria. Collectively, our results suggest that S. marcescens UENF-22GI is a strong candidate to be used in the enrichment of substrates for plant growth promotion or as part of bioinoculants for Agriculture.

microbiology

Accurate Filtering of Privacy-Sensitive Information in Raw Genomic Data

Sequencing thousands of human genomes has enabled breakthroughs in many areas, among them precision medicine, the study of rare diseases, and forensics. However, mass collection of such sensitive data entails enormous risks if not protected to the highest standards. In this article, we follow the position and argue that post-alignment privacy is not enough and that data should be automatically protected as early as possible in the genomics workflow, ideally immediately after the data is produced. We show that a previous approach for filtering short reads cannot extend to long reads and present a novel filtering approach that classifies raw genomic data (i.e., whose location and content is not yet determined) into privacy-sensitive (i.e., more affected by a successful privacy attack) and non-privacy-sensitive information. Such a classification allows the fine-grained and automated adjustment of protective measures to mitigate the possible consequences of exposure, in particular when relying on public clouds. We present the first filter that can be indistinctly applied to reads of any length, i.e., making it usable with any recent or future sequencing technologies. The filter is accurate, in the sense that it detects all known sensitive nucleotides except those located in highly variable regions (less than 10 nucleotides remain undetected per genome instead of 100,000 in previous works). It has far less false positives than previously known methods (10% instead of 60%) and can detect sensitive nucleotides despite sequencing errors (86% detected instead of 56% with 2% of mutations). Finally, practical experiments demonstrate high performance, both in terms of throughput and memory consumption.

bioinformatics

Classification of early and late stage Liver Hepatocellular Carcinoma patients from their genomics and epigenomics profiles.

BackgroundLiver Hepatocellular Carcinoma (LIHC) is the second major cancer worldwide, responsible for millions of premature deaths every year. Prediction of clinical staging is vital to implement optimal therapeutic strategy and prognostic prediction in cancer patients. However, to date, no method has been developed for predicting stage of LIHC from genomic profile of samples.\n\nResultsIn current study, in silico models have been developed for classifying LIHC patients in early and late stage using RNA expression and DNA methylation data. The Cancer Genome Atlas (TCGA) dataset contains 173 early and 177 late stage samples of LIHC, was extensively analysed to identify differentially expressed RNA transcripts and methylated CpG sites that can discriminate early and late stages of LIHC samples with high precision. Naive Bayes model developed using 51 features that combine 21 CpG methylation sites and 30 RNA transcripts achieved maximum MCC 0.58 with accuracy 78.87% on validation dataset. Further, we also analysed genomics and epigenomics profiles of normal and LIHC samples and developed model to classify LIHC samples with AUROC 0.99. In addition, multiclass models developed for classifying samples in normal, early and late stage of cancer and achieved accuracy of 76.54% and AUROC of 0.86.\n\nConclusionOur study reveals stage prediction of LIHC samples with high accuracy based on genomics and epigenomics profiling is a challenging task in comparison to classification of LIHC and normal samples. Comprehensive analysis, differentially expressed RNA transcripts, methylated CpG sites in LIHC samples and prediction models are available from CancerLSP (http://webs.iiitd.edu.in/raghava/cancerlsp/).

bioinformatics

Comparative Genome Analysis Reveals Important Genetic Factors Associated with Probiotic Property in Enterococcus faecium strains

Enterococcus faecium though commensals in human gut, few strains provide beneficial effect to humans as probiotics, few are responsible for nosocomial infection and few as non-pathogens. Comparative genomics of E. faecium will help to reveal the genomic differences responsible for the said properties. In this study, we compared E. faecium strain 17OM39 with a marketed probiotic, non-pathogenic non-probiotic (NPNP) and pathogenic strains. The core genome analysis revealed, 17OM39 was closely related with marketed probiotic strain T110. Strain 17OM39 was found to be devoid of known vancomycin, tetracycline resistance genes and functional virulence genes. Moreover, 17OM39 is \"less open\" due to absence of frequently found transposable elements. Genes imparting beneficial functional properties were observed to be present in marketed probiotic T110 and 17OM39 strains. Additional, genes associated with colonization within gastrointestinal tract were detected across all the strains. Beyond shared genetic features; this study particularly identified genes that are responsible to impart probiotic, non-pathogenic and pathogenic features to the strains of E. faecium. The study also provides insights into the acquired and intrinsic drug resistance genes, which will be helpful for better understanding of the physiology of antibiotic resistance in E. faecium strains. In addition, we could identify genes contributing to the intrinsic ability of 17OM39 E. faecium isolate to be a potential probiotic.\n\nThe study has comprehensively characterized genome sequence of each strain to find the genetic variation and understand effects of these on functionality, phenotypic complexity. Further the evolutionary relationship of species along with adaptation strategies have been including in this study.

bioinformatics

Improvements to the rice genome annotation through large-scale analysis of RNA-Seq and proteomics datasets

Rice (Oryza sativa) is one of the most important worldwide crops. The genome has been available for over 10 years and has undergone several rounds of annotation. We created a comprehensive database of transcripts from 29 public RNA sequencing datasets, officially predicted genes from Ensembl plants, and common contaminants in which to search for protein-level evidence. We re-analysed nine publicly accessible rice proteomics datasets. In total, we identified 420K peptide spectrum matches from 47K peptides and 8,187 protein groups. 4168 peptides were initially classed as putative novel peptides (not matching official genes). Following a strict filtration scheme to rule out other possible explanations, we discovered 1,584 high confidence novel peptides. The novel peptides were clustered into 692 genomic loci where our results suggest annotation improvements. 80% of the novel peptides had an ortholog match in the curated protein sequence set from at least one other plant species. For the peptides clustering in intergenic regions (and thus potentially new genes), 101 loci were identified, for which 43 had a high-confidence hit for a protein domain. Our results can be displayed as tracks on the Ensembl genome or other browsers supporting Track Hubs, to support re-annotation of the rice genome.

plant biology

FORGe: prioritizing variants for graph genomes

There is growing interest in using genetic variants to augment the reference genome into a \"graph genome\" to improve read alignment accuracy and reduce allelic bias. While adding a variant has the positive effect of removing an undesirable alignment-score penalty, it also increases both the ambiguity of the reference genome and the cost of storing and querying the genome index. We introduce methods and a software tool called FORGe for modeling these effects and prioritizing variants accordingly. We show that FORGe enables a range of advantageous and measurable trade-offs between accuracy and computational overhead.

bioinformatics

Skyhawk: An Artificial Neural Network-based discriminator for reviewing clinically significant genomic variants

MotivationMany rare diseases and cancers are fundamentally diseases of the genome. In the past several years, genome sequencing has become one of the most important tools in clinical practice for rare disease diagnosis and targeted cancer therapy. However, variant interpretation remains the bottleneck as is not yet automated and may take a specialist several hours of work per patient. On average, one-fifth of this time is spent on visually confirming the authenticity of the candidate variants. ResultsWe developed Skyhawk, an artificial neural network-based discriminator that mimics the process of expert review on clinically significant genomics variants. Skyhawk runs in less than one minute to review ten thousand variants, and about 30 minutes to review all variants in a typical whole-genome sequencing sample. Among the false positive singletons identified by GATK HaplotypeCaller, UnifiedGenotyper and 16GT in the HG005 GIAB sample, 79.7% were rejected by Skyhawk. Worked on the Variants with Unknown Significance (VUS), Skyhawk marked most of the false positive variants for manual review and most of the true positive variants no need for review. AvailabilitySkyhawk is easy to use and freely available at https://github.com/aquaskyline/Skyhawk

bioinformatics

Combining landscape genomics and ecological modelling to investigate local adaptation of indigenous Ugandan cattle to East Coast fever

East Coast fever (ECF) is a fatal sickness affecting cattle populations of eastern, central, and southern Africa. The disease is transmitted by the tick Rhipicephalus appendiculatus, and caused by the protozoan Theileria parva parva, which invades host lymphocytes and promotes their clonal expansion. Importantly, indigenous cattle show tolerance to infection in ECF-endemically stable areas. Here, the putative genetic bases underlying ECF-tolerance were investigated using molecular data and epidemiological information from 823 indigenous cattle from Uganda. Vector distribution and host infection risk were estimated over the study area and subsequently tested as triggers of local adaptation by means of landscape genomics analysis. We identified 41 and seven candidate adaptive loci for tick resistance and infection tolerance, respectively. Among the genes associated with the candidate adaptive loci are PRKG1 and SLA2. PRKG1 was already described as associated with tick resistance in indigenous South African cattle, due to its role into inflammatory response. SLA2 is part of the regulatory pathways involved into lymphocytes proliferation. Additionally, local ancestry analysis suggested the zebuine origin of the genomic region candidate for tick resistance.\n\nAuthor summaryThe tick-borne parasite Theileria parva parva infects cattle populations of eastern, central and southern Africa, by causing a highly fatal pathology called \"East Coast fever\". The disease is especially severe for the exotic breeds imported to Africa, as well as outside the endemic areas of East Africa. In these regions, indigenous cattle populations can survive to infection, and this tolerance might result from unique adaptations evolved to fight the disease. We investigated this hypothesis by using a method named \"landscape genomics\", with which we compared the genetic characteristics of indigenous Ugandan cattle coming from areas at different infection risk, and located genomic sites potentially attributable to tolerance. In particular, the method pinpointed two genes, one (PRKG1) involved into inflammatory response and potentially affecting East Coast fever vector attachment, the other (SLA2) involved into lymphocytes proliferation, a process activated by T. parva parva infection. Our findings can orientate future research on the genetic basis of East Coast fever-tolerance, and derive from a general method that can be applied to investigate adaptation in analogous host-vector-parasite systems. Characterization of the genetic factors underlying East Coast-fever-tolerance represents an essential step towards enhancing sustainability and productivity of local agroecosystems.

evolutionary biology

Kaposi’s sarcoma-associated herpesvirus ORF68 is a DNA binding protein required for viral genome cleavage and packaging

Herpesviral DNA packaging into nascent capsids requires multiple conserved viral proteins that coordinate genome encapsidation. Here, we investigated the role of the ORF68 protein of Kaposis sarcoma-associated herpesvirus (KSHV), a protein required for viral DNA encapsidation whose function remains largely unresolved across the herpesviridae. We found that KSHV ORF68 is expressed with early kinetics and localizes predominantly to viral replication compartments, although it is dispensable for viral DNA replication and gene expression. However, in agreement with its proposed role in viral DNA packaging, KSHV-infected cells lacking ORF68 failed to cleave viral DNA concatemers, accumulated exclusively immature B-capsids, and released no infectious progeny virions. ORF68 has no predicted domains aside from a series of putative zinc finger motifs. However, in vitro biochemical analyses of purified ORF68 protein revealed that it robustly binds DNA and is associated with nuclease activity. These activities provide new insights into the role of KSHV ORF68 in viral genome encapsidation.\n\nImportanceKaposis sarcoma-associated herpesvirus (KSHV) is the etiologic agent of Kaposis sarcoma and several B-cell cancers, causing significant morbidity and mortality in immunocompromised individuals. A critical step in the production of infectious viral progeny is the packaging of the newly replicated viral DNA genome into the capsid, which involves coordination between at least seven herpesviral proteins. While the majority of these packaging factors have been well studied in related herpesviruses, the role of the KSHV ORF68 protein and its homologs remains unresolved. Here, using a KSHV mutant lacking ORF68, we confirm its requirement for viral DNA processing and packaging in infected cells. Furthermore, we show that the purified ORF68 protein directly binds DNA and is associated with a metal-dependent cleavage activity on double stranded DNA in vitro. These activities suggest a novel role for ORF68 in herpesviral genome processing and encapsidation.

molecular biology

Learning the sequence of influenza A genome assembly during viral replication using point process models and fluorescence in situ hybridization

Within influenza virus infected cells, viral genomic RNA are selectively packed into progeny virions, which predominantly contain a single copy of 8 viral RNA segments. Intersegmental RNA-RNA interactions are thought to mediate selective packaging of each viral ribonucleoprotein complex (vRNP). Clear evidence of a specific interaction network culminating in the full genomic set has yet to be identified. Using multi-color fluorescence in situ hybridization to visualize four vRNP segments within a single cell, we developed image-based models of vRNP-vRNP spatial dependence. These models were used to construct likely sequences of vRNP associations resulting in the full genomic set. Our results support the notion that selective packaging occurs during cytoplasmic transport and identifies the formation of multiple distinct vRNP sub-complexes that likely form as intermediate steps toward full genomic inclusion into a progeny virion. The methods employed demonstrate a statistically driven, model based approach applicable to other interaction and assembly problems. Author SummaryInfluenza virus consists of eight viral ribonucleoproteins (vRNPs) that are assembled by infected cells to produce new virions. The process by which all eight vRNPs are assembled is not yet understood. We therefore used images from a previous study in which up to four vRNPs had been visualized in the same cell to construct spatial point process models that measure how well the subcellular distribution of one vRNP can be predicted from one or more other vRNPs. We used the likelihood of these models as an estimate of the extent of association between vRNPs and thereby constructed likely sequences of vRNP assembly that would produce full virions. Our work identifies the formation of multiple distinct vRNP sub-complexes that likely form as intermediate steps toward production of a virion. The results may be of use in designing strategies to interfere with virus assembly. We also anticipate that the approach may be useful for studying other assembly processes, especially for complexes with modest affinities and more components than can be visualized simultaneously.

bioinformatics

Full Bayesian comparative phylogeography from genomic data

A challenge to understanding biological diversification is accounting for community-scale processes that cause multiple, co-distributed lineages to co-speciate. Such processes predict non-independent, temporally clustered divergences across taxa. Approximate-likelihood Bayesian computation (ABC) approaches to inferring such patterns from comparative genetic data are very sensitive to prior assumptions and often biased toward estimating shared divergences. We introduce a full-likelihood Bayesian approach, ecoevolity, which takes full advantage of information in genomic data. By analytically integrating over gene trees, we are able to directly calculate the likelihood of the population history from genomic data, and efficiently sample the model-averaged posterior via Markov chain Monte Carlo algorithms. Using simulations, we find that the new method is much more accurate and precise at estimating the number and timing of divergence events across pairs of populations than existing approximate-likelihood methods. Our full Bayesian approach also requires several orders of magnitude less computational time than existing ABC approaches. We find that despite assuming unlinked characters (e.g., unlinked single-nucleotide polymorphisms), the new method performs better if this assumption is violated in order to retain the constant characters of whole linked loci. In fact, retaining constant characters allows the new method to robustly estimate the correct number of divergence events with high posterior probability in the face of character-acquisition biases, which commonly plague loci assembled from reduced-representation genomic libraries. We apply our method to genomic data from four pairs of insular populations of Gekko lizards from the Philippines that are not expected to have co-diverged. Despite all four pairs diverging very recently, our method strongly supports that they diverged independently, and these results are robust to very disparate prior assumptions.

evolutionary biology

Genome-centric metagenomics revealed the spatial distribution and the diverse metabolic functions of lignocellulose degrading uncultured bacteria

The mechanisms by which specific anaerobic microorganisms remain firmly attached to lignocellulosic material allowing them to efficiently decompose the organic matter are far to be elucidated. To circumvent this issue, the microbiomes collected from anaerobic digesters treating pig manure and meadow grass were fractionated to separate the planktonic microbes from those adhered to lignocellulosic substrate. Assembly of shotgun reads followed by binning process recovered 151 population genomes, 80 out of which were completely new and were not previously deposited in any database. Genome coverage allowed the identification of microbial spatial distribution into the engineered ecosystem. Moreover, a composite bioinformatic analysis using multiple databases for functional annotation revealed that uncultured members of Bacteroidetes and Firmicutes follow diverse metabolic strategies for polysaccharide degradation. The structure of cellulosome in Firmicutes can vary depending on the number and functional roles of carbohydrate-binding modules. On contrary, members of Bacteroidetes are able to adhere and degrade lignocellulose due to the presence of multiple carbohydrate-binding family 6 modules in beta-xylosidase and endoglucanase proteins or S-layer homology modules in unknown proteins. This study combines the concept of variability in spatial distribution with genome-centric metagenomics allowing a functional and taxonomical exploration of the biogas microbiome.\n\nImportanceThis work contributes new knowledge about lignocellulose degradation in engineered ecosystems. Specifically, the combination of the spatial distribution of uncultured microbes with genome-centric metagenomics provides novel insights into the metabolic properties of planktonic and firmly attached to plant biomass bacteria. Moreover, the knowledge obtained in this study enabled us to understand the diverse metabolic strategies for polysaccharide degradation in different species of Bacteroidetes and Clostridiales. Even though structural elements of cellulosome were restricted to Clostridiales, our study identified in Bacteroidetes a putative mechanism for biomass decomposition based on a gene cluster responsible for cellulose degradation, disaccharide cleavage to glucose and transport to cytoplasm.

microbiology

Identifying Plasmids in Bacterial Genome Assemblies

Despite their importance to bacterial pathogenesis, plasmids are rarely identified in incomplete genome sequences. The free FindPlasmids (FP) package for Windows, Mac OS X, and Linux facilitates identification of plasmids in incomplete genome sequences. FP found plasmids in 98.8% of complete genomes in which they were present, correctly identifying plasmids ranging in number from 1 to 10 plasmids. In a sample of 50 E. coli genome assemblies it has identified from zero to three plasmids ranging in size from 1,549 to 133,843 bp and present in 42% of the assemblies examined.

microbiology

VARAN-GIE: Curation of Genomic Interval Sets

Genomic interval sets are fundamental elements of genome annotation and are the output of countless bioinformatics applications. Nevertheless, tool support for the manual curation of these data is currently limited. We developed VARAN-GIE, an extension of the popular Integrative Genomics Viewer (IGV) that adds functionality to edit, annotate and merge genomic interval sets. Data can easily be shared with other users and imported/exported from/to multiple common data formats. VARAN-GIE binary releases, source-code, user guides and tutorials are available at https://github.com/popitsch/varan-gie/

bioinformatics

Removal of alleles by genome editing -- RAGE against the deleterious load

BackgroundIn this paper, we simulate deleterious load in an animal breeding program, and compare the efficiency of genome editing and selection for decreasing load. Deleterious variants can be identified by bioinformatics screening methods that use sequence conservation and biological prior information about protein function. Once deleterious variants have been identified, how can they be used in breeding?\n\nResultsWe simulated a closed animal breeding population subject to both natural selection against deleterious load and artificial selection for a quantitative trait representing the breeding goal. Deleterious load was polygenic and due to either codominant or recessive variants. We compared strategies for removal of deleterious alleles by genome editing (RAGE) to selection against carriers. Each strategy varied in how animals and variants were prioritized for editing or selection.\n\nConclusionsGenome editing of deleterious alleles reduces deleterious load, but requires simultaneous editing of multiple deleterious variants in the same sire to be effective when deleterious variants are recessive. In the short term, selection against carriers is a possible alternative to genome editing when variants are recessive. The dominance of deleterious variants affects both the efficiency of genome editing and selection against carriers, and which variant prioritization strategy is the most efficient. Our results suggest that in the future, there is the potential to use RAGE against deleterious load in animal breeding.

genetics

POPSICLE: A Software Suite to Study Population Structure and Ancestral Determinants of Phenotypes using Whole Genome Sequencing Data

The advent of new sequencing technologies has provided access to genome-wide markers which may be evaluated for their association with phenotypes. Recent studies have leveraged these technologies and sequenced hundreds and sometimes thousands of strains to improve the accuracy of genotype-phenotype predictions. Sequencing of thousands of strains is not practical for many research groups which argues for the formulation of new strategies to improve predictability using lower sample sizes and more cost-effective methods. We introduce here a novel computational algorithm called POPSICLE that leverages the local genetic variations to infer blocks of shared ancestries to construct complex evolutionary relationships. These evolutionary relationships are subsequently visualized using chromosome painting, as admixtures and as clades to acquire general as well as specific ancestral relationships within a population. In addition, POPSICLE evaluates the ancestral blocks for their association with phenotypes thereby bridging two powerful methodologies from population genetics and genome-wide association studies. In comparison to existing tools, POPSICLE offers substantial improvements in terms of accuracy, speed and automation. We evaluated POPSICLEs ability to find genetic determinants of Artemisinin resistance within P. falciparum using 57 randomly selected strains, out of 1,612 that were used in the original study. POPSICLE found Kelch, a gene implicated in the original study, to be significant (p-value 0) towards resistance to Artemisinin. We further extended this analysis to find shared ancestries among closely related P. falciparum, P. reichenowi and P. gaboni species from the Laverania subgenus of Plasmodium. POPSICLE was able to accurately infer the population structure of the Laverania subgenus and detected 4 strains from a chimpanzee in Koulamoutou with significant shared ancestries with P. falciparum and P. gaboni. We simulated 4 datasets to asses if these shared ancestries indicated a hybrid or mixed infections involving P. falciparum and P. gaboni. The analysis based on the simulated data and genome-wide heterozygosity profiles of the strains indicate these are most likely mixed infections although the possibility of hybrids cannot be ruled out. POPSICLE is a java-based utility that requires no installation and can be downloaded freely from https://popsicle-admixture.sourceforge.io/\n\nAuthor SummaryThe associations between genotypes and phenotypes have traditionally been performed using markers such as single nucleotide polymorphisms. Often, these markers are independently evaluated for their association with phenotypes. A genomic region is deemed significant if multiple markers with significance colocalize. However, multiple markers that are in linkage disequilibrium can sometimes work synergistically and contribute to phenotypic variations. These synergistic associations across markers and across subpopulations have traditionally been captured by population genetic approaches that determine local ancestries. We sought to bridge these two powerful but independent methodologies to improve genotype-phenotype predictions. We developed a new software called POPSICLE that employs an innovative approach to determine local ancestries and evaluates them for their association with phenotypes. Validity of POPSICLE in determining the genes that are responsible for Plasmodium Falciparums resistance to Artemisinin and in determining the population structure of Laverania subgenus of Plasmodium are discussed.

bioinformatics

A fast mrMLM algorithm for multi-locus genome-wide association studies

BackgroundRecent developments in technology result in the generation of big data. In genome-wide association studies (GWAS), we can get tens of million SNPs that need to be tested for association with a trait of interest. Indeed, this poses a great computational challenge. There is a need for developing fast algorithms in GWAS methodologies. These algorithms must ensure high power in QTN detection, high accuracy in QTN estimation and low false positive rate.\n\nResultsHere, we accelerated mrMLM algorithm by using GEMMA idea, matrix transformations and identities. The target functions and derivatives in vector/matrix forms for each marker scanning are transformed into some simple forms that are easy and efficient to evaluate during each optimization step. All potentially associated QTNs with P-values [≤] 0.01 are evaluated in a multi-locus model by LARS algorithm and/or EM-Empirical Bayes. We call the algorithm FASTmrMLM. Numerical simulation studies and real data analysis validated the FASTmrMLM. FASTmrMLM reduces the running time in mrMLM by more than 50%. FASTmrMLM also shows high statistical power in QTN detection, high accuracy in QTN estimation and low false positive rate as compared to GEMMA, FarmCPU and mrMLM. Real data analysis shows that FASTmrMLM was able to detect more previously reported genes than all the other methods: GEMMA/EMMA, FarmCPU and mrMLM.\n\nConclusionsFASTmrMLM is a fast and reliable algorithm in multi-locus GWAS and ensures high statistical power, high accuracy of estimates and low false positive rate.\n\nAuthor SummaryThe current developments in technology result in the generation of a vast amount of data. In genome-wide association studies, we can get tens of million markers that need to be tested for association with a trait of interest. Due to the computational challenge faced, we developed a fast algorithm for genome-wide association studies. Our approach is a two stage method. In the first step, we used matrix transformations and identities to quicken the testing of each random marker effect. The target functions and derivatives which are in vector/matrix forms for each marker scanning are transformed into some simple forms that are easy and efficient to evaluate during each optimization step. In the second step, we selected all potentially associated SNPs and evaluated them in a multi-locus model. From simulation studies, our algorithm significantly reduces the computing time. The new method also shows high statistical power in detecting significant markers, high accuracy in marker effect estimation and low false positive rate. We also used the new method to identify relevant genes in real data analysis. We recommend our approach as a fast and reliable method for carrying out a multi-locus genome-wide association study.

bioinformatics