Search bioRxivSearch

Biology subjects

Xia, X.

Publications and source records attributed to Xia, X..

14 recordsLinked to original sources

A starless bias in the maximum likelihood phylogenetic methods

I analyzed various site pattern combinations in a 4-OTU case to identify sources of starless bias and parameter-estimation bias in likelihood-based phylogenetic methods, and reported three significant contributions. First, the likelihood method is odd in that it may not generate a star tree with sequences that are equidistant from each other. This behaviour, dubbed starless bias, happens in a 4-OTU tree when there is an excess (i.e., more than expected from a star tree and a substitution model) of conflicting phylogenetic signals supporting the three resolved topologies equally. Special site pattern combinations leading to rejection of a star tree, when sequences are equidistant from each other, were identified. Second, fitting gamma distribution to model rate heterogeneity over sites is strongly confounded with tree topology, especially in conjunction with the starless bias. I present examples to show dramatic differences in the estimated shape parameter between a star tree and a resolved tree. There may be no rate heterogeneity over sites (with the estimated > 10000) when a star tree is imposed, but < 1 (suggesting strong rate heterogeneity over sites) when an (incorrect) resolved tree is imposed. Thus, the dependence of "rate heterogeneity" on tree topology implies that "rate heterogeneity" is not a sequence-specific feature, cautioning against interpreting a small to mean that some sites are under strong purifying selection and others not. Thirdly, because there is no existing (and working) likelihood method for evaluating a star tree with continuous gamma-distributed rate, I have implemented the method for JC69 in a self-contained R script for a four-OTU tree (star or resolved), in addition to another R script assuming a constant rate over sites. These R scripts should be useful for teaching and exploring likelihood methods in phylogenetics.

bioinformatics

Comprehensive proteomic and metabolomic profiling of mcr-1 mediated colistin resistance in Escherichia coli

The spread of mcr-1 in human and veterinary medicine has jeopardized the use of polymyxins, the last-resort antibiotics against life-threatening multidrug-resistant Gram-negative bacteria. As a lipid-modified gene, whether mcr-1 brings proteomic and metabolomic changes in the bacteria and affects the corresponding metabolic pathway is largely unknown. Herein, we used label-free quantitative proteomics and untargeted metabolomics to profile comprehensive proteome and metabolome characteristics of mcr-1-mediated colistin-resistant and -sensitive Escherichia coli and further insight the resistant mechanism of colistin. We identified large sets of differential expression proteins and metabolites that contributed to mcr-1-mediated antibiotic resistance predominantly in the different growth conditions with and without colistin. mcr-1 could cause the down-regulated expression of most proteins to adapt drug pressure. Pathway analysis showed that metabolic process was significantly affected, mainly related to glycerophospholipid metabolism, thiamine metabolism, and lipopolysaccharide biosynthesis. The substrate phosphatidylethanolamine for mcr-1 to mediate colistin resistance is accumulated in colistin-resistant E. coli. Notably, mcr-1 can not only cause the phosphoethanolamine modification of bacterial cell membrane lipid A, but also affect the biosynthesis and transport of lipoprotein in colistin resistance through disturbing the expression of efflux pump proteins involved in cationic antibacterial peptide resistance pathway. Overall, the disturbed glycerophospholipid metabolism, lipopolysaccharide biosynthesis and the accumulation of the substrate phosphatidylethanolamine is closely related with mcr-1-mediated colistin resistance and these findings can further provide valuable information to inhibit colistin resistance by blocking this metabolic process.

pharmacology and toxicology

Wheat avenin-like protein and its significant Fusarium Head Blight resistant functions

Wheat Avenin-like proteins (TaALP) are atypical storage proteins belonging to the Prolamin superfamily. Previous studies on ALPs have focused on the proteins positive effects on dough strength, whilst no correlation has been made between TaALPs and the plant immune system. Here, we performed genome-wide characterization of ALP encoding genes in bread wheat. In silico analyses indicated the presence of critical peptides in TaALPs that are active in the plant immune system. Pathogenesis-related nucleotide motifs were also identified in the putative promoter regions of TaALP encoding genes. RT-PCR was performed on TaALP and previously characterised pathogenesis resistance genes in developing wheat caryopses under control and Fusarium graminearum infection conditions. The results showed that TaALP and NMT genes were upregulated upon F. graminearum inoculation. mRNA insitu hybridization showed that TaALP genes were expressed in the embryo, aleurone and sub-aleurone layer cells. Seven TaALP genes were cloned for the expression of recombinant proteins in Escherichia coli, which displayed significant inhibitory function on F. graminearum under anti-fungal tests. In addition, FHB index association analyses showed that allelic variations of two ALP genes on chromosome 7A were significantly correlated with FHB symptoms. Over-expression of an ALP gene on chromosome 7A showed an enhanced resistance to FHB. Yeast two Hybridization results revealed that ALPs have potential proteases inhibiting effect on metacaspases and beta-glucosidases. A vital infection process related pathogen protein, F. graminearum Beta-glucosidase was found to interact with ALPs. Our study is the first to report a class of wheat storage protein or gluten protein with biochemical functions. Due to its abundance in the grain and the important multi-functions, the results obtained in the current study are expected to have a significant impact on wheat research and industry.

molecular biology

The wheat Sr22, Sr33, Sr35 and Sr45 genes confer resistance against stem rust in barley

In the last 20 years, stem rust caused by the fungus Puccinia graminis f. sp. tritici (Pgt), has re-emerged as a major threat to wheat and barley cultivation in Africa and Europe. In contrast to wheat with 82 designated stem rust (Sr) resistance genes, barleys genetic variation for stem rust resistance is very narrow with only seven resistance genes genetically identified. Of these, only one locus consisting of two genes is effective against Ug99, a strain of Pgt which emerged in Uganda in 1999 and has since spread to much of East Africa and parts of the Middle East. The objective of this study was to assess the functionality, in barley, of cloned wheat Sr genes effective against Ug99. Sr22, Sr33, Sr35 and Sr45 were transformed into barley cv. Golden Promise using Agrobacterium-mediated transformation. All four genes were found to confer effective stem rust resistance. The barley transgenics remained susceptible to the barley leaf rust pathogen Puccinia hordei, indicating that the resistance conferred by these wheat Sr genes was specific for Pgt. Cloned Sr genes from wheat are therefore a potential source of resistance against wheat stem rust in barley.

plant biology

Extensive Expansion of the Speedy gene Family in Homininae and Functional Differentiation in Humans

BackgroundThe cell cycle plays important roles in physiology and disease. The Speedy/RINGO family of atypical cyclins regulates the cell cycle. However, the origin, evolution and function of the Speedy family are not completely understood. Understanding the origins and evolution of Speedy family would shed lights on the evolution of complexity of cell cycles in eukaryotes.\n\nResultsHere, we performed a comprehensive identification of Speedy genes in 258 eukaryotic species and found that the Speedy subfamily E was extensively expanded in Homininae, characterized by emergence of a low-Spy1-identify domain. Furthermore, the Speedy gene family show functional differentiation in humans and have a distinct expression pattern, different regulation network and co-expressed gene networks associated with cell cycle and various signaling pathways. Expression levels of the Speedy gene family are prognostic biomarkers among different cancer types.\n\nConclusionsOverall, we present a comprehensive view of the Speedy genes and highlight their potential function.

evolutionary biology

High-throughput identification and marker development of perfect SSR for cultivated genus of passion fruit (Passiflora edulis)

Simple sequence repeat (SSR) markers are characterized by high polymorphism, good reproducibility and co-dominance etc. They can be easily applied to develop efficient, simple and practical molecular markers. In the present study, bioinformatics methods were applied to identify high-throughput perfect SSRs of cultivar Passiflora genome. A total of 13104 perfect SSRs were obtained. SSR core sequence structure is mainly 2-4 bases, the maximum numbers are TA, AT, TC and AG. The maximum numbers of repetitions were up to 20 times. A total of 12934 pairs of SSR markers were developed by using bioinformatics software, and 20 pairs of markers were selected for amplification specificity assessment of MTX and WJ10, and the polymorphism rate was as high as 60%. The large-scale development of the SSR markers of Passiflora cultivar has paved a foundation for the efficient utilization of the germplasm resources of passion fruit, genetic improvement of the varieties and molecular breeding.

genomics

BK channel inhibition by strong extracellular acidification

BK-type voltage- and Ca2+-dependent K+ channels are found in a wide range of cells and intracellular organelles. Among different loci, the composition of the extracellular microenvironment, including pH, may differ substantially. For example, it has been reported that BK channels are expressed in lysosomes with their extracellular side facing the strongly acidified lysosomal lumen (pH ~ 4.5). Here we show that BK activation is strongly and reversibly inhibited by extracellular H+, with its conductance-voltage relationship shifted by more than +100 mV at pHO 4. Our results reveal that this inhibition is mainly caused by H+ inhibition of BK voltage-sensor (VSD) activation through three acidic residues on the extracellular side of BK VSD. Given that these key residues (D133, D147, D153) are highly conserved among members in the voltage-dependent cation channel superfamily, the mechanism underlying BK inhibition by extracellular acidification might also be applicable to other members in the family.

biophysics

Transcriptome Landscape of Human Oocytes and Granulosa Cells Throughout Folliculogenesis

Folliculogenesis is a highly regulated process that involves bidirectional interactions of the oocytes and surrounding granulosa cells (GCs). Little is unknown, however, about the transcriptomic profiles of human oocytes and GCs throughout folliculogenesis. Here we performed a high resolution RNA-Seq of human oocytes and GCs at each follicular stage, which revealed unique transcriptional profiles, stage-specific signature genes, oocyte- and GC-derived genes that reflect ovarian reserve. We identified reciprocal cell-to-cell interactions between oocytes and GCs, including NOTCH, TGF-{beta} signaling and gap junctions and determined the expression patterns of maternal-effect genes involved in folliculogenesis and early embryogenesis. Finally, we demonstrated robust differences between human and mice oocyte transcriptomes. This is the first comprehensive overview of the transcriptomic signatures governing the stepwise human folliculogenesis in-vivo that provides a valuable resource for basic and translational research in human reproductive biology.

cell biology

An improved method for fitting gamma distribution to substitution rate variation among sites

Gamma distribution has been used to fit substitution rate variation over site. One simple method to estimate the shape parameter of the gamma distribution is to 1) reconstruct a phylogenetic tree and the ancestral states of internal nodes, 2) perform pairwise comparison between nodes on each side of each branch to count the number of \"observed\" substitutions for each site, and apply correction of multiple hits to derive the estimated number of substitutions for each site, and 3) fit the site-specific substitution data to gamma distribution to obtain the shape parameter This method is fast but its accuracy depends much on the accuracy of the estimated site-specific number of substitutions. The existing method has three shortcomings. First, it uses Poisson correction which is inadequate for almost any nucleotide sequences. Second, it does independent estimation for the number of substitutions at each site without making use of information at all sites. Third, the program implementing the method has never been made publically available. I have implemented in DAMBE software a new method based on the F84 substitution model with simultaneous estimation that uses information from all sites in estimating the number of substitutions at each site. DAMBE is freely available at available at http://dambe.bio.uottawa.ca

bioinformatics

Imputing missing distances in molecular phylogenetics

Missing data are frequently encountered in molecular phylogenetics and need to be imputed. For a distance matrix with missing distances, the least-squares approach is often used for imputing the missing values. Here I develop a method, similar to the expectation-maximization algorithm, to impute multiple missing distance in a distance matrix. I show that, for inferring the best tree and missing distances, the minimum evolution criterion is not as desirable as the least-squares criterion. I also discuss the problem involving cases where the missing values cannot be uniquely determined, e.g., when a missing distance involve two sister taxa. The new method has the advantage over the existing one in that it does not assume a molecular clock. I have implemented the function in DAMBE software which is freely available at available at http://dambe.bio.uottawa.ca

bioinformatics

Identification of candidate genes for gelatinization temperature, gel consistency and pericarp color by GWAS in rice based on SLAF-sequencing

Rice is an important cereal in the world, uncovering the genetic basis of agronomic traits in rice landraces genes associated with agronomically important traits is indispensable for both understanding the genetic basis of phenotypic variation and efficient crop improvement. Gelatinization temperature, gel consistency and pericarp color are important indices of rice cooking and eating quality evaluation and potential nutritional importance, which attract wide attentions in the application of genetic and breeding. To dissect the genetic basis of gelatinization temperature (GT), gel consistency (GC) and pericarp color (PC), a total of 419 rice landraces core germplasm collections consisting of 330 indica lines, 78 japonica lines and 11 uncertain varieties were grown, collected, then GT, GC, PC were measured for two years, and sequenced using Specific Locus Amplified Fragment Sequencing (SLAF) technology. In this study, 261,385,070 clean reads and 56,768 polymorphic SLAF tags were obtained, which a total of 211,818 single nucleotide polymorphisms (SNPs) were discovered. With 208,993 SNPs meeting the criterion of minor allele frequency (MAF) > 0.05 and integrity> 0.5, the phylogenetic tree and population structure analysis were performed for all 419 rice landraces, and the whole panel mainly separated into six subpopulations based on population structure analysis. Genome-wide association study (GWAS) was carried out for the whole panel, indica subpanel and japonica subpanel with subset SNPs respectively. One quantitative trait locus (QTL) on chromosome 6 for GT was detected in the whole panel and indica subpanel, and one QTL associated with GC was located on chromosome 6 in the whole panel and indica subpanel. For the PC trait, 8 QTLs were detected in the whole panel on chromosome 1, 3, 4, 7, 8, 10 and 11, and 7 QTLs in the indica subpanel on chromosome 3, 4, 7, 8, 10 and 11. The loci on chromosome 3, 8, 10 and 11 have not been identified previously, and they may be the candidate genes of pericarp color. For the three traits, no QTL was detected in japonica subpanel probably because of the polymorphism repartition between the subpanel, or small population size of japonica subpanel. This paper provides new gene resources and insights into the molecular mechanisms of important agricultural trait of rice phenotypic variation and genetic improvement of rice quality variety breeding.

genomics

Genetic Diversity and Distributional Pattern of Ammonia Oxidizing Archaea Lineages in the Global Oceans

In the study, we used miTAG approach to analyse the distributional pattern of the ammonium oxidizing archaea (AOA) lineages in the global oceans using the metagenomics datasets of the Tara Oceans global expedition (2009-2013). Using ammonium monooxygenase alpha subunit gene as biomarker, the AOA communities were obviously segregated with water depth, except the upwelling regions. Besides, the AOA communities in the euphotic zones are more heterogeneous than in the mesopelagic zones (MPZs). Overall, water column A clade (WCA) distributes more evenly and widely in the euphotic zone and MPZs, while water column B clade (WCB) and SCM-like clade mainly distribute in MPZ and high latitude waters, respectively. At fine-scale genetic diversity, SCM1-like and 2 WCA subclades showed distinctive niche separation of distributional pattern. The AOA subclades were further divided into ecological significant taxonomic units (ESTUs), which were delineated from the distribution pattern of their corresponding subclades. For examples, ESTUs of WCA have different correlation with depth, nitrate to silicate ratio and salinity; SCM1-like-A was negatively correlated with irradiation; the other SCM-like ESTUs preferred low temperature and high nutrient conditions, etc. Our study provides new insight to the genetic diversity of AOA in global scale and its connections with environmental factors.

ecology

ARSDA: A new approach for storing, transmitting and analyzing high-throughput sequencing data

Two major stumbling blocks exist in high-throughput sequencing (HTS) data analysis. The first is the sheer file size typically in gigabytes when uncompressed, causing problems in storage, transmission and analysis. However, these files do not need to be so large and can be reduced without loss of information. Each HTS file, either in compressed .SRA or plain text .fastq format, contains numerous identical reads stored as separate entries. For example, among 44603541 forward reads in the SRR4011234.sra file (from a Bacillus subtilis transcriptomic study) deposited at NCBIs SRA database, one read has 497027 identical copies. Instead of storing them as separate entries, one can and should store them as a single entry with the SeqID_NumCopy format (which I dub as FASTA+ format). The second is the proper allocation reads that map equally well to paralogous genes. I illustrate in detail a new method for such allocation. I have developed ARSDA software that implement these new approaches. A number of HTS files for model species are in the process of being processed and deposited at http://coevol.rdc.uottawa.ca to demonstrate that this approach not only saves a huge amount of storage space and transmission bandwidth, but also dramatically reduces time in downstream data analysis. Instead of matching the 497027 identical reads separately against the Bacillus subtilis genome, one only needs to match it once. ARSDA includes functions to take advantage of HTS data in the new sequence format for downstream data analysis such as gene expression characterization. ARSDA can be run on Windows, Linux and Macintosh computers and is freely available at http://dambe.bio.uottawa.ca/ARSDA/ARSDA.aspx.

bioinformatics

Exploring the mutational robustness of nucleic acidsby searching genotype neighbourhoods in sequencespace

To assess the mutational robustness of nucleic acids, many genome- and protein-level studies have been performed; in these investigations, nucleic acids are treated as genetic information carriers and transferrers. However, the molecular mechanism through which mutations alter the structural, dynamic and functional properties of nucleic acids is poorly understood. Here, we performed SELEX in silico study to investigate the fitness distribution of the nucleic acid genotype neighborhood in a sequence space for L-Arm binding aptamer. Although most mutants of the L-Arm-binding aptamer failed to retain their ligand-binding ability, two novel functional genotype neighborhoods were isolated by SELEX in silico and experimentally verified to have similar binding affinity (Kd = 69.3 M and 110.7 M) as the wild-type aptamer (Kd = 114.4 M). Based on data from the current study and previous research, mutational robustness is strongly influenced by the local base environment and ligand-binding mode, whereas bases distant from the binding pocket provide potential evolutionary pathways to approach global fitness maximum. Our work provides an example of successful application of SELEX in silico to optimize an aptamer and demonstrates the strong sensitivity of mutational robustness to the site of genetic variation.

evolutionary biology