Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Multiscale analysis of structurally conserved motifs

This work develops a generic framework to perform a multiscale structural analysis of two structures (homologous proteins, conformations) undergoing conformational changes. Practically, given a seed structural alignment, we identify structural motifs with a hierarchical structure, characterized by three unique properties. First, the hierarchical structure sheds light on the trade-off between size and flexibility. Second, motifs can be combined to perform an overall comparison of the input structures in terms of combined RMSD - an improvement over the classical least RMSD. Third, motifs can be used to seed iterative aligners, and to design hybrid sequence-structure profile HMM characterizing protein families.\n\nFrom the methods standpoint, our framework is reminiscent from the bootstrap and combines concepts from rigidity analysis (distance difference matrices), graph theory, computational geometry (space filling diagrams), and topology (topological persistence).\n\nOn challenging cases (class II fusion proteins, flexible molecules) we illustrate the ability of our tools to localize conformational changes, shedding light of commonalities of structures which would otherwise appear as radically different.\n\nOur tools are available within the Structural Bioinformatics Library (http://sbl.inria.fr). We anticipate that they will be of interest to perform structural comparisons at large, and for remote homology detection.

bioinformatics

Hybrid sequence-structure based HMM models leverage theidentification of homologous proteins: the example of class II fusion

We present a sequence-structure based method characterizing a set of functionally related proteins exhibiting low sequence identity and loose structural conservation. Given a (small) set of structures, our method consists of three main steps. First, pairwise structural alignments are combined with multi-scale geometric analysis to produce structural motifs i.e. regions structurally more conserved than the whole structures. Second, the sub-sequences of the motifs are used to build profile hidden Markov models (HMM) biased towards the structurally conserved regions. Third, these HMM are used to retrieve from UniProtKB proteins harboring signatures compatible with the function studied, in a bootstrap fashion.\n\nWe apply these hybrid HMM to investigate two questions related to class II fusion proteins, an especially challenging class since known structures exhibit low sequence identity (less than 15%) and loose structural similarity (of the order of 15[A] in lRMSD). In a first step, we compare the performances of our hybrid HMM against those of sequence based HMM. Using various learning sets, we show that both classes of HMM retrieve unique species. The number of unique species reported by both classes of methods are comparable, stressing the novelty brought by our hybrid models. In a second step, we use our models to identify 17 plausible HAP2-GSC1 candidate sequences in 10 different drosophila melanogaster species. These models are not identified by the PF[A]M family HAP2-GCS1 (PF10699), stressing the ability of our structural motifs to capture signals more subtle than whole Pfam domains.\n\nIn a more general setting, our method should be of interest for all cases functional families with low sequence identity and loose structural conservation.\n\nOur software tools are available from the FunChaT package of the Structural Bioinformatics Library (http://sbl.inria.fr).

bioinformatics

Lemon: a modern C++ tool for the rapid development of structural benchmarking datasets

MotivationThe protein data bank (PDB) currently holds over 140,000 biomolecular structures and continues to release new structures on a weekly basis. The PDB is an essential resource to the structural bioinformatics community to develop software that mine, use, categorize, and analyze such data. New computational biology methods are evaluated using custom benchmarking sets derived as subsets of 3D experimentally determined structures and structural features from the PDB. Currently, such benchmarking features are manually curated with custom scripts in a non-standardized manner that results in slow distribution and updates with new experimental structures. Finally, there is a scarcity of standardized tools to rapidly query 3D descriptors of the entire PDB. ApproachOur solution is the Lemon framework, a C++11 library with Python bindings, which provides a consistent workflow methodology for selecting biomolecular interactions based on user criterion and computing desired 3D structural features. This framework can parse and characterize the entire PDB in less than ten minutes on modern, multithreaded hardware. The speed in parsing is obtained by using the recently developed MacroMolecule Transmission Format (MMTF) to reduce the computational cost of reading text-based PDB files. The use of C++ lambda functions and Python binds provide extensive flexibility for analysis and categorization of the PDB by allowing the user to write custom functions to suite their objective. We think Lemon will become a one-stop-shop to quickly mine the entire PDB to generate desired structural biology features. The Lemon software is available as a C++ header library along with example functions at https://github.com/chopralab/lemon.

bioinformatics

Performance of gene expression analyses using de novo assembled transcripts in polyploid species

MotivationQuality of gene expression analyses using de novo assembled transcripts in species experienced recent polyploidization is yet unexplored.\n\nResultsFive plant species with various polyploidy history were used for differential gene expression (DGE) analyses. DGE analyses using putative genes inferred by Trinity performed similar to or better than Corset and Grouper in precision, but lower in sensitivity. In species that lack polyploidy event in the past few million years, DGE analyses using de novo assembled transcriptome identified 50-76% of the differentially expressed genes recovered by mapping reads to the reference genes. However, in species with more recent polyploidy event, the percentage decreased to 7-30%. In addition, 7-89% of differentially expressed genes from de novo assembly are contaminations. Gene co-expression network analyses using de novo assemblies vs. mapping to the reference genes recovered the same module that significantly correlated with treatment in one of the five species tested.\n\nAvailability and ImplementationCommands and scripts used in this study are available at https://bitbucket.org/lychen83/chen_et_al_2018_benchmark_dge/; Analysis files are available at Dryad doi: XXXXXX.\n\nContactlychen83@qq.com\n\nSupplementary informationSupplementary data are available at Bioinformatics online

bioinformatics

On the Depth of Deep Learning Models for Splice Site Identification

The success of deep learning has been shown in various fields including computer vision, speech recognition, natural language processing and bioinformatics. The advance of Deep Learning in Computer Vision has been an important source of inspiration for other research fields. The objective of this work is to adapt known deep learning models borrowed from computer vision such as VGGNet, Resnet and AlexNet for the classification of biological sequences. In particular, we are interested by the task of splice site identification based on raw DNA sequences. We focus on the role of model architecture depth on model training and classification performance.\n\nWe show that deep learning models outperform traditional classification methods (SVM, Random Forests, and Logistic Regression) for large training sets of raw DNA sequences. Three model families are analyzed in this work namely VGGNet, AlexNet and ResNet. Three depth levels are defined for each model family. The models are benchmarked using the following metrics: Area Under ROC curve (AUC), Number of model parameters, number of floating operations. Our extensive experimental evaluation show that shallow architectures have an overall better performance than deep models. We introduced a shallow version of ResNet, named S-ResNet. We show that it gives a good trade-off between model complexity and classification performance.\n\nAuthor summaryDeep Learning has been widely applied to various fields in research and industry. It has been also succesfully applied to genomics and in particular to splice site identification. We are interested in the use of advanced neural networks borrowed from computer vision. We explored well-known models and their usability for the problem of splice site identification from raw sequences. Our extensive experimental analysis shows that shallow models outperform deep models. We introduce a new model called S-ResNet, which gives a good trade-off between computational complexity and classification accuracy.

bioinformatics

FASTCAR: Rapid alignment-free prediction of sequence alignment identity scores

MotivationPairwise alignment is a predominant algorithm in the field of bioinformatics. This algorithm is quadratic -- slow especially on long sequences. Many applications utilize identity scores without the corresponding alignments. For these applications, we propose FASTCAR. It produces identity scores for pairs of DNA sequences using alignment-free methods and two self-supervised general linear models. ResultsFor the first time, the new tool can predict the pair-wise identity score in linear time and space. On two large-scale sequence databases, FASTCAR provided the best compromise between sensitivity and precision while being faster than BLAST by 40% and faster than USEARCH by 6-10 times. Further, FASTCAR is capable of producing the pair-wise identity scores of long DNA sequences -- millions-of-nucleotides-long bacterial genomes; this task cannot be accomplished by any alignment-based tool. AvailabilityFASTCAR is available at https://github.com/TulsaBioinformaticsToolsmith/FASTCAR and as the Supplementary Dataset 1. Contacthani-girgis@utulsa.edu Supplementary informationSupplementary data are available online.

bioinformatics

FactorialHMM: Fast and exact inference in factorial hidden Markov models

MotivationHidden Markov models (HMMs) are powerful tools for modeling processes along the genome. In a standard genomic HMM, observations are drawn, at each genomic position, from a distribution whose parameters depend on a hidden state; the hidden states evolve along the genome as a Markov chain. Often, the hidden state is the Cartesian product of multiple processes, each evolving independently along the genome. Inference in these so-called Factorial HMMs has a naive running time that scales as the square of the number of possible states, which by itself increases exponentially with the number of subchains; such a running time scaling is impractical for many applications. While faster algorithms exist, there is no available implementation suitable for developing bioinformatics applications.\n\nResultsWe developed FactorialHMM, a Python package for fast exact inference in Factorial HMMs. Our package allows simulating either directly from the model or from the posterior distribution of states given the observations. Additionally, we allow the inference of all key quantities related to HMMs: (1) the (Viterbi) sequence of states with the highest posterior probability; (2) the likelihood of the data; and (3) the posterior probability (given all observations) of the marginal and pairwise state probabilities. The running time and space requirement of all procedures is linearithmic in the number of possible states. Our package is highly modular, providing the user with maximal flexibility for developing downstream applications.\n\nAvailabilityhttps://github.com/regevs/factorialhmm

bioinformatics

Encircling the regions of the pharmacogenomic landscape that determine drug response

The integration of large-scale drug sensitivity screens and genome-wide experiments is changing the field of pharmacogenomics, revealing molecular determinants of drug response without the need for a priori, hypothesis-driven assumptions about drug action. In particular, transcriptomic signatures of drug sensitivity may guide drug repositioning, the discovery of synergistic drug combinations and suggest new therapeutic biomarkers. However, the inherent complexity of transcriptomic signatures, with thousands of genes differentially expressed, makes them hard to interpret, giving poor mechanistic insights and hampering translation to the clinics. Here we show how network biology can help simplify transcriptomic drug signatures, filtering out irrelevant genes, accounting for tissue-specific biases and ultimately yielding functionally-coherent, less noisy drug modules. We successfully analyzed 170 drugs tested in 637 cancer cell lines, proving a broad applicability of our approach and evincing an intimate relationship between modules gene expression levels and drugs mechanisms of action. Further, we have characterized multiple aspects of our transcriptomic modules. As a result, the drugs included in this study are now annotated well beyond the reductionist (target-centered) view.\n\nAuthor SummaryLarge scale pharmacogenomics studies performed with hundreds of cell lines offer a means to link the molecular features of the cells to their response to drug treatments. Unfortunately, simple drug-gene correlations are usually not enough to consistently identify what gene expression patterns will determine drug sensitivity, as the tissue of origin of the cells, together with the expression of e.g. membrane transporters can greatly confound the analysis. To ameliorate these biases, we have devised a network-based strategy that selects genes that are both well correlated to drug response and closeby in the human protein interaction network. Reassuringly, we have confirmed that our identified drug sensitivity modules are tightly connected to the mechanisms of action of the drugs. Moreover, while our modules have no more than 100 genes, they retain the predictive power of the much larger gene signatures that are typically obtained by drug-gene correlations alone. Here, we release the characteristic modules for almost 200 drugs, in a format that is suitable for downstream bioinformatics analyses such as gene-set enrichment analysis.

bioinformatics

rawMSA: Proper Deep Learning makes protein sequence profiles and feature extraction obsolete

In the last few decades, huge efforts have been made in the bioinformatics community to develop machine learning-based methods for the prediction of structural features of proteins in the hope of answering fundamental questions about the way proteins function and about their involvement in several illnesses. The recent advent of Deep Learning has renewed the interest in neural networks, with dozens of methods being developed in the hope of taking advantage of these new architectures. On the other hand, most methods are still based on heavy pre-processing of the input data, as well as the extraction and integration of multiple hand-picked, manually designed features. Since Multiple Sequence Alignments (MSA) are almost always the main source of information in de novo prediction methods, it should be possible to develop Deep Networks to automatically refine the data and extract useful features from it. In this work, we propose a new paradigm for the prediction of protein structural features called rawMSA. The core idea behind rawMSA is borrowed from the field of natural language processing to map amino acid sequences into an adaptively learned continuous space. This allows the whole MSA to be input into a Deep Network, thus rendering sequence profiles and other pre-calculated features obsolete. We developed rawMSA in three different flavors to predict secondary structure, relative solvent accessibility and inter-residue contact maps. We have rigorously trained and benchmarked rawMSA on a large set of proteins and have determined that it outperforms classical methods based on position-specific scoring matrices (PSSM) when predicting secondary structure and solvent accessibility, while performing on a par with the top ranked CASP12 methods in the inter-residue contact map prediction category. We believe that rawMSA represents a promising, more powerful approach to protein structure prediction that could replace older methods based on protein profiles in the coming years.\n\nAvailabilitydatasets, dataset generation code, evaluation code and models are available at: https://bitbucket.org/clami66/rawmsa

bioinformatics

In silico REPOSITIONING OF APPROVED DRUGS AGAINST Schistosoma mansoni ENERGY METABOLISM TARGETS

Schistosomiasis is a neglected parasitosis caused by Schistosoma spp. Praziquantel is used for the chemoprophylaxis and treatment of this disease. Although this monotherapy is effective, the risk of resistance and its low efficiency against immature worms compromises its effectiveness. Therefore, it is necessary to develop new schistosomicide drugs. However, the development of new drugs is a long and expensive process. The repositioning of approved drugs has been proposed as a quick, cheap, and effective alternative to solve this problem. This study employs chemogenomic analysis with use of bioinformatics tools to search, identify, and analyze data on approved drugs with the potential to inhibit Schistosoma mansoni energy metabolism enzymes. The TDR Targets Database, Gene DB, Protein, DrugBank, Therapeutic Targets Database (TTD), Promiscuous, and PubMed databases were used. Fifty-nine target proteins were identified, of which 18 had one or more approved drugs. The results identified 20 potential drugs for schistosomiasis treatment, all approved for use in humans.

bioinformatics

HTSlib-js: Developer toolkit for a client-side JavaScript interface for HTSlib

BackgroundAn increasing number of bioinformatics tools are developed in JavaScript to provide an interactive visual interface to users. However, there are no tools yet that enable fast client-side analysis of large, high-throughput sequencing file formats, such as BAM and VCF files.\n\nResultsWe present HTSlib-js, an asm.js-based JavaScript wrapper for HTSlib, the de facto standard for processing BAM and VCF files. HTSlib-js exploits recent technological advances in web browser engines that dramatically increase the performance of browser-based tools to enable swift processing of files in the aforementioned formats. HTSlib-js enables quick development of JavaScript-based applications that include processing of aligned sequence reads and variant calling data.\n\nConclusionsHTSlib-js constitutes a toolkit for developers to easily write fast, accessible, browser-based applications that place an emphasis on data visualization, interactivity and privacy. Real-world examples demonstrate the capabilities of HTSlib-js and serve as guide for own developments.

bioinformatics

Integrated modeling of peptide digestion and detection for the prediction of proteotypic peptides in targeted proteomics

MotivationThe selection of proteotypic peptides, i.e., detectable unique representatives of proteins of interest, is a key step in targeted shotgun proteomics. To date, much effort has been made to predict proteotypic peptides in the absence of mass spectrometry data. However, the performance of existing tools is still unsatisfactory. One crucial reason is their neglect of the close relationship between protein proteolytic digestion and peptide detection.\n\nResultsWe present an algorithm (named AP3) that firstly considers peptide digestion probability as a feature for proteotypic peptide prediction and demonstrated peptide digestion probability is the most important feature for accurate prediction of proteotypic peptides. AP3 showed higher accuracy than existing tools and accurately predicted the proteotypic peptides for a targeted proteomics assay, showing its great potential for assisting the design of targeted proteomics experiments.\n\nAvailability and ImplementationFreely available at http://fugroup.amss.ac.cn/software/AP3/AP3.html.\n\nContactyfu@amss.ac.cn or zhuyunping@gmail.com\n\nSupplementary InformationSupplementary data are available at Bioinformatics online.

bioinformatics

Learning Protein Structural Fingerprints under the Label-Free Supervision of Domain Knowledge

Finding homologous proteins is the indispensable first step in many protein biology studies. Thus, building highly efficient \"search engines\" for protein databases is a highly desired function in protein bioinformatics. As of August 2018, there are more than 140,000 protein structures in PDB, and this number is still increasing rapidly. Such a big number introduces a big challenge for scanning the whole structure database with high speeds and high sensitivities at the same time. Unfortunately, classic sequence alignment tools and pairwise structure alignment tools are either not sensitive enough to remote homologous proteins (with low sequence identities) or not fast enough for the task. Therefore, specifically designed computational methods are required for quickly scanning structure databases for homologous proteins.\n\nHere, we propose a novel ContactLib-DNN method to quickly scan structure databases for homologous proteins. The core idea is to build structure fingerprints for proteins, and to perform alignment-free comparisons with the fingerprints. Specifically, the fingerprints are low-dimensional vectors representing the contact groups within the proteins. Notably, the Cartesian distance between two fingerprint vectors well matches the RMSD between the two corresponding contact groups. This is done by using RMSD as the domain knowledge to supervise the deep neural network learning. When comparing to existing methods, ContactLib-DNN achieves the highest average AUROC of 0.959. Moreover, the best candidate found by ContactLib-DNN has a probability of 70.0% to be a true positive. This is a significant improvement over 56.2%, the best result produced by existing methods.\n\nGitHub: https://github.com/Chenyao2333/contactlib/\n\nIndex Termshomologous proteins, protein structures, remote protein homolog detection, alignment-free comparisons

bioinformatics

RAFTS3G - An efficient and versatile clustering software to analyses in large protein datasets

The need to develop computational tools and techniques that can predict efficiently consistent groups of family proteins in large volume of biological information is still a great perspective in Bioinformatic studies. Besides that, it is difficult to increase speed demanding low computational processing to minimize the information complexity. Tools already consolidated as the CD-HIT and UCLUST generates very compact data that makes the Data Mining difficult and have low efficiency when used for detect homology among proteins requiring manual intervention, therefore it is necessary a tool that is also efficient in low similarity. Here we present a new approach for the Data Mining and analysis of homology in large dataset of protein sequences, the RAFTS3G. We used the UniProtKB/Swiss-Prot database with the most popular clustering tools and RAFTS3G proved to be more than 10 times faster than CD-HIT and its strategy increases the performance in low similarity to detect protein families.\n\nContact: raittz@ufpr.br

bioinformatics

Prediction of Deleterious Single Nucleotide Polymorphisms in Human p53 Gene

With a variety of accessible Single Nucleotide Polymorphisms (SNPs) data on human p53 gene, this investigation is intended to deal with detrimental SNPs in p53 gene by executing diverse valid computational tools, including Filter, SIFT, PredictSNP, Fathmm, UTRScan, ConSurf, Phyre, Tm-Adjust, I-Mutant, Task Seek after practical and basic appraisal, dissolvable openness, atomic progression, and analysing the energy minimization. Of 581 p53 SNPs, 420 SNPs are found to be missense or non-synonymous and 435 SNPs are in the 3 prime UTR and 112 SNPs are of every 5 prime UTR from which 16 non synonymous SNPs (nsSNPs) as non-tolerable while PredictSNP package predicted 14 (taking consideration SNP colored green by two or more than 2 analyses is neutral). By concentrating on six bioinformatics tools of various dimensions a combined output is generated where 14 nsSNPs are prone to exert a deleterious effect. By using diverse SNP analysing tools we have found 5 missense SNPs in the 3 crucial amino acids position in the DNA binding domain. The underlying discoveries are fortified by I-Mutant and Project HOPE. The ExPASy-PROSITE tools characterized whether the mutations located in the functional part of the protein or not. This study provides a decisive outcome concluding the accessible SNPs information by recognizing the five harming nsSNPs: rs28934573 (S241F), rs11540652 (R248Q), rs121913342 (R248W), rs121913343 (R273C) and rs28934576 (R273H). The findings of this investigation recognize the detrimental nsSNPs which enhance the danger of various kinds of oncogenesis in patients of different populations in genome-wide studies (GWS).

bioinformatics

SCDT: Detecting CNVs of low chimeric ratio in cf-DNA

MotivationSequencing of cell-free DNA (cf-DNA) has enabled Noninvasive Prenatal Testing (NIPT) and\"liquid biopsy\" of cancers. However, while the aneuploidy and point mutations were focused on by most of NITP and liquid biopsy studies, detecting sub-chromosome CNVs that affect a few to dozens of megabases was rarely reported, likely attributable to the difficulty in accurately identifying them, especially for those present in a small fraction of cf-DNA.\n\nResultsWe developed a somatic CNV detection tool (SCDT), for detecting sub-chromosome CNVs in cf-DNA using whole genome sequencing (WGS) data or off-target reads in target sequencing data. Additional to using control samples for correcting genome position specific bias, two GC correction steps were performed, which regressed GC content of DNA fragments and that of genome bins, respectively. After GC correction, the coefficients of variation of copy ratios approximated the lower boundary of theoretical values, suggesting removing of almost all systematic errors. Finally, CNVs were detected by a piecewise least squares fitting based segmentation algorithm, which outperformed other segmentation methods. We applied SCDT on simulated and real maternal plasma samples, and target cf-DNA sequencing of 118 normal individuals and 240 cancer patients, and demonstrated high sensitivity and specificity.\n\nAvailabilitySCDT is available at https://github.com/Martiantian/Somatic_cnv_detect_tool.\n\nContactzhuhongmei@genomics.cn\n\nSupplementary InformationSupplementary data are available at Bioinformatics online

bioinformatics

eQTL network analysis reveals that regulatory genes are evolutionarily older and bearing more types of PTM sites in Coprinopsis cinerea

Understanding the DNA variation in regulation of carbohydrate-active enzymes (CAZymes) is fundamental to the use of wood-decaying basidiomycetes in lignocellulose conversion into renewable energy. Our goal is to identify the regulators of lignocellulolytic enzymes in Coprinopsis cinerea, of which the genome harbors high number of Auxiliary Activities enzymes.\n\nThe DNA sequence of C. cinerea family including 46 single spore isolates (SSIs) from crosses of two homozygous strains are used to develop a panel of SNP markers. Then the RNA sequence were used to characterize the gene expression profiles. The RNA were extracted from cultures grown on softwood-enriched sawdust to induce lignocellulolytic enzymes and CCR de-repression genes. To assess the genetic contribution to enzyme expression variations among the 46 SSIs, associations between SNPs and gene expressions were examined genome-widely. 5148 local eQTLs and 7738 distant eQTLs were obtained. By analyzing these eQTLs, the potential regulatory factors of the CAZymes expression and the de-repression of Carbon Catabolism Repression (CCR) were identified.\n\nThe eQTL network is characterized in terms of hotspots, evolutionary age and post-translational modifications (PTMs). In the eQTL network of C. cinerea, the non-regulatory genes are younger than the regulatory genes. The proteins regulated by combinational multiple types of PTMs are more likely to function as super regulatory hotspots in protein-protein interactions. The evolutionary age analysis and the PTMome analysis could serve as alternative methods to identify master regulators from genomic data.\n\nThis work demonstrates a comprehensive bioinformatics approach to identify regulatory factors with next-generation sequencing data. The results provide candidate genes for bioengineering to increase the enzyme production, which will practically benefit the bioethanol production from lignocellulose.\n\nSignificanceThis eQTL analysis is designed to study the fungal CAZymes and carbon catabolism repression, especially during the mycelium stage.\n\nO_LIIn Coprinopsis cinerea, only the regions near two ends of the chromosomes have high recombination rate, and suitable for family based eQTL analysis.\nC_LIO_LIA sugar transporter is a hotspot controlling many CCR genes.\nC_LIO_LICAZymes are not regulated by a master regulator, but by individual regulators. This indicates that CAZymes are under specific regulatory pathways, so can response to specific conditions.\nC_LIO_LIIn the eQTL network, the rGenes are evolutionarily older, with more types of PTM sites than eGenes.\nC_LIO_LIIn the eQTL network, the proteins with more types of PTM sites are more likely associated with Information Storage and Processing, and act as super-hub in the network.\nC_LI

bioinformatics

Integrated Analysis Revealed Hub Genes in Breast Cancer

The aim of this study was to identify the hub genes in breast cancer and provide further insight into the tumorigenesis and development of breast cancer. To explore the hub genes in breast cancer, we performed an integrated bioinformatics analysis. Two gene expression profiles were downloaded from the GEO database. The differentially expressed genes (DEGs) were identified by using the \"limma\" package. Then, we performed Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis to explore the functional annotation and potential pathways of the DEGs. Next, protein-protein interaction (PPI) network analysis and weighted gene coexpression network analysis (WGCNA) were conducted to screen for hub genes. To confirm the reliability of the identified hub genes, we obtained TCGA-BRCA data by using WGCNA to screen for genes that were strongly related to breast cancer. By combining the results from the GEO and TCGA datasets, we finally identified 15 real hub genes in breast cancer. Finally, we performed an overall survival analysis to explore the connection between the expression of hub genes and the overall survival time of breast cancer patients. We found that for all hub genes, higher expression was associated with significantly shorter overall survival times among breast cancer patients.

bioinformatics