Search bioRxivSearch

Biology subjects

Gyllensten, U.

Publications and source records attributed to Gyllensten, U..

5 recordsLinked to original sources

Genome-wide CNV study and functional evaluation identified CTDSPL as tumour suppressor gene for cervical cancer

We have investigated copy number variations (CNVs) in relation to cervical cancer by analyzing 731,422 single-nucleotide polymorphisms (SNPs) in 1,034 cervical cancer cases and 3,948 controls, followed by replication in 1,396 cases and 1,057 controls. We found that a 6367bp deletion in intron 1 of the CTD small phosphatase like gene (CTDSPL) was associated with 2.54-fold increased risk of cervical cancer (odds ratio =2.54, 95% confidence interval =2.08-3.12, P=2.0x10-19). This CNV is one of the strongest genetic risk variants identified so far for cervical cancer. The deletion removes the binding sites of zinc finger protein 263, binding protein 2 and interferon regulatory factor 1, and hence downregulates the transcription of CTDSPL. HeLa cells expressing CTDSPL showed a significant decrease in colony-forming ability. Compared with control groups, mice injected with HeLa cells expressing CTDSPL exhibited a significant reduction in tumour volume. Furthermore, CTDSPL-depleted immortalized End1/E6E7 could form tumours in NOD-SCID mice.

cancer biology

High throughput proteomics identifies 484 high-accuracy plasma protein biomarker signatures for ovarian cancer

Ovarian cancer is usually detected at a late stage with the 5-year survival at only 30-40%. Additional means for early detection and improved diagnosis are acutely needed. To search for novel biomarkers, we compared circulating plasma levels of 981 proteins in patients with ovarian cancer and benign tumours, using the proximity extension assay. A novel combinatorial strategy was developed for identification of multivariate biomarker signatures, resulting in 484 mutually exclusive models out of which 448 did not contain the present biomarker MUCIN-16. The top-ranking model consisted of 14 proteins and had a AUC=0.95, PPV=1.0, sensitivity=0.99 and specificity=1.0 for detection of stage III-IV ovarian cancer in the discovery data, and an AUC=0.89, PPV=0.93, sensitivity=0.89 and specificity=0.95 in the replication data. The novel plasma protein signature could be used to improve the diagnosis of women with adnexal ovarian mass or in screening to identify women that should be referred to specialized examination.

cancer biology

De novo assembly of two Swedish genomes reveals missing segments from the human GRCh38 reference and improves variant calling of population-scale sequencing data

We have performed de novo assembly of two Swedish genomes using long-read sequencing and optical mapping, resulting in total assembly sizes of nearly 3 Gb and hybrid scaffold N50 values of over 45 Mb. A further analysis revealed over 10 Mb of sequences absent from the human GRCh38 reference in each individual. Around 6 Mb of these novel sequences (NS) are shared with a Chinese personal genome. The NS are highly repetitive, have elevated GC-content and are primarily located in centromeric or telomeric regions. A BLAST search showed that 31% of the NS are different from any sequences deposited in nucleotide databases. The remaining NS correspond to human (62%) or primate (6%) nucleotide entries, while 1% of hits show the highest similarity to other species, including mouse and a few different classes of parasitic worms. Up to 1 Mb of NS can be assigned to chromosome Y, and large segments are missing from GRCh38 also at chromosomes 14, 17 and 21. Inclusion of these novel sequences into the GRCh38 reference radically improves the alignment and variant calling of whole-genome sequencing data at several genomic loci. Through a re-analysis of 200 samples from a Swedish population-scale sequencing project, we obtained over 75,000 putative novel SNVs per individual when using a custom version of GRCh38 extended with 17.3 Mb of NS. In addition, about 10,000 false positive SNV calls per individual were removed from the GRCh38 autosomes and sex chromosomes in the re-analysis, with some of them located in protein coding regions.

genomics

Amplification-free, CRISPR-Cas9 Targeted Enrichment and SMRT Sequencing of Repeat-Expansion Disease Causative Genomic Regions

Targeted sequencing has proven to be an economical means of obtaining sequence information for one or more defined regions of a larger genome. However, most target enrichment methods require amplification. Some genomic regions, such as those with extreme GC content and repetitive sequences, are recalcitrant to faithful amplification. Yet, many human genetic disorders are caused by repeat expansions, including difficult to sequence tandem repeats.\n\nWe have developed a novel, amplification-free enrichment technique that employs the CRISPR-Cas9 system for specific targeting multiple genomic loci. This method, in conjunction with long reads generated through Single Molecule, Real-Time (SMRT) sequencing and unbiased coverage, enables enrichment and sequencing of complex genomic regions that cannot be investigated with other technologies. Using human genomic DNA samples, we demonstrate successful targeting of causative loci for Huntingtons disease (HTT; CAG repeat), Fragile X syndrome (FMR1; CGG repeat), amyotrophic lateral sclerosis (ALS) and frontotemporal dementia (C9orf72; GGGGCC repeat), and spinocerebellar ataxia type 10 (SCA10) (ATXN10; variable ATTCT repeat). The method, amenable to multiplexing across multiple genomic loci, uses an amplification-free approach that facilitates the isolation of hundreds of individual on-target molecules in a single SMRT Cell and accurate sequencing through long repeat stretches, regardless of extreme GC percent or sequence complexity content. Our novel targeted sequencing method opens new doors to genomic analyses independent of PCR amplification that will facilitate the study of repeat expansion disorders.

genomics

SweGen: A whole-genome map of genetic variability in a cross-section of the Swedish population

Here we describe the SweGen dataset, a high-quality map of genetic variation in the Swedish population. This data represents a basic resource for clinical genetics laboratories as well as for sequencing-based association studies, by providing information on the frequencies of genetic variants in a cohort that is well matched to national patient cohorts. To select samples for this study, we first examined the genetic structure of the Swedish population using high-density SNP-array data from a nation-wide population based cohort of over 10,000 individuals. From this sample collection, 1,000 individuals, reflecting a cross-section of the population and capturing the main genetic structure, were selected for whole genome sequencing (WGS). Analysis pipelines were developed for automated alignment, variant calling and quality control of the sequencing data. This resulted in a whole-genome map of aggregated variant frequencies in the Swedish population that we hereby release to the scientific community.

genetics