Search bioRxivSearch

Biology subjects

Smith, M. L.

Publications and source records attributed to Smith, M. L..

6 recordsLinked to original sources

The humungous fungus of Michigan three decades on

In the late 1980s, a genetic individual of the fungus Armillaria gallica that extended over at least 37 hectares of forest floor and encompassed hundreds of tree root systems was discovered on the Upper Peninsula of Michigan. Based on observed growth rates, the individual was estimated to be at least 1500 years old with a mass of more than 105 kg. Nearly three decades on, we returned to the site of individual for new sampling. We report here that the same genetic individual of A. gallica is still alive on its original site, but we estimated that it is older and larger than originally estimated, at least 2,500 years and 4 x 105 kg, respectively. We also show that mutation has occurred within the somatic cells of the individual, reflecting its historical pattern of growth from a single point. The overall rate of mutation, however, was extremely low. The large individual of A. gallica has been remarkably resistant to genomic change as it has persisted in place.

evolutionary biology

Tensional homeostasis in multicellular clusters: effects of geometry and traction force dynamics

The ability of cells to maintain a constant level of their cytoskeletal tension in response to external and internal disturbances is referred to as tensional homeostasis. It is essential for the normal physiological function of cells and tissues, and for protection against disease progression, including atherosclerosis and cancer. It has been shown recently that some cell types, such as endothelial cells, can maintain tensional homeostasis only when they form multicellular clusters, whereas other cell types, such as fibroblasts, do not require clustering for tensional homeostasis. For example, measurements of cell-extracellular matrix traction forces have shown that temporal fluctuations of the traction field in clusters of endothelial cells become progressively attenuated with increasing number of cells in the cluster, whereas in fibroblasts cell clustering does not influence traction field variability. Mechanisms that are responsible for these observations are largely unknown. In this study, a theoretical analysis and mathematical modeling have been applied to analyze experimental data obtained previously from traction microscopy measurements in order to investigate possible physical mechanisms that influence temporal variability of the traction field in multicellular forms. The focus of the analysis is on the contribution of dynamics and distribution of focal adhesion traction forces in conjunction with geometrical shape and size of multicellular clusters. Results of the analysis revealed that cluster size, magnitude and temporal fluctuations of focal adhesion traction forces have a major influence on traction field variability, whereas the influence of cluster shape appears to be minor.

biophysics

Disentangling the process of speciation using machine learning

Historically, investigations into the processes driving speciation have largely been isolated from systematic investigations into species limits. Recent advances in sequencing technology have led to a rapid increase in the availability of genomic data, and this, in turn, has led to the introduction of many novel methods for species delimitation. However, these methods have been limited to divergence-only scenarios and have not attempted to evaluate complex modes of speciation, such as those that include gene flow during early stages of divergence (sympatric speciation) or population size changes (founder effect speciation). To address this shortcoming, we introduce delimitR, an approach that enables biologists to infer species boundaries and evaluate the demographic processes that may have led to speciation. delimitR uses the binned multidimensional Site Frequency Spectrum and a machine-learning algorithm (Random Forests) to compare speciation models. We use simulations to evaluate the accuracy of delimitR. When comparing models that include lineage divergence and gene flow for three populations, error rates are near zero with recent divergence times (<100,000 generations) and a modest number of Single Nucleotide Polymorphisms (SNPs; 1,500). When applied to a more complex model set (including divergence, gene flow, and population size changes), error rates are moderate (~0.15 with 10,000 SNPs), and misclassifications are generally between highly similar models. We also evaluate the utility of delimitR using three previously published datasets and find results that corroborate previous findings. Our analyses indicate that delimitR can serve as an important conceptual bridge uniting various investigations into the process of speciation.

evolutionary biology

Amplification-free, CRISPR-Cas9 Targeted Enrichment and SMRT Sequencing of Repeat-Expansion Disease Causative Genomic Regions

Targeted sequencing has proven to be an economical means of obtaining sequence information for one or more defined regions of a larger genome. However, most target enrichment methods require amplification. Some genomic regions, such as those with extreme GC content and repetitive sequences, are recalcitrant to faithful amplification. Yet, many human genetic disorders are caused by repeat expansions, including difficult to sequence tandem repeats.\n\nWe have developed a novel, amplification-free enrichment technique that employs the CRISPR-Cas9 system for specific targeting multiple genomic loci. This method, in conjunction with long reads generated through Single Molecule, Real-Time (SMRT) sequencing and unbiased coverage, enables enrichment and sequencing of complex genomic regions that cannot be investigated with other technologies. Using human genomic DNA samples, we demonstrate successful targeting of causative loci for Huntingtons disease (HTT; CAG repeat), Fragile X syndrome (FMR1; CGG repeat), amyotrophic lateral sclerosis (ALS) and frontotemporal dementia (C9orf72; GGGGCC repeat), and spinocerebellar ataxia type 10 (SCA10) (ATXN10; variable ATTCT repeat). The method, amenable to multiplexing across multiple genomic loci, uses an amplification-free approach that facilitates the isolation of hundreds of individual on-target molecules in a single SMRT Cell and accurate sequencing through long repeat stretches, regardless of extreme GC percent or sequence complexity content. Our novel targeted sequencing method opens new doors to genomic analyses independent of PCR amplification that will facilitate the study of repeat expansion disorders.

genomics

beachmat: a Bioconductor C++ API for accessing single-cell genomics data from a variety of R matrix types

Recent advances in single-cell RNA sequencing have dramatically increased the number of cells that can be profiled in a single experiment. This provides unparalleled resolution to study cellular heterogeneity within biological processes such as differentiation. However, the explosion of data that are generated from such experiments poses a challenge to the existing computational infrastructure for statistical data analysis. In particular, large matrices holding expression values for each gene in each cell require sparse or file-backed representations for manipulation with the popular R programming language. Here, we describe a C++ interface named beachmat, which enables agnostic data access from various matrix representations. This allows package developers to write efficient C++ code that is interoperable with simple, sparse and HDF5-backed matrices, amongst others. We perform simulations to examine the performance of beachmat on each matrix representation, and we demonstrate how beachmat can be incorporated into the code of other packages to drive analyses of a very large single-cell data set.

bioinformatics

Telomerecat: A Ploidy-Agnostic Method For Estimating Telomere Length From Whole Genome Sequencing Data

Telomere length is a risk factor in disease and the dynamics of telomere length are crucial to our understanding of cell replication and vitality. The proliferation of whole genome sequencing represents an unprecedented opportunity to glean new insights into telomere biology on a previously unimaginable scale. To this end, a number of approaches for estimating telomere length from whole-genome sequencing data have been proposed. Here we present Telomerecat, a novel approach to the estimation of telomere length. Previous methods have been dependent on the number of telomeres present in a cell being known, which may be problematic when analysing aneuploid cancer data and non-human samples. Telomerecat is designed to be agnostic to the number of telomeres present, making it suited for the purpose of estimating telomere length in cancer studies. Telomerecat also accounts for interstitial telomeric reads and presents a novel approach to dealing with sequencing errors. We show that Telomerecat performs well at telomere length estimation when compared to leading experimental and computational methods. Furthermore, we show that it detects expected patterns in longitudinal data, technical replicates, and cross-species comparisons. We also apply the method to a cancer cell data, uncovering an interesting relationship with the underlying telomerase genotype.

bioinformatics