Search bioRxiv⌕ Search

Biology subjects

Fowler, E. E. E.

Publications and source records attributed to Fowler, E. E. E..

2 recordsLinked to original sources

Breast Density Analysis of Digital Breast Tomosynthesis

We evaluated an automated percentage of breast density (BD) technique (PDa) with digital breast tomosynthesis (DBT) data. The approach is based on the wavelet expansion followed by analyzing signal dependent noise. Several measures were investigated as risk factors: normalized volumetric; total dense volume; average of the DBT slices (slice-mean); a two-dimensional (2D) metric applied to the synthetic images; and the mean and standard deviations of the pixel values. Volumetric measures were derived theoretically, and PDa was modeled as a function of compressed breast thickness. An alternative method for constructing synthetic 2D mammograms was investigated using the volume results. A matched case-control study (n = 426 pairs) was analyzed. Conditional logistic regression modeling, controlling body mass index and ethnicity, was used to estimate odds ratios (ORs) for each measure with 95% confidence intervals provided parenthetically. There were several significant findings: volumetric measure [OR = 1.43 (1.18, 1.72)], which produced an identical OR as the slice-mean measure as predicted; [OR =1.44 (1.18, 1.75)] when applied to the synthetic images; and mean of the pixel values (volume or 2D synthetic) [ORs [~] 1.31 (1.09, 1.57)]. PDa was modeled as 2nd degree polynomial (concave-down): its maximum value occurred at 0.41x(compressed breast thickness), which was similar across case-control groups, and was significant from this position [OR = 1.47 (1.21, 1.78)]. A standardized 2D synthetic image was produced, where each pixel value represents the percentage of BD above its location. The significant findings indicate the validity of the technique. Derivations supported by empirical analyses produced a new synthetic 2D standardized image technique. Ancillary to the objectives, the results provide evidence for understanding the percentage of BD measure applied to 2D mammograms. Notwithstanding the findings, the study design provides a template for investigating other measures such as texture.

biophysics↗

Techniques to Produce and Evaluate Realistic Multivariate Synthetic Data

BackgroundData modeling in biomedical-healthcare research requires a sufficient sample size for exploration and reproducibility purposes. A small sample size can inhibit model performance evaluations (i.e., the small sample problem). ObjectiveA synthetic data generation technique addressing the small sample size problem is evaluated. We show: (1) from the space of arbitrarily distributed samples, a subgroup (class) has a latent multivariate normal characteristic; (2) synthetic populations (SPs) of unlimited size can be generated from this class with univariate kernel density estimation (uKDE) followed by standard normal random variable generation techniques; and (3) samples drawn from these SPs are statistically like their respective samples. MethodsThree samples (n = 667), selected pseudo-randomly, were investigated each with 10 input variables (i.e., X). uKDE (optimized with differential evolution) was used to augment the sample size in X (i.e., the input variables). The enhanced sample size was used to construct maps that produced univariate normally distributed variables in Y (mapped input variables). Principal component analysis in Y produced uncorrelated variables in T, where the univariate probability density functions (pdfs) were approximated as normal with specific variances; a given SP in T was generated with normally distributed independent random variables with these specified variances. Reversing each step produced the respective SPs in Y and X. Synthetic samples of the same size were drawn from these SPs for comparisons with their respective samples. Multiple tests were deployed: to assess univariate and multivariate normality; to compare univariate and multivariate pdfs; and to compare covariance matrices. ResultsOne sample was approximately multivariate normal in X and all samples were approximately multivariate normal in Y, permitting the generation of unlimited sized SPs. Uni/multivariate pdf and covariance comparisons (in X, Y and T) showed similarity between samples and synthetic samples. ConclusionsThe work shows that a class of multivariate samples has a latent normal characteristic; for such samples, our technique is a simplifying mechanism that offers an approximate solution to the small sample problem by generating similar synthetic data. Further studies are required to understand this latent normal class, as two samples exhibited this characteristic in the study.

bioinformatics↗