Search bioRxiv⌕ Search

Biology subjects

Bohlman, S. A.

Publications and source records attributed to Bohlman, S. A..

3 recordsLinked to original sources

Continental-scale Hyperspectral tree species classification in the National Ecological Observatory Network

Advances in remote sensing imagery and machine learning applications unlock the potential for developing algorithms for species classification at the level of individual tree crowns at unprecedented scales. However, most approaches to date focus on site-specific applications and a small number of taxonomic groups. Little is known about how well these approaches generalize across broader geographic areas and ecosystems. Leveraging field surveys and hyperspectral remote sensing data from the National Ecological Observatory Network (NEON), we developed a continental-extent model for tree species classification that can be applied to the network, including a wide range of US terrestrial ecosystems. We compared the performance of a model trained with data from 27 NEON sites to models trained with data from each individual site, evaluating advantages and challenges posed by training species classifiers at the US scale. We evaluated the effect of geographic location, topography, and ecological conditions on the accuracy and precision of species predictions (72 out of 77 species). On average, the general model resulted in good overall classification accuracy (micro-F1 score), with better accuracy than site-specific classifiers (average individual tree level accuracy of 0.77 for the general model and 0.70 for site-specific models). Aggregating species to the genus-level increased accuracy to 0.83. Regions with more species exhibited lower classification accuracy. Predicted species were more likely to be confused with congeneric and co-occurring species and confusion was highest for trees with structural damage and in complex closed-canopy forests. The model produced accurate estimates of uncertainty, correctly identifying trees where confusion was likely. Using only data from NEON, this single integrated classifier can make predictions for 20% of all tree species found in forest ecosystems across the entire US, which make up to roughly 90% of the upper canopy of the studied ecosystems. This suggests the potential for integrating information from multiple datasets and locations to develop broad scale general models for species classification from hyperspectral imaging.

ecology↗

Data science competition for cross-site delineation and classification of individual trees from airborne remote sensing data

Delineating and classifying individual trees in remote sensing data is challenging. Many tree crown delineation methods have difficulty in closed-canopy forests and do not leverage multiple datasets. Methods to classify individual species are often accurate for common species, but perform poorly for less common species and when applied to new sites. We ran a data science competition to help identify effective methods for delineation of individual crowns and classification to determine species identity. This competition included data from multiple sites to assess the methods ability to generalize learning across multiple sites simultaneously, and transfer learning to novel sites where the methods were not trained. Six teams, representing 4 countries and 9 individual participants, submitted predictions. Methods from a previous competition were also applied and used as the baseline to understand whether the methods are changing and improving over time. The best delineation method was based on an instance segmentation pipeline, closely followed by a Faster R-CNN pipeline, both of which outperformed the baseline method. However, the baseline (based on a growing region algorithm) still performed well as did the Faster R-CNN. All delineation methods generalized well and transferred to novel forests effectively. The best species classification method was based on a two-stage fully connected neural network, which significantly outperformed the baseline (a random forest and Gradient boosting ensemble). The classification methods generalized well, with all teams training their models using multiple sites simultaneously, but the predictions from these trained models generally failed to transfer effectively to a novel site. Classification performance was strongly influenced by the number of field-based species IDs available for training the models, with most methods predicting common species well at the training sites. Classification errors (i.e., species misidentification) were most common between similar species in the same genus and different species that occur in the same habitat. The best methods handled class imbalance well and learned unique spectral features even with limited data. Most methods performed better than baseline in detecting new (untrained) species, especially in the site with no training data. Our experience further shows that data science competitions are useful for comparing different methods through the use of a standardized dataset and set of evaluation criteria, which highlights promising approaches and common challenges, and therefore advances the ecological and remote sensing field as a whole.

ecology↗

Disentangling the roles of inter and intraspecific variation on leaf trait distributions across the eastern United States

Functional traits are influenced by phylogenetic constraints and environmental conditions, but previous large-scale studies modeled traits either as species weighted averages or directly from the environment, precluding analyses of the relative contributions of inter- and intraspecific variation across regions. We developed a joint model integrating phylogenetic and environmental information to understand and predict the distribution of eight leaf traits across the eastern USA. This model explained 68% of trait variation, outperforming both species-only and environment-only models, with variance attributable to species alone (23%), the environment alone (13%), and their combined effects (25%). The importance of the two drivers varied by trait. Predictions for the eastern USA produced accurate estimates of intraspecific variation and deviated from both species-only and environment-only models. Predictions revealed that intraspecific variation holds information across scales, affects relationships in the leaf economic spectrum and is key for interpreting trait distributions and ecosystem processes within and across ecoregions.

ecology↗