Faculty & Staff Scholarship

Effects of Training Set Size on Supervised Machine-Learning Land-Cover Classification of Large-Area High-Resolution Remotely Sensed Data

Author ORCID Identifier

https://orcid.org/0000-0001-9580-9213

https://orcid.org/0000-0002-0414-9748

https://orcid.org/0000-0002-4412-5599

https://orcid.org/0000-0002-0619-3347

Document Type

Article

Publication Date

2021

College/Unit

Chambers College of Business and Economics

Department/Program/Center

Management Information Systems

Abstract

The size of the training data set is a major determinant of classification accuracy. Neverthe- less, the collection of a large training data set for supervised classifiers can be a challenge, especially for studies covering a large area, which may be typical of many real-world applied projects. This work investigates how variations in training set size, ranging from a large sample size (n = 10,000) to a very small sample size (n = 40), affect the performance of six supervised machine-learning algo- rithms applied to classify large-area high-spatial-resolution (HR) (1–5 m) remotely sensed data within the context of a geographic object-based image analysis (GEOBIA) approach. GEOBIA, in which adjacent similar pixels are grouped into image-objects that form the unit of the classification, offers the potential benefit of allowing multiple additional variables, such as measures of object geometry and texture, thus increasing the dimensionality of the classification input data. The six supervised machine-learning algorithms are support vector machines (SVM), random forests (RF), k-nearest neighbors (k-NN), single-layer perceptron neural networks (NEU), learning vector quantization (LVQ), and gradient-boosted trees (GBM). RF, the algorithm with the highest overall accuracy, was notable for its negligible decrease in overall accuracy, 1.0%, when training sample size decreased from 10,000 to 315 samples. GBM provided similar overall accuracy to RF; however, the algorithm was very expensive in terms of training time and computational resources, especially with large training sets. In contrast to RF and GBM, NEU, and SVM were particularly sensitive to decreasing sample size, with NEU classifications generally producing overall accuracies that were on average slightly higher than SVM classifications for larger sample sizes, but lower than SVM for the smallest sample sizes. NEU however required a longer processing time. The k-NN classifier saw less of a drop in overall accuracy than NEU and SVM as training set size decreased; however, the overall accuracies of k-NN were typically less than RF, NEU, and SVM classifiers. LVQ generally had the lowest overall accuracy of all six methods, but was relatively insensitive to sample size, down to the smallest sample sizes. Overall, due to its relatively high accuracy with small training sample sets, and minimal variations in overall accuracy between very large and small sample sets, as well as relatively short processing time, RF was a good classifier for large-area land-cover classifications of HR remotely sensed data, especially when training data are scarce. However, as performance of different supervised classifiers varies in response to training set size, investigating multiple classification algorithms is recommended to achieve optimal accuracy for a project.

Digital Commons Citation

Ramezan, Christopher A.; Warner, Timothy A.; Maxwell, Aaron E.; and Price, Bradley S., "Effects of Training Set Size on Supervised Machine-Learning Land-Cover Classification of Large-Area High-Resolution Remotely Sensed Data" (2021). Faculty & Staff Scholarship. 2976.
https://researchrepository.wvu.edu/faculty_publications/2976

Source Citation

Ramezan,C.A.;Warner, T.A.; Maxwell, A.E.; Price, B.S. Effects of Training Set Size on Supervised Machine-Learning Land-Cover Classification of Large-Area High-Resolution Remotely Sensed Data. Remote Sens. 2021, 13, 368. https://doi.org/10.3390/rs13030368

Comments

Copyright: © 2021 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/).

This article received support from the WVU Libraries' Open Access Author Fund.

Download

Included in

Management Information Systems Commons, Remote Sensing Commons

COinS

Faculty & Staff Scholarship

Effects of Training Set Size on Supervised Machine-Learning Land-Cover Classification of Large-Area High-Resolution Remotely Sensed Data

Author ORCID Identifier

Document Type

Publication Date

College/Unit

Department/Program/Center

Abstract

Digital Commons Citation

Source Citation

Comments

Included in

Browse

Resources

Search

Author Corner

Faculty & Staff Scholarship

Effects of Training Set Size on Supervised Machine-Learning Land-Cover Classification of Large-Area High-Resolution Remotely Sensed Data

Authors

Author ORCID Identifier

Document Type

Publication Date

College/Unit

Department/Program/Center

Abstract

Digital Commons Citation

Source Citation

Comments

Included in

Share

Browse

Resources

Search

Author Corner