EconPapers    
Economics at your fingertips  
 

Sensitivity of machine learning regression models to data structure and quality in crop yield prediction

Mahoukpégo Luc Zinzinhédo, Mèvognon Firmin Mitchozounnou, Kolawolé Valère Salako and Romain Glèlè Kakaï

PLOS ONE, 2026, vol. 21, issue 9, 1-25

Abstract: Accurate crop yield prediction is essential for agricultural planning, yet machine learning (ML) models remain highly sensitive to the quality and structure of input data. This study uses simulated datasets to systematically investigate how data structure (sample size and number of predictors), data imperfections (missing values and multicollinearity), and pre-processing methods (imputation techniques and principal component analysis) influence ML regression performance. The performance of ML-based models (RF, SVM, MLR, XGBoost, LightGBM, NNet, and kNN) on the pre-processed data was evaluated. In total, each algorithm was tested on 1,728 datasets and pre-processing scenarios, yielding 12,096 model-scenario evaluations across the seven ML algorithms. Results show that missing data decreases performance (R2 drops by up to 20.63%; MAE increases by 12.67%), while multicollinearity may inflate R2 values despite poorer MAE performance. Larger sample sizes consistently improve prediction accuracy (R2 = +18.87%; MAE = −8.16%), whereas more predictors generally reduce it (R2 = −8.47%; MAE = +44.74%). Regression-based imputation improved both R2 and MAE the most, while RF demonstrated greater robustness across varying conditions. These simulation findings were further validated on five real-world crop datasets (Maize, Yam, Cassava, Sorghum, and Peanuts), confirming that despite RF remains the safest and most robust default, no single pre-processing strategy is universally optimal and that the selection of imputation method and dimensionality reduction must be strictly contingent upon the dataset’s specific missingness rate and correlation structure. Overall, this study highlights the complex interplay between data characteristics and pre-processing, urging the development of clearer guidelines for applying ML to agricultural datasets.

Date: 2026
References: Add references at CitEc
Citations:

Downloads: (external link)
https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0353938 (text/html)
https://journals.plos.org/plosone/article/file?id= ... 53938&type=printable (application/pdf)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:plo:pone00:0353938

DOI: 10.1371/journal.pone.0353938

Access Statistics for this article

More articles in PLOS ONE from Public Library of Science
Bibliographic data for series maintained by plosone ().

 
Page updated 2026-09-06
Handle: RePEc:plo:pone00:0353938