نوع مقاله : مقاله پژوهشی
عنوان مقاله English
نویسندگان English
Introduction
Accurate prediction of maize (Zea mays L.) grain yield is critical for efficient resource management and enhancing productivity in sustainable agriculture, particularly in low-input systems. The integration of biofertilizers, such as plant growth-promoting rhizobacteria (PGPR) and arbuscular mycorrhizal fungi (AMF), offers a promising avenue to improve crop performance while reducing environmental impacts. However, the complexity of ecophysiological data and their interactions necessitates advanced modeling techniques for precise yield prediction. Machine learning (ML) algorithms, combined with optimization methods, provide robust tools to address this challenge (Ingole et al., 2025). This study proposes a novel hybrid approach combining the Non-dominated Sorting Genetic Algorithm II (NSGA-II) for feature selection and the eXtreme Gradient Boosting (XGBoost) algorithm for predictive modeling. By leveraging ecophysiological data, this approach aims to optimize maize yield predictions under biofertilizer applications, supporting precision agriculture and climate adaptation strategies. The objectives were to identify key predictive features, enhance model accuracy, and interpret the contributions of selected variables to yield outcomes.
Materials and Methods
The study utilized data from 96 experimental plots at the Ferdowsi University of Mashhad's research fields, collected over two years. A comprehensive dataset comprising 75 features (32 primary and 43 interaction terms) was compiled, including ecophysiological variables such as chlorophyll content index (SPAD), canopy temperature, maximum photosynthesis rate, specific root length, and soil phosphorus percentage. The methodology consisted of two main phases: feature selection and predictive modeling. First, NSGA-II, a multi-objective optimization algorithm, was employed to select an optimal subset of features by simultaneously maximizing prediction accuracy (coefficient of determination, R²) and minimizing the number of features. NSGA-II iteratively evaluated feature combinations to identify a Pareto-optimal set, balancing model simplicity and performance. Subsequently, the selected features were used to train an XGBoost model, a gradient-boosting framework known for its robustness in handling complex datasets. The model was validated using 5-fold cross-validation to ensure generalizability. Model interpretability was enhanced through SHapley Additive exPlanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME) methods. SHAP analysis provided global and local feature importance, while LIME elucidated individual prediction contributions. Counterfactual analysis was also conducted to assess the impact of removing key features on model predictions.
Results and Discussion
NSGA-II successfully identified nine key features critical for yield prediction, including canopy temperature at the dough stage (Canopy Temp_4), chlorophyll content index at the milk stage (SPAD_3), soil respiration rate, and maximum photosynthesis rate. These features were selected for their ability to maximize R² while maintaining model parsimony. The trained XGBoost model achieved an R² of 0.63 and a root mean square error (RMSE) of 2.171 tons per hectare, indicating moderate predictive accuracy suitable for agricultural applications. SHAP analysis revealed that SPAD_3 (mean SHAP value: 0.2603 t/ha) and Canopy Temp_4 (0.2512 t/ha) were the most influential predictors, positively contributing to yield predictions. Conversely, complex interactions, such as the interaction between leaf area index at the silking stage and the squared mean canopy temperature, had a slight negative effect (-0.0674 t/ha). LIME analysis further highlighted the negative influence of ear length (weight: -0.9252), suggesting its role in reducing predicted yields in specific cases. The Pareto front generated by NSGA-II corroborated the trade-off between feature count and prediction accuracy, while feature importance plots aligned with SHAP and LIME findings. Counterfactual analysis demonstrated that excluding key features significantly altered predictions, underscoring their importance. The integration of NSGA-II and XGBoost proved highly effective in optimizing maize yield predictions under biofertilizer influence, offering a robust framework for low-input agriculture. The selection of nine key features by NSGA-II reduced model complexity while maintaining predictive power, aligning with the principles of precision agriculture. The high influence of SPAD_3 and Canopy Temp_4 underscores the importance of physiological and environmental factors in yield determination, particularly in biofertilizer-enhanced systems. The negative contribution of certain interaction terms and ear length suggests the need for careful consideration of variable interactions in model development. The use of SHAP and LIME enhanced model transparency, providing actionable insights for farmers and researchers. Compared to traditional models, this hybrid approach outperformed baseline methods, achieving a higher R² and lower RMSE. The findings support the adoption of advanced ML techniques in sustainable agriculture, enabling precise resource allocation and adaptation to climatic variability. Future research could explore additional biofertilizer types and larger datasets to further refine predictive accuracy and generalizability.
Conclusion
In conclusion, this study successfully developed and validated a hybrid modeling framework integrating NSGA-II for multi-objective feature selection with XGBoost for predictive modeling, enabling accurate maize grain yield prediction under biofertilizer applications in low-input systems. By identifying a parsimonious set of nine key ecophysiological features—most notably chlorophyll content index at the milk stage (SPAD_3) and canopy temperature at the dough stage—the approach achieved a respectable R² of 0.63 and RMSE of 2.171 t/ha, demonstrating moderate yet practical predictive power suitable for precision agriculture in variable environments. Interpretability analyses via SHAP and LIME underscored the dominant positive contributions of physiological indicators while highlighting nuanced negative effects from certain interactions and traits like ear length, providing valuable insights into biofertilizer-mediated yield dynamics. This methodology not only outperformed traditional baselines by balancing model complexity and accuracy but also supports sustainable resource management and climate-resilient strategies in arid and semi-arid regions. Future efforts should expand datasets, incorporate diverse biofertilizer regimes, and test generalizability across broader agroecological zones to further enhance predictive robustness and real-world applicability.
کلیدواژهها English
Authors retain the copyright. This is an open access article distributed under Creative Commons Attribution 4.0 International License (CC BY 4.0)