بوم شناسی کشاورزی

بوم شناسی کشاورزی

واسنجی، بهینه‌سازی و ارزیابی الگوریتم تلفیقی NSGA-II و XGBOOST برای پیش‌بینی عملکرد دانه ذرت (Zea mays L.) تحت تأثیر کودهای زیستی

نوع مقاله : مقاله پژوهشی

نویسندگان
گروه اگروتکنولوژی، دانشکده کشاورزی، دانشگاه فردوسی مشهد، مشهد، ایران
چکیده
پیش‌بینی دقیق عملکرد دانه ذرت (Zea mays L.) برای مدیریت منابع و افزایش بهره‌وری در کشاورزی پایدار ضروری است. در این مطالعه ابتدا، الگوریتمNSGA-II  برای انتخاب بهینه ویژگی‌‌ها به کار گرفته شد و سپس الگوریتم  XGBoostبا ویژگی‌‌های انتخاب‌شده برای پیش‌بینی آموزش داده شد. داده‌ها از ۹۶ نمونه از مزرعه‌ تحقیقاتی دانشگاه فردوسی مشهد شامل ۷۵ ویژگی (۳۲ اصلی و ۴۳ تعاملی) طی دو سال آزمایش جمع‌آوری شدند. داده‌های اکوفیزیولوژیکی شامل متغیرهایی نظیر شاخص کلروفیل (SPAD)، دمای تاج‌پوشش، سرعت فتوسنتز بیشینه، طول ویژه ریشه، درصد فسفر خاک و غیره بودند. ابتدا، NSGA-II برای شناسایی مجموعه بهینه ویژگی‌ها به‌منظور حداکثر کردن دقت پیش‌بینی (R²) و کمینه کردن تعداد ویژگی‌ها به‌کار گرفته شد. NSGA-II با بهینه‌سازی همزمان تعداد ویژگی‌ها و دقت، نُه ویژگی کلیدی (مانند دمای تاج‌پوشش در مرحله خمیری دانه (Canopy Temp_4)، شاخص کلروفیل برگ در مرحله شیری دانه (SPAD_3)، سرعت تنفس خاک و سرعت فتوسنتز بیشینه را انتخاب کرد. مدل XGBoost با اعتبارسنجی (پنج‌تایی) آموزش داده شد و ضریب تبیین (R²) 63/0 و جذر میانگین مربعات خطا (RMSE) 171/2 تن در هکتار به دست آمد. برای تفسیر مدل، از روش‌های شیپ (SHAP) و لایم (LIME) استفاده شد. نمودارهای شیپ (نمودارهای آبشاری و فورس پلات) نشان دادند که SPAD_3 با میانگین شیپ 2603/0 تن در هکتار و Canopy Temp_4 با 2512/0 تن در هکتار، نقش کلیدی در افزایش پیش‌بینی داشتند، درحالی‌که تعاملات پیچیده مثل اثر متقابل شاخص سطح برگ در مرحله کاکل‌دهی و مجذور میانگین دمای تاج‌پوشش اثر منفی اندکی (0674/0-) ایجاد کردند. تحلیل لایم نیز تأیید کرد که طول بلال با وزن 9252/0- تأثیر منفی قابل ‌توجهی دارد. این یافته‌ها با جبهه پاراتو و نمودار اهمیت ویژگی‌ها هم‌راستا بودند. علاوه‌بر این، تحلیل کانترفکچوآل نشان داد که حذف ویژگی‌های کلیدی، تغییرات قابل ‌توجهی در پیش‌بینی ایجاد می‌کند. نتایج این پژوهش نشان‌دهنده کارآیی بالای تلفیق NSGA-II و XGBoost در بهینه‌سازی مدل‌های پیش‌بینی کشاورزی است که به‌نوبه خود می‌تواند به مدیریت دقیق زراعی، انتخاب نهاده‌‌های طبیعی و سازگاری با تغییرات اقلیمی کمک کند.
کلیدواژه‌ها
موضوعات

عنوان مقاله English

Calibration, Optimization, and Evaluation of the Integrated NSGA-II and XGBoost Algorithm for Predicting Corn (Zea mays L.) Yield Performance under the Influence of Biofertilizers: A Novel Approach in Low-Input Agriculture

نویسندگان English

Mohsen Jahan
Mehdi Nassiri Mahallati
Department of Agrotechnology, Faculty of Agriculture, Ferdowsi University of Mashhad, Mashhad, Iran
چکیده English

Introduction
Accurate prediction of maize (Zea mays L.) grain yield is critical for efficient resource management and enhancing productivity in sustainable agriculture, particularly in low-input systems. The integration of biofertilizers, such as plant growth-promoting rhizobacteria (PGPR) and arbuscular mycorrhizal fungi (AMF), offers a promising avenue to improve crop performance while reducing environmental impacts. However, the complexity of ecophysiological data and their interactions necessitates advanced modeling techniques for precise yield prediction. Machine learning (ML) algorithms, combined with optimization methods, provide robust tools to address this challenge (Ingole et al., 2025). This study proposes a novel hybrid approach combining the Non-dominated Sorting Genetic Algorithm II (NSGA-II) for feature selection and the eXtreme Gradient Boosting (XGBoost) algorithm for predictive modeling. By leveraging ecophysiological data, this approach aims to optimize maize yield predictions under biofertilizer applications, supporting precision agriculture and climate adaptation strategies. The objectives were to identify key predictive features, enhance model accuracy, and interpret the contributions of selected variables to yield outcomes.
 
Materials and Methods
The study utilized data from 96 experimental plots at the Ferdowsi University of Mashhad's research fields, collected over two years. A comprehensive dataset comprising 75 features (32 primary and 43 interaction terms) was compiled, including ecophysiological variables such as chlorophyll content index (SPAD), canopy temperature, maximum photosynthesis rate, specific root length, and soil phosphorus percentage. The methodology consisted of two main phases: feature selection and predictive modeling. First, NSGA-II, a multi-objective optimization algorithm, was employed to select an optimal subset of features by simultaneously maximizing prediction accuracy (coefficient of determination, R²) and minimizing the number of features. NSGA-II iteratively evaluated feature combinations to identify a Pareto-optimal set, balancing model simplicity and performance. Subsequently, the selected features were used to train an XGBoost model, a gradient-boosting framework known for its robustness in handling complex datasets. The model was validated using 5-fold cross-validation to ensure generalizability. Model interpretability was enhanced through SHapley Additive exPlanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME) methods. SHAP analysis provided global and local feature importance, while LIME elucidated individual prediction contributions. Counterfactual analysis was also conducted to assess the impact of removing key features on model predictions.
 
Results and Discussion
NSGA-II successfully identified nine key features critical for yield prediction, including canopy temperature at the dough stage (Canopy Temp_4), chlorophyll content index at the milk stage (SPAD_3), soil respiration rate, and maximum photosynthesis rate. These features were selected for their ability to maximize R² while maintaining model parsimony. The trained XGBoost model achieved an R² of 0.63 and a root mean square error (RMSE) of 2.171 tons per hectare, indicating moderate predictive accuracy suitable for agricultural applications. SHAP analysis revealed that SPAD_3 (mean SHAP value: 0.2603 t/ha) and Canopy Temp_4 (0.2512 t/ha) were the most influential predictors, positively contributing to yield predictions. Conversely, complex interactions, such as the interaction between leaf area index at the silking stage and the squared mean canopy temperature, had a slight negative effect (-0.0674 t/ha). LIME analysis further highlighted the negative influence of ear length (weight: -0.9252), suggesting its role in reducing predicted yields in specific cases. The Pareto front generated by NSGA-II corroborated the trade-off between feature count and prediction accuracy, while feature importance plots aligned with SHAP and LIME findings. Counterfactual analysis demonstrated that excluding key features significantly altered predictions, underscoring their importance. The integration of NSGA-II and XGBoost proved highly effective in optimizing maize yield predictions under biofertilizer influence, offering a robust framework for low-input agriculture. The selection of nine key features by NSGA-II reduced model complexity while maintaining predictive power, aligning with the principles of precision agriculture. The high influence of SPAD_3 and Canopy Temp_4 underscores the importance of physiological and environmental factors in yield determination, particularly in biofertilizer-enhanced systems. The negative contribution of certain interaction terms and ear length suggests the need for careful consideration of variable interactions in model development. The use of SHAP and LIME enhanced model transparency, providing actionable insights for farmers and researchers. Compared to traditional models, this hybrid approach outperformed baseline methods, achieving a higher R² and lower RMSE. The findings support the adoption of advanced ML techniques in sustainable agriculture, enabling precise resource allocation and adaptation to climatic variability. Future research could explore additional biofertilizer types and larger datasets to further refine predictive accuracy and generalizability.
 
Conclusion
In conclusion, this study successfully developed and validated a hybrid modeling framework integrating NSGA-II for multi-objective feature selection with XGBoost for predictive modeling, enabling accurate maize grain yield prediction under biofertilizer applications in low-input systems. By identifying a parsimonious set of nine key ecophysiological features—most notably chlorophyll content index at the milk stage (SPAD_3) and canopy temperature at the dough stage—the approach achieved a respectable R² of 0.63 and RMSE of 2.171 t/ha, demonstrating moderate yet practical predictive power suitable for precision agriculture in variable environments. Interpretability analyses via SHAP and LIME underscored the dominant positive contributions of physiological indicators while highlighting nuanced negative effects from certain interactions and traits like ear length, providing valuable insights into biofertilizer-mediated yield dynamics. This methodology not only outperformed traditional baselines by balancing model complexity and accuracy but also supports sustainable resource management and climate-resilient strategies in arid and semi-arid regions. Future efforts should expand datasets, incorporate diverse biofertilizer regimes, and test generalizability across broader agroecological zones to further enhance predictive robustness and real-world applicability.



 

 




 
 

کلیدواژه‌ها English

Mycorrhizal fungi
Machine learning
Non-dominated sorting genetic algorithm
Rhizobacteria
SHAP analysis

Authors retain the copyright. This is an open access article distributed under Creative Commons Attribution 4.0 International License (CC BY 4.0)

  1. Amin, S., Hesam, M., Jabbari, A., & Abdolhosseini, M. (2023). Multi-objective optimization of cropping patterns with emphasis on economic benefits and ensuring food supply chain security (A case study: Gonbad-e Kavus - Golestan Dam). Journal of Water and Soil Conservation, 30(2), 119-139. https://doi.org/10.22069/jwsc.2023.21203.3637
  2. Baio, F.H.R., Santana, D.C., Teodoro, L.P.R., Oliveira, I.C.D., Gava, R., de Oliveira, J.L.G., Silva Junior, C.A.D., Teodoro, P.E., & Shiratsuchi, L.S. (2023). Maize yield prediction with machine learning, spectral variables and irrigation management. Remote Sensing, 15(1), 79. https://doi.org/10.3390/rs15010079
  3. Barikloo, A., alamdar, P., Moravej, K., & Servati, M. (2017). Prediction of irrigated wheat yield by using hybrid algorithm methods of artificial neural networks and genetic algorithm. Journal of Water and Soil, 31(3), 715-726. https://doi.org/10.22067/jsw.v31i3.56158
  4. Bergstra, J., & Bengio, Y. (2012). Random search for hyper-parameter optimization. Journal of Machine Learning Research, 13(10), 281-305.
  5. Bhupenchandra, I., Chongtham, S. K., Devi, A. G., Dutta, P., Sahoo, M. R., Mohanty, S., Kumar, S., Choudhary, A. K., Devi, E. L., Sinyorita, S., Devi, S. H., Mahanta, M., Kumari, A., Devi, H. L., Josmee, R. K., Pusparani, A., Pathaw, N., Gupta, S., Meena, M., ... Swapnil, P. (2024). Unlocking the potential of arbuscular mycorrhizal fungi: Exploring role in plant growth promotion, nutrient uptake mechanisms, biotic stress alleviation, and sustaining agricultural production systems. Journal of Plant Growth Regulators. 44, 6802–6840. https://doi.org/10.1007/s00344-024-11467-9
  6. Bracho-Mujica, G., Rötter, R. P., Haakana, M., Palosuo, T., Fronzek, S., Asseng, S., Yi, C., Ewert, F., Gaiser, T., Kassie, B., Paff, K., Rezaei, E. E., Rodríguez, A., Ruiz-Ramos, M., Srivastava, A. K., Stratonovitch, P., Tao, F., & Semenov, M. A. (2024). Effects of changes in climatic means, variability, and agro-technologies on future wheat and maize yields at 10 sites across the globe. Agricultural and Forest Meteorology, 346, Article 109887. https://doi.org/10.1016/j.agrformet.2024.109887
  7. Brar, B., Bala, K., Saharan, B. S., Sadh, P. K., & Duhan, J. S. (2024). Bio-boosting agriculture: Harnessing the potential of fungi-bacteria-plant synergies for crop improvement. Discover Plants, 1, 21 https://doi.org/10.1007/s44372-024-00023-0
  8. Cheng, E., Zhang, B., Peng, D., Zhong, L., Yu, L., Liu, Y., Xiao, C., Li, C., Li, X., Chen, Y., Ye, H., Wang, H., Yu, R., Hu, J., & Yang, S. (2022). Wheat yield estimation using remote sensing data based on machine learning approaches. Frontiers in Plant Science, 13, 2022. https://doi.org/10.3389/fpls.2022.1090970
  9. Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences, 2nd Ed. New York: Routledge.
  10. Corrales, C., Schoving, C., Raynal, H., et al. (2022). Rogate model based on feature selection techniques and regression learners to improve soybean yield prediction in southern France. Computers and Electronics in Agriculture, 192, 106578. https://doi.org/10.1016/j.compag.2021.106578
  11. Deb, K., Pratap, A., Agarwal, S., & Meyarivan, T. (2002). A Fast and elitist multiobjective genetic algorithm: NSGA-II. IEEE Transactions on Evolutionary Computation, 6(2), 182-197. https://doi.org/10.1109/4235.996017
  12. Du, J., Liu, R., Cheng, D., Wang, X., Zhang, T., & Yu, F. (2024). Enhancing NSGA-II algorithm through hybrid strategy for optimizing maize water and fertilizer irrigation simulation. Symmetry, 16(8), 1062. https://doi.org/10.3390/sym16081062
  13. Fashoto, S., Mbunge, E., Opeyemi, O.G., & Van Den Burg, J. (2021). Implementation of machine learning for predicting maize crop yields using multiple linear regression and backward elimination. Malaysian Journal of Computing, 6, 679-697.
  14. Hastie, T., Tibshirani, R., & Friedman, J. (2009). The Elements of Statistical Learning. Springer.
  15. Kakavand, S., Mazandarani Zadeh, H., & Ramezani Etedali, H. (2023). Optimal redistribution of water among agricultural sector operators using a fuzzy multi-objective optimization model. Irrigation Sciences and Engineering, 46(1), 77-93. https://doi.org/10.22055/jise.2021.37122.1966
  16. Khoso, M.A., Wagan, S., Alam, I., Hussain, A., Ali, Q., Saha, S., Poudel, T.R., Manghwar, H., & Liu, F. (2024). Impact of plant growth-promoting rhizobacteria (PGPR) on plant nutrition and root characteristics: Current perspective. Plant Stress, 11, https://doi.org/10.1016/j.stress.2023.100341.
  17. Kuhn, M., & Johnson, K. (2013). Applied Predictive Modeling. Springer Science & Business Media. 600 p. https://doi.org/10.1007/978-1-4614-6849-3
  18. Ingole, V. S., Kshirsagar, U. A., Singh, V., Yadav, M. V., Krishna, B., & Kumar, R. (2025). A hybrid model for soybean yield prediction integrating convolutional neural networks, recurrent neural networks, and graph convolutional networks. Computation, 13(1), 4. https://doi.org/10.3390/computation13010004
  19. Lyu, J., Jiang, Y., Xu, C., Liu, Y., Su, Z., Liu, J., & He, J. (2022). Multi-objective winter wheat irrigation strategies optimization based on coupling AquaCrop-OSPy and NSGA-III: A case study in Yangling, China. Science of The Total Environment, 843, 157104. https://doi.org/10.1016/j.scitotenv.2022.157104
  20. Mitra, D., Panneerselvam, P., Senapati, A., Chidambaranathan, P., Nayak, A.K., & Mohapatra, P.K.D. (2023). Arbuscular mycorrhizal fungi response on soil phosphorus utilization and enzymes activities in aerobic rice under phosphorus-deficient conditions. Life (Basel), 13(5), 1118. https://doi.org/10.3390/life13051118
  21. Moore, D.S., Notz, W.I., & Flinger, M.A. (2013). The Basic Practice of Statistics (6th Ed.). New York, NY: W. H. Freeman and Company. p. 138.
  22. Mukhopadhyay, A., Maulik, U., Bandyopadhyay, S., & Coello, C.A.C. (2014). A survey of multi-objective evolutionary algorithms for data mining. Applied Soft Computing, 23, 184-199. https://doi.org/10.1109/TEVC.2013.2290086
  23. Probst, P., Boulesteix, A.L., & Bischl, B. (2019). Tunability: Importance of hyperparameters of machine learning algorithms. Journal of Machine Learning Research, 20(53), 1-32.
  24. Sahabifard, F. Z., Shahnazari, A., & Sadeghi, S. (2024). Increasing the productivity of agricultural water under the optimization scenario of water resource allocation using the algorithm (NSGA-II). Journal of Water and Irrigation Management, 14(2), 405-419. DOI: https://doi.org/10.22059/jwim.2024.369097.1122
  25. Salvador, G.L.O., Araju, F.F., Pereira, A.P.A., Bonofacio, A., & Araujo, A.S.F. (2022). Rhizobacteria and arbuscular mycorrhizal fungus presented distinct and specific effects on soybean growth when inoculated with organic compost. Rhizosphere, 22, 100513. https://doi.org/10.1016/j.rhisph.2022.100513
  26. Shayegan, M., Alimohammadi, A., & Mansourian, A. (2011). Multi objective optimization of land use assignment using NSGA II. Remote Sensing and GIS in IRAN, 4(2), 1-18. Available Online at: https://scj.sbu.ac.ir/article_94927.html
  27. Slafer, G. A., Foulkes, M. J., Reynolds, M. P., Murchie, E. H., Carmo-Silva, E., Flavell, R., Gwyn, J., Sawkins, M., & Griffiths, S. (2023). A ‘wiring diagram’ for sink strength traits impacting wheat yield potential. Journal of Experimental Botany, 74(1), 40–71. https://doi.org/10.1093/jxb/erac410
  28. Tarbiat Modares University. (2016). Optimization of Water Resource Allocation for Agricultural Use Using Fuzzy Data Envelopment Analysis (FDEA) and NSGA-II Genetic Algorithm (Case Study: Varamin Plain). Available Online at: https://civilica.com/doc/1283657/
  29. Taiz, L., E. Zeiger, I. M. Moller, et al. (2018). Fundamentals of plant physiology. New York, USA: Oxford University. Press. ISBN 978160535790
  30. Xie, X., Wu, T., Zhu, M., Jiang, G., Xu, Y., Wang, X., & Pu, L. (2021). Comparison of random forest and multiple linear regression models for estimation of soil extracellular enzyme activities in agricultural reclaimed coastal saline land. Ecological Indicators, 120, https://doi.org/10.1016/j.ecolind.2020.106925
  31. Yang, C., Liu, C., & Ma, X (2025). Multi-objective irrigation strategies and production prediction for winter wheat in China in future dry years using CERES-wheat model and non-dominated sorting genetic algorithm II . Computers and Electronics in Agriculture, 230, 109888. https://doi.org/10.1016/j.compag.2024.109888
  32. Yang, C., Liu, C., Q.Z., Wang, S., Xing, X., & Ma, X. (2025). Multi-objective irrigation strategies and production prediction for winter wheat in China in future dry years using CERES-wheat model and non-dominated sorting genetic algorithm II. Computers and Electronics in Agriculture, 230, https://doi.org/10.1016/j.compag.2024.109888

 

ارسال نظر در مورد این مقاله
نام را وارد کنید.
نشانی پست الکترونیکی را به درستی وارد کنید.
وابستگی سازمانی را به درستی وارد کنید.
توضیحات را وارد کنید (حداقل 50 حرف)
CAPTCHA Image
شناسه امنیتی را به درستی وارد کنید.

  • تاریخ دریافت 22 تیر 1404
  • تاریخ بازنگری 12 دی 1404
  • تاریخ پذیرش 15 دی 1404
  • تاریخ اولین انتشار 01 بهمن 1404