S-SIRUS: an explainability algorithm for spatial regression Random Forest

Random Forest (RF) is a widely used machine learning algorithm known for its flexibility, user-friendliness, and high predictive performance across various domains. However, it is non-interpretable. This can limit its usefulness in applied sciences, where understanding the relationships between predictors and response variable is crucial from a decision-making perspective. In the literature, several methods have been proposed to explain RF, but none of them addresses the challenge of explaining RF in the context of spatially dependent data. Therefore, this work aims to explain regression RF in the case of spatially dependent data by extracting a compact and simple list of rules from an RF that explicitly takes into account the spatial correlation, i.e. RF-GLS. In this respect, we propose S-SIRUS, a spatial extension of SIRUS, the latter being a well-established regression rule algorithm able to extract a stable and short list of rules from the classical regression RF algorithm. To our knowledge, S-SIRUS is the only explainability tool proposed to open an RF-GLS, which, in turn, is the only random forest algorithm in the literature that accounts for spatial correlation internally in the algorithm. A simulation study was conducted to evaluate the explainability capability of the proposed S-SIRUS, by considering different levels of spatial dependence among the data. The results suggest that S-SIRUS exhibits a higher test predictive accuracy than SIRUS when spatial correlation is present. We encourage the use of SIRUS in the absence of spatial correlation and recommend adopting S-SIRUS when such correlation is present.

(2025). S-SIRUS: an explainability algorithm for spatial regression Random Forest [journal article - articolo]. In STATISTICS AND COMPUTING. Retrieved from https://hdl.handle.net/10446/306767

S-SIRUS: an explainability algorithm for spatial regression Random Forest

Patelli, Luca;Golini, Natalia;Ignaccolo, Rosaria;Cameletti, Michela

2025-07-04

Abstract

Random Forest (RF) is a widely used machine learning algorithm known for its flexibility, user-friendliness, and high predictive performance across various domains. However, it is non-interpretable. This can limit its usefulness in applied sciences, where understanding the relationships between predictors and response variable is crucial from a decision-making perspective. In the literature, several methods have been proposed to explain RF, but none of them addresses the challenge of explaining RF in the context of spatially dependent data. Therefore, this work aims to explain regression RF in the case of spatially dependent data by extracting a compact and simple list of rules from an RF that explicitly takes into account the spatial correlation, i.e. RF-GLS. In this respect, we propose S-SIRUS, a spatial extension of SIRUS, the latter being a well-established regression rule algorithm able to extract a stable and short list of rules from the classical regression RF algorithm. To our knowledge, S-SIRUS is the only explainability tool proposed to open an RF-GLS, which, in turn, is the only random forest algorithm in the literature that accounts for spatial correlation internally in the algorithm. A simulation study was conducted to evaluate the explainability capability of the proposed S-SIRUS, by considering different levels of spatial dependence among the data. The results suggest that S-SIRUS exhibits a higher test predictive accuracy than SIRUS when spatial correlation is present. We encourage the use of SIRUS in the absence of spatial correlation and recommend adopting S-SIRUS when such correlation is present.

Scheda breve

Scheda completa

Scheda completa (DC)

	Tipo di articolo
	
				articolo
			
	Data di pubblicazione
	
				4-lug-2025
			
	Rivista in ANCE
	
				STATISTICS AND COMPUTING
			
	Tutti gli autori
	
						Patelli, Luca; Golini, Natalia; Ignaccolo, Rosaria; Cameletti, Michela
					
	Citazione
	
				(2025). S-SIRUS: an explainability algorithm for spatial regression Random Forest  [journal article - articolo]. In STATISTICS AND COMPUTING. Retrieved from https://hdl.handle.net/10446/306767
			
	Nelle collezioni:
	
				1.1.01 Articoli/Saggi in rivista - Journal Articles/Essays

File allegato/i alla scheda:

File	Dimensione del file	Formato
2025 Patelli et al - S-SIRUS an explainability algorithm for spatial regression Random Forest - Stat Comput.pdf Solo gestori di archivio Versione: publisher's version - versione editoriale Licenza: Licenza default Aisberg Dimensione del file 2.42 MB Formato Adobe PDF Visualizza/Apri	2.42 MB	Adobe PDF	Visualizza/Apri

Pubblicazioni consigliate

Aisberg ©2008 Servizi bibliotecari, Università degli studi di Bergamo | Terms of use/Condizioni di utilizzo

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10446/306767

Citazioni

0

0

social impact