Background and objective: Pediatric celiac disease (CD) is characterized by marked clinical and immunological heterogeneity that is not fully captured by the current symptom-based Oslo classification. Data-driven computational phenotyping may enable the identification of more homogeneous patient subgroups by integrating multimodal clinical information. This study proposes a topological data analysis (TDA)-based computational framework (pheTDA) to identify and characterize novel sub-phenotypes of pediatric CD from multicentric clinical data. Methods: A retrospective multicentric cohort of 2,922 pediatric CD patients described by 34 demographic, clinical, serological, genetic, and histological variables was analyzed. Mixed-type data were processed using Gower distance and mapped through a semi-supervised TDA Mapper pipeline, followed by Louvain community detection. The resulting communities were characterized through statistical testing and supervised machine learning. pheTDA was benchmarked against t-SNE followed by DBSCAN and agglomerative hierarchical clustering (AHC), and sensitivity analyses repeated pheTDA using clinical variables only and diagnostic-test variables only. Results: pheTDA identified seven relatively balanced sub-phenotypes with distinct combinations of symptoms, serological profiles, histological damage, growth delay, and autoimmune comorbidities. Compared with pheTDA, DBSCAN generated a markedly imbalanced nine-cluster partition, whereas the four-cluster AHC solution was dominated by sex. pheTDA showed a clinically broader effect-size profile, with its leading discriminators spanning clinical, serological, genetic, and histological domains. Clinical-only pheTDA communities were strongly aligned with the symptom-based Oslo categories, whereas diagnostic-test-only communities were more fragmented and reflected specific test combinations. Random Forest classifiers achieved a higher mean AUC for pheTDA communities than for Oslo classes (0.94 vs. 0.88). Conclusions: Multimodal pheTDA provides an exploratory and reproducible framework for identifying clinically interpretable pediatric CD sub-phenotypes. Benchmarking and feature-domain sensitivity analyses support the value of integrating clinical and diagnostic information while emphasizing that the resulting communities are data-dependent computational representations rather than definitive disease classes.
(2026). A new computational phenotyping framework for the clinical characterization of pediatric celiac disease [journal article - articolo]. In COMPUTER METHODS AND PROGRAMS IN BIOMEDICINE. Retrieved from https://hdl.handle.net/10446/335427
A new computational phenotyping framework for the clinical characterization of pediatric celiac disease
Brembilla, Valentina;Lenzi, Erika;Maffioletti, Simone;Medolago, Emanuele;Sirtoli, Chiara;Ferramosca, Antonio;Pala, Daniele
2026-09-17
Abstract
Background and objective: Pediatric celiac disease (CD) is characterized by marked clinical and immunological heterogeneity that is not fully captured by the current symptom-based Oslo classification. Data-driven computational phenotyping may enable the identification of more homogeneous patient subgroups by integrating multimodal clinical information. This study proposes a topological data analysis (TDA)-based computational framework (pheTDA) to identify and characterize novel sub-phenotypes of pediatric CD from multicentric clinical data. Methods: A retrospective multicentric cohort of 2,922 pediatric CD patients described by 34 demographic, clinical, serological, genetic, and histological variables was analyzed. Mixed-type data were processed using Gower distance and mapped through a semi-supervised TDA Mapper pipeline, followed by Louvain community detection. The resulting communities were characterized through statistical testing and supervised machine learning. pheTDA was benchmarked against t-SNE followed by DBSCAN and agglomerative hierarchical clustering (AHC), and sensitivity analyses repeated pheTDA using clinical variables only and diagnostic-test variables only. Results: pheTDA identified seven relatively balanced sub-phenotypes with distinct combinations of symptoms, serological profiles, histological damage, growth delay, and autoimmune comorbidities. Compared with pheTDA, DBSCAN generated a markedly imbalanced nine-cluster partition, whereas the four-cluster AHC solution was dominated by sex. pheTDA showed a clinically broader effect-size profile, with its leading discriminators spanning clinical, serological, genetic, and histological domains. Clinical-only pheTDA communities were strongly aligned with the symptom-based Oslo categories, whereas diagnostic-test-only communities were more fragmented and reflected specific test combinations. Random Forest classifiers achieved a higher mean AUC for pheTDA communities than for Oslo classes (0.94 vs. 0.88). Conclusions: Multimodal pheTDA provides an exploratory and reproducible framework for identifying clinically interpretable pediatric CD sub-phenotypes. Benchmarking and feature-domain sensitivity analyses support the value of integrating clinical and diagnostic information while emphasizing that the resulting communities are data-dependent computational representations rather than definitive disease classes.Pubblicazioni consigliate
Aisberg ©2008 Servizi bibliotecari, Università degli studi di Bergamo | Terms of use/Condizioni di utilizzo

