Return
PEXP: A Scalable Parallel Tree-Based Framework for Interpreting Models on Big Data
DOI:10.1109/tbdata.2026.3668673.png)
Abstract
En 中文
The proliferation of Big Data has fueled the success of deep learning; however, its inherent opaque decision-making nature poses significant challenges for its adoption in safety-critical domains. Existing interpretable machine learning methods offer partial solutions but often struggle with model fidelity, inconsistent explanations, a lack of holistic model understanding, and critically, computational inefficiency, especially when applied to models trained on large-scale datasets. To overcome these hurdles, this paper introduces Parallel Explainer (PEXP), an innovative parallel tree-based interpretation framework designed for scalability and comprehensive understanding. PEXP initiates by generating a localized sample set around a target instance through data distribution-aware perturbations. It then computes similarity scores and employs a kernel function to weight these samples effectively. Leveraging concepts from Bagging and Boosting, PEXP efficiently constructs Parallel Ensemble Trees as its core interpretable model. This model provides feature importance-based explanations and aggregates insights across all samples to achieve a global understanding of the model’s behavior on the entire dataset. Experimental results demonstrate PEXP’s significant advantages over mainstream interpretable methods in both runtime efficiency and the quality of explanations, particularly crucial for Big Data analytics. Furthermore, a case study illustrates PEXP’s application in enhancing the interpretability of video anomaly detection systems within smart cities, a domain characterized by large volumes of data and offers insights for improving Transformer-based architectures.
Keywords:
Big Data
explanation
interpretable machine learning
systems management
Journal
I
IF:
5.7
Papers:
860
Citations:
3.0K

