Return
Evaluating Feature-Based Machine-Learning Models with Post Hoc Explainability for Eye-Tracking-Based Task Type and Workload Inference
DOI:10.3390/ai7080325.png)
Abstract
En 中文
Eye tracking is a valuable behavioral signal for human-centered AI, yet the reliability of feature-based machine-learning models for inferring task type and workload across users and tasks remains uncertain, because experimentally defined workload labels may reflect task type and visual structure as much as cognitive demand. The practical problem is that designers of gaze-adaptive systems need to know which inferences are dependable enough to act on, and reported accuracies alone do not answer this, because the choice of prediction target and validation split can determine the result. This study systematically evaluates feature-based machine-learning models with post hoc explainability across three prediction targets: task type, binary load-versus-rest, and three-level workload. Eye-movement features derived from fixations, saccades, pupils, and blinks were extracted from short temporal windows collected from 54 participants performing attention, visual-spatial, and memory tasks under rest, easy, and difficult conditions, and evaluated using leave-one-subject-out (LOSO) and leave-one-group-out (LOGO) validation. Task type was classified most reliably (85.9% LOSO, 83.4% LOGO), binary load-versus-rest showed moderate, validation-sensitive robustness (81.4% LOSO, 63.9% LOGO), and three-level workload classification was substantially more challenging (56.4% LOSO, 44.3% LOGO). SHAP and statistical analyses consistently identified fixation dispersion, pupil-related measures, and subject-normalized features as the strongest contributors across all three targets. These findings show that prediction target definition, validation strategy, and post hoc explainability jointly determine what can be reliably inferred from gaze-based machine-learning models. Eye tracking alone therefore appears promising for task-type recognition and may support coarse engagement-related inference when the deployment task family is represented during model development, whereas task-independent fine-grained workload estimation remains unsupported by the present evidence.
Keywords:
eye tracking
task-type inference
workload inference
feature-based machine learning
post hoc explainability
human-centered AI
cross-task generalization
Journal
A
IF:
5
Papers:
973
Citations:
941

