Return
Task Graph Maximum Likelihood Estimation for Procedural Activity Understanding in Egocentric Videos
L
G
A
DOI:10.1109/tpami.2026.3689721.png)
Abstract
En 中文
Humans engage daily in procedural activities - goal-oriented sequences of key-steps following certain ordering constraints. Task graphs mined from videos or textual descriptions have recently gained popularity as a human-readable, holistic representation of procedural activities encoding a partial ordering over keysteps, and have shown promise in supporting downstream video understanding tasks. This paper introduces an approach based on gradient-based maximum likelihood optimization of edge weights, which can be used to directly estimate an adjacency matrix and can also be naturally plugged into more complex neural network architectures. We validate the proposed approach on CaptainCook4D, EgoPER, and EgoProceL, which we manually annotate with task graphs as an additional contribution. The three datasets together constitute a new benchmark for task graph learning, where our approach obtains improvements of +14.5%, +10.2% and +13.6% in <inline-formula><tex-math notation="LaTeX">$F_1$</tex-math></inline-formula> score over previous approaches. Thanks to the differentiability of the proposed framework, we also introduce a feature-based approach for predicting task graphs from key-step textual or video embeddings, which exhibits emerging video understanding abilities. Task graphs learned with our approach obtain top performance in the Ego-Exo4D procedure understanding benchmark, including 5 different downstream tasks, with gains of up to +4.61%, +0.10%, +5.02%, +8.62%, and +15.16% in finding Previous Keysteps, Optional Keysteps, Procedural Mistakes, Missing Keysteps, and Future Keysteps, respectively. We finally show significant enhancements to the task of online mistake detection in procedural egocentric videos, achieving gains of +19.8% and +6.4% in the Assembly101-O and EPIC-Tent-O datasets, respectively, compared to the state of the art.
Keywords:
Task graphs
procedural sequences
online mistake detection
video understanding
Journal
IF:
18.6
Papers:
831
Citations:
9.8W
