Action-Driven Semantic Representation and Aggregation for Video Captioning2025-04-010 PRE AI DOI:10.1109/tcsvt.2024.3502736原文链接原文求助分享收藏摘要 En