Journal of System Simulation ›› 2026, Vol. 38 ›› Issue (7): 2053-2067.doi: 10.16182/j.issn1004731x.joss.25-0879

• Papers • Previous Articles     Next Articles

Research on Temporal Action Localization Methods for Cross-modal Understanding

Li Jinwei, Liu Xiaoyang, Ju Rusheng   

  1. College of Systems Engineering, National University of Defense Technology, Changsha 410073, China
  • Received:2025-09-12 Revised:2025-12-25 Online:2026-07-28 Published:2026-07-31
  • Contact: Ju Rusheng

Abstract:

To address the problems of insufficient localization accuracy and high model complexity in the temporal action localization (TAL) task for video-text cross-modal understanding, an anchor-free action transformer (AFAT) model was proposed. Based on the anchor-free framework, the local self-attention mechanism of Transformer was introduced to enhance the global modeling capability of temporal features. A multi-scale feature pyramid structure was combined to strengthen the representation of actions with different durations, and a lightweight predictor was adopted to reduce computational redundancy. Experimental results show that the average precision of this model on the THUMOS14 dataset is significantly improved compared with the baseline model, and the model complexity is effectively reduced while maintaining high detection performance.

Key words: cross-modal understanding, temporal action localization, local self-attention, feature pyramid network, prototype system

CLC Number: