Journal of System Simulation ›› 2026, Vol. 38 ›› Issue (7): 2053-2067.doi: 10.16182/j.issn1004731x.joss.25-0879
• Papers • Previous Articles Next Articles
Li Jinwei, Liu Xiaoyang, Ju Rusheng
Received:2025-09-12
Revised:2025-12-25
Online:2026-07-28
Published:2026-07-31
Contact:
Ju Rusheng
CLC Number:
Li Jinwei, Liu Xiaoyang, Ju Rusheng. Research on Temporal Action Localization Methods for Cross-modal Understanding[J]. Journal of System Simulation, 2026, 38(7): 2053-2067.
Table 1
Experimental results on THUMOS14 dataset
| 模型 | t-IoU阈值 | |||||
|---|---|---|---|---|---|---|
| 0.3 | 0.4 | 0.5 | 0.6 | 0.7 | ||
| TCN[ | 33.3 | 25.6 | 15.9 | 9.4 | 21.1 | |
| CDC[ | 40.1 | 29.4 | 23.3 | 13.1 | 7.9 | 22.8 |
| CBR[ | 50.1 | 41.2 | 31.1 | 19.0 | 9.9 | 30.3 |
| SSAD[ | 43.0 | 35.6 | 24.0 | 34.2 | ||
| R-C3D[ | 44.8 | 35.6 | 28.9 | 36.5 | ||
| BSN[ | 53.5 | 45.2 | 36.9 | 28.4 | 20.3 | 36.9 |
| MGG[ | 53.9 | 46.8 | 37.4 | 29.5 | 21.3 | 37.8 |
| BMN[ | 56.1 | 47.4 | 38.8 | 29.7 | 20.6 | 38.5 |
| G-TAD[ | 54.6 | 47.6 | 40.3 | 30.8 | 23.4 | 39.3 |
| TAL[ | 53.2 | 48.5 | 42.8 | 33.8 | 20.8 | 39.8 |
| GNN[ | 57.1 | 49.1 | 40.4 | 31.2 | 23.1 | 40.2 |
| DBG[ | 57.8 | 49.4 | 42.8 | 33.8 | 21.8 | 41.1 |
| A2Net[ | 58.6 | 54.1 | 45.5 | 32.5 | 17.2 | 41.6 |
| BUTA[ | 53.9 | 50.7 | 45.4 | 38.1 | 28.4 | 43.3 |
| GTAN[ | 57.8 | 47.2 | 38.8 | 47.9 | ||
| AFSD[ | 67.3 | 62.4 | 55.5 | 43.7 | 31.1 | 52.0 |
| GAP[ | 69.1 | 57.4 | 32.0 | 53.0 | ||
| TALLFormer[ | 76.0 | 63.2 | 34.5 | 59.2 | ||
| Re2TAL[ | 77.0 | 71.5 | 62.4 | 49.7 | 36.3 | 59.4 |
| Actionformer[ | 82.1 | 77.8 | 71.0 | 59.4 | 43.9 | 66.8 |
| AFAT | 82.5 | 78.7 | 71.6 | 59.6 | 44.8 | 67.5 |
Table 2
Experimental results on ActivityNet-1.3 dataset
| 模型 | t-IoU阈值 | |||
|---|---|---|---|---|
| 0.5 | 0.75 | 0.95 | ||
| TAL[ | 38.2 | 18.3 | 1.3 | 19.3 |
| SCC[ | 40.0 | 17.9 | 4.7 | 20.9 |
| SSN[ | 39.1 | 23.5 | 5.5 | 22.7 |
| BSN[ | 39.1 | 23.5 | 5.5 | 22.7 |
| CDC[ | 45.2 | 26.5 | 0.2 | 23.9 |
| A2Net[ | 43.6 | 28.7 | 3.7 | 25.3 |
| TALLFormer[ | 41.3 | 27.3 | 6.3 | 27.2 |
| RTD-Net[ | 47.2 | 30.7 | 8.6 | 28.8 |
| BMN[ | 50.1 | 34.8 | 8.3 | 31.0 |
| GTAD[ | 50.3 | 34.6 | 9.1 | 31.3 |
| AFSD[ | 52.4 | 35.3 | 6.5 | 31.4 |
| GNN[ | 50.6 | 34.8 | 9.4 | 31.6 |
| GTAN[ | 52.6 | 34.1 | 8.9 | 31.9 |
| TCANet[ | 52.3 | 36.7 | 6.9 | 31.9 |
| VSGN[ | 52.4 | 36.1 | 8.4 | 32.3 |
| Actionformer[ | 53.5 | 36.2 | 8.2 | 35.6 |
| AFAT | 54.6 | 37.2 | 8.3 | 36.2 |
| GAP[ | 56.7 | 37.2 | 9.8 | 36.7 |
| Re2TAL[ | 54.8 | 37.8 | 9.0 | 36.8 |
Table 9
Localization accuracy and performance metrics
| 指标类别 | 指标名称 | 数值 |
|---|---|---|
| 时间分辨率 | 基础分辨率(4帧@30fps)/s | 0.133 |
| level 0(步长1)精度/s | 0.133 | |
| level 1(步长2)精度/s | 0.267 | |
| level 2(步长4)精度/s | 0.533 | |
| level 3(步长8)精度/s | 1.067 | |
| level 4(步长16)精度/s | 2.133 | |
| level 5(步长32)精度/s | 4.267 | |
| 定位能力 | 定位误差(半步长)/s | ±0.067 |
| 最短检测片段/s | 0.05 | |
| 最长检测片段/s | 214.45 | |
| 有效感受野/s | 153 | |
| 推理延迟/s | 0.432 | |
| 定位延迟/s | 18 | |
| 检测输出 | NMS后候选片段数 | 200 |
| 动作类别数 | 20 | |
| Top-K候选数 | 2 000 | |
| 置信度阈值 | 0.001 |
| [1] | Lowe D G. Object Recognition from Local Scale-invariant Features[C]//Proceedings of the Seventh IEEE International Conference on Computer Vision. Piscataway: IEEE, 1999: 1150-1157. |
| [2] | Wang Wenzhuang, Di Xiaoguang, Liu Maozhen, et al. Multi-level Symmetric Semantic Alignment Network for Image-text Matching[J]. Neurocomputing, 2024, 599: 128082. |
| [3] | Liu Meng, Nie Liqiang, Wang Yunxiao, et al. A Survey on Video Moment Localization[J]. ACM Computing Surveys, 2023, 55(9): 188. |
| [4] | Xia Huifen, Zhan Yongzhao. A Survey on Temporal Action Localization[J]. IEEE Access, 2020, 8: 70477-70487. |
| [5] | Yuan Yitian, Mei Tao, Zhu Wenwu. To Find Where You Talk: Temporal Sentence Localization in Video with Attention Based Location Regression[C]//Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence and Thirty-First Innovative Applications of Artificial Intelligence Conference and Ninth AAAI Symposium on Educational Advances in Artificial Intelligence. Palo Alto: AAAI Press, 2019: 9159-9166. |
| [6] | Zhang Chenlin, Wu Jianxin, Li Yin. ActionFormer: Localizing Moments of Actions with Transformers[C]//Computer Vision - ECCV 2022. Cham: Springer Nature Switzerland, 2022: 492-510. |
| [7] | Cheng Feng, Bertasius G. TallFormer: Temporal Action Localization with a Long-memory Transformer[C]//Computer Vision - ECCV 2022. Cham: Springer Nature Switzerland, 2022: 503-521. |
| [8] | Zhao Kunpeng, Miyazaki Asahi, Okita Tsuyoshi. Detecting Informative Channels: ActionFormer[J]. International Journal of Activity and Behavior Computing, 2025, 2025(1): 1-36. |
| [9] | Dosovitskiy A, Beyer L, Kolesnikov A, et al. An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale[C]//ICLR 2021 Conference. New York: ICLR, 2021: 1-21. |
| [10] | Carreira João, Zisserman A. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2017: 4724-4733. |
| [11] | Alwassel Humam, Giancola Silvio, Ghanem Bernard. TSP: Temporally-sensitive Pretraining of Video Encoders for Localization Tasks[C]//2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW). Piscataway: IEEE, 2021: 3166-3176. |
| [12] | Zheng Zhaohui, Wang Ping, Liu Wei, et al. Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression[C]//Proceedings of the Thirty-Fourth AAAI Conference on Artificial Intelligence and the Thirty-Second Conference on Innovative Applications of Artificial Intelligence and the Tenth Symposium on Educational Advances in Artificial Intelligence. Palo Alto: AAAI Press, 2020: 12993-13000. |
| [13] | Lin Chuming, Xu Chengming, Luo Donghao, et al. Learning Salient Boundary Feature for Anchor-free Temporal Action Localization[C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2021: 3319-3328. |
| [14] | Dai Xiyang, Singh B, Zhang Guyue, et al. Temporal Context Network for Activity Localization in Videos[C]//2017 IEEE International Conference on Computer Vision (ICCV). Piscataway: IEEE, 2017: 5727-5736. |
| [15] | Shou Zheng, Chan J, Zareian A, et al. CDC: Convolutional-de-convolutional Networks for Precise Temporal Action Localization in Untrimmed Videos[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2017: 1417-1426. |
| [16] | Gao Jiyang, Yang Zhenheng, Nevatia R. Cascaded Boundary Regression for Temporal Action Detection[C]//Proceedings of the British Machine Vision Conference (BMVC). Durham: BMVA Press, 2017: 52.1-52.11. |
| [17] | Lin Tianwei, Zhao Xu, Shou Zheng. Single Shot Temporal Action Detection[C]//Proceedings of the 25th ACM International Conference on Multimedia. New York: ACM, 2017: 988-996. |
| [18] | Xu Huijuan, Das A, Saenko K. R-C3D: Region Convolutional 3D Network for Temporal Activity Detection[C]//2017 IEEE International Conference on Computer Vision (ICCV). Piscataway: IEEE, 2017: 5794-5803. |
| [19] | Lin Tianwei, Zhao Xu, Su Haisheng, et al. BSN: Boundary Sensitive Network for Temporal Action Proposal Generation[C]//Computer Vision - ECCV 2018. Cham: Springer International Publishing, 2018: 3-21. |
| [20] | Liu Yuan, Ma Lin, Zhang Yifeng, et al. Multi-granularity Generator for Temporal Action Proposal[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2019: 3599-3608. |
| [21] | Lin Tianwei, Liu Xiao, Li Xin, et al. BMN: Boundary-matching Network for Temporal Action Proposal Generation[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV). Piscataway: IEEE, 2019: 3888-3897. |
| [22] | Xu Mengmeng, Zhao Chen, Rojas David S, et al. G-TAD: Sub-graph Localization for Temporal Action Detection[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2020: 10153-10162. |
| [23] | Chao Yuwei, Vijayanarasimhan S, Seybold B, et al. Rethinking the Faster R-CNN Architecture for Temporal Action Localization[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE, 2018: 1130-1139. |
| [24] | Bai Yueran, Wang Yingying, Tong Yunhai, et al. Boundary Content Graph Neural Network for Temporal Action Proposal Generation[C]//Computer Vision – ECCV 2020. Cham: Springer International Publishing, 2020: 121-137. |
| [25] | Lin Chuming, Li Jian, Wang Yabiao, et al. Fast Learning of Temporal Action Proposal via Dense Boundary Generator[C]//Proceedings of the Thirty-Fourth AAAI Conference on Artificial Intelligence and the Thirty-Second Conference on Innovative Applications of Artificial Intelligence and the Tenth Symposium on Educational Advances in Artificial Intelligence. Palo Alto: AAAI Press, 2020: 11499-11506. |
| [26] | Yang Le, Peng Houwen, Zhang Dingwen, et al. Revisiting Anchor Mechanisms for Temporal Action Localization[J]. IEEE Transactions on Image Processing, 2020, 29: 8535-8548. |
| [27] | Long Fuchen, Yao Ting, Qiu Zhaofan, et al. Gaussian Temporal Awareness Networks for Action Localization[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2019: 344-353. |
| [28] | Nag S, Zhu Xiatian, Song Yizhe, et al. Post-processing Temporal Action Detection[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2023: 18837-18845. |
| [29] | Zhao Chen, Liu Shuming, Mangalam K, et al. Re2TAL: Rewiring Pretrained Video Backbones for Reversible Temporal Action Localization[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2023: 10637-10647. |
| [30] | Laurens van der Maaten, Hinton Geoffrey. Visualizing Data Using t-SNE[J]. Journal of Machine Learning Research, 2008, 9(86): 2579-2605. |
| [31] | Fabian Caba Heilbron, Barrios Wayner, Escorcia Victor, et al. SCC: Semantic Context Cascade for Efficient Action Detection[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2017: 3175-3184. |
| [32] | Zhao Yue, Xiong Yuanjun, Wang Limin, et al. Temporal Action Detection with Structured Segment Networks[C]//2017 IEEE International Conference on Computer Vision (ICCV). Piscataway: IEEE, 2017: 2933-2942. |
| [33] | Tan Jing, Tang Jiaqi, Wang Limin, et al. Relaxed Transformer Decoders for Direct Action Proposal Generation[C]//2021 IEEE/CVF International Conference on Computer Vision (ICCV). Piscataway: IEEE, 2021: 13506-13515. |
| [34] | Zhiwu Qing, Su Haisheng, Gan Weihao, et al. Temporal Context Aggregation Network for Temporal Action Proposal Refinement[C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2021: 485-494. |
| [35] | Zhao Chen, Thabet Ali, Ghanem Bernard. Video Self-stitching Graph Network for Temporal Action Localization[C]//2021 IEEE/CVF International Conference on Computer Vision (ICCV). Piscataway: IEEE, 2021: 13638-13647. |
| [1] | Lili An, Tian Xia, Wenbin Yang, Xinbo Wu. Modeling and Simulation of Spaceborne, Near-Spaceborne, and Airborne Integrated Collaborative Remote Sensing System Based on DoDAF [J]. Journal of System Simulation, 2023, 35(5): 936-948. |
| [2] | Liu Wang, Sun Jinyu, Ma Shiwei. A Temporal Action Detection Algorithm Based on Spatio-Temporal Feature Pyramid Network [J]. Journal of System Simulation, 2019, 31(11): 2382-2387. |
| Viewed | ||||||
|
Full text |
|
|||||
|
Abstract |
|
|||||