| [1] |
Lowe D G. Object Recognition from Local Scale-invariant Features[C]//Proceedings of the Seventh IEEE International Conference on Computer Vision. Piscataway: IEEE, 1999: 1150-1157.
|
| [2] |
Wang Wenzhuang, Di Xiaoguang, Liu Maozhen, et al. Multi-level Symmetric Semantic Alignment Network for Image-text Matching[J]. Neurocomputing, 2024, 599: 128082.
|
| [3] |
Liu Meng, Nie Liqiang, Wang Yunxiao, et al. A Survey on Video Moment Localization[J]. ACM Computing Surveys, 2023, 55(9): 188.
|
| [4] |
Xia Huifen, Zhan Yongzhao. A Survey on Temporal Action Localization[J]. IEEE Access, 2020, 8: 70477-70487.
|
| [5] |
Yuan Yitian, Mei Tao, Zhu Wenwu. To Find Where You Talk: Temporal Sentence Localization in Video with Attention Based Location Regression[C]//Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence and Thirty-First Innovative Applications of Artificial Intelligence Conference and Ninth AAAI Symposium on Educational Advances in Artificial Intelligence. Palo Alto: AAAI Press, 2019: 9159-9166.
|
| [6] |
Zhang Chenlin, Wu Jianxin, Li Yin. ActionFormer: Localizing Moments of Actions with Transformers[C]//Computer Vision - ECCV 2022. Cham: Springer Nature Switzerland, 2022: 492-510.
|
| [7] |
Cheng Feng, Bertasius G. TallFormer: Temporal Action Localization with a Long-memory Transformer[C]//Computer Vision - ECCV 2022. Cham: Springer Nature Switzerland, 2022: 503-521.
|
| [8] |
Zhao Kunpeng, Miyazaki Asahi, Okita Tsuyoshi. Detecting Informative Channels: ActionFormer[J]. International Journal of Activity and Behavior Computing, 2025, 2025(1): 1-36.
|
| [9] |
Dosovitskiy A, Beyer L, Kolesnikov A, et al. An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale[C]//ICLR 2021 Conference. New York: ICLR, 2021: 1-21.
|
| [10] |
Carreira João, Zisserman A. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2017: 4724-4733.
|
| [11] |
Alwassel Humam, Giancola Silvio, Ghanem Bernard. TSP: Temporally-sensitive Pretraining of Video Encoders for Localization Tasks[C]//2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW). Piscataway: IEEE, 2021: 3166-3176.
|
| [12] |
Zheng Zhaohui, Wang Ping, Liu Wei, et al. Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression[C]//Proceedings of the Thirty-Fourth AAAI Conference on Artificial Intelligence and the Thirty-Second Conference on Innovative Applications of Artificial Intelligence and the Tenth Symposium on Educational Advances in Artificial Intelligence. Palo Alto: AAAI Press, 2020: 12993-13000.
|
| [13] |
Lin Chuming, Xu Chengming, Luo Donghao, et al. Learning Salient Boundary Feature for Anchor-free Temporal Action Localization[C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2021: 3319-3328.
|
| [14] |
Dai Xiyang, Singh B, Zhang Guyue, et al. Temporal Context Network for Activity Localization in Videos[C]//2017 IEEE International Conference on Computer Vision (ICCV). Piscataway: IEEE, 2017: 5727-5736.
|
| [15] |
Shou Zheng, Chan J, Zareian A, et al. CDC: Convolutional-de-convolutional Networks for Precise Temporal Action Localization in Untrimmed Videos[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2017: 1417-1426.
|
| [16] |
Gao Jiyang, Yang Zhenheng, Nevatia R. Cascaded Boundary Regression for Temporal Action Detection[C]//Proceedings of the British Machine Vision Conference (BMVC). Durham: BMVA Press, 2017: 52.1-52.11.
|
| [17] |
Lin Tianwei, Zhao Xu, Shou Zheng. Single Shot Temporal Action Detection[C]//Proceedings of the 25th ACM International Conference on Multimedia. New York: ACM, 2017: 988-996.
|
| [18] |
Xu Huijuan, Das A, Saenko K. R-C3D: Region Convolutional 3D Network for Temporal Activity Detection[C]//2017 IEEE International Conference on Computer Vision (ICCV). Piscataway: IEEE, 2017: 5794-5803.
|
| [19] |
Lin Tianwei, Zhao Xu, Su Haisheng, et al. BSN: Boundary Sensitive Network for Temporal Action Proposal Generation[C]//Computer Vision - ECCV 2018. Cham: Springer International Publishing, 2018: 3-21.
|
| [20] |
Liu Yuan, Ma Lin, Zhang Yifeng, et al. Multi-granularity Generator for Temporal Action Proposal[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2019: 3599-3608.
|
| [21] |
Lin Tianwei, Liu Xiao, Li Xin, et al. BMN: Boundary-matching Network for Temporal Action Proposal Generation[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV). Piscataway: IEEE, 2019: 3888-3897.
|
| [22] |
Xu Mengmeng, Zhao Chen, Rojas David S, et al. G-TAD: Sub-graph Localization for Temporal Action Detection[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2020: 10153-10162.
|
| [23] |
Chao Yuwei, Vijayanarasimhan S, Seybold B, et al. Rethinking the Faster R-CNN Architecture for Temporal Action Localization[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE, 2018: 1130-1139.
|
| [24] |
Bai Yueran, Wang Yingying, Tong Yunhai, et al. Boundary Content Graph Neural Network for Temporal Action Proposal Generation[C]//Computer Vision – ECCV 2020. Cham: Springer International Publishing, 2020: 121-137.
|
| [25] |
Lin Chuming, Li Jian, Wang Yabiao, et al. Fast Learning of Temporal Action Proposal via Dense Boundary Generator[C]//Proceedings of the Thirty-Fourth AAAI Conference on Artificial Intelligence and the Thirty-Second Conference on Innovative Applications of Artificial Intelligence and the Tenth Symposium on Educational Advances in Artificial Intelligence. Palo Alto: AAAI Press, 2020: 11499-11506.
|
| [26] |
Yang Le, Peng Houwen, Zhang Dingwen, et al. Revisiting Anchor Mechanisms for Temporal Action Localization[J]. IEEE Transactions on Image Processing, 2020, 29: 8535-8548.
|
| [27] |
Long Fuchen, Yao Ting, Qiu Zhaofan, et al. Gaussian Temporal Awareness Networks for Action Localization[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2019: 344-353.
|
| [28] |
Nag S, Zhu Xiatian, Song Yizhe, et al. Post-processing Temporal Action Detection[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2023: 18837-18845.
|
| [29] |
Zhao Chen, Liu Shuming, Mangalam K, et al. Re2TAL: Rewiring Pretrained Video Backbones for Reversible Temporal Action Localization[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2023: 10637-10647.
|
| [30] |
Laurens van der Maaten, Hinton Geoffrey. Visualizing Data Using t-SNE[J]. Journal of Machine Learning Research, 2008, 9(86): 2579-2605.
|
| [31] |
Fabian Caba Heilbron, Barrios Wayner, Escorcia Victor, et al. SCC: Semantic Context Cascade for Efficient Action Detection[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2017: 3175-3184.
|
| [32] |
Zhao Yue, Xiong Yuanjun, Wang Limin, et al. Temporal Action Detection with Structured Segment Networks[C]//2017 IEEE International Conference on Computer Vision (ICCV). Piscataway: IEEE, 2017: 2933-2942.
|
| [33] |
Tan Jing, Tang Jiaqi, Wang Limin, et al. Relaxed Transformer Decoders for Direct Action Proposal Generation[C]//2021 IEEE/CVF International Conference on Computer Vision (ICCV). Piscataway: IEEE, 2021: 13506-13515.
|
| [34] |
Zhiwu Qing, Su Haisheng, Gan Weihao, et al. Temporal Context Aggregation Network for Temporal Action Proposal Refinement[C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE, 2021: 485-494.
|
| [35] |
Zhao Chen, Thabet Ali, Ghanem Bernard. Video Self-stitching Graph Network for Temporal Action Localization[C]//2021 IEEE/CVF International Conference on Computer Vision (ICCV). Piscataway: IEEE, 2021: 13638-13647.
|