系统仿真学报 ›› 2026, Vol. 38 ›› Issue (9): 2518-2534.doi: 10.16182/j.issn1004731x.joss.26-0200

• 专栏:自动驾驶与智能交通系统仿真 • 上一篇    

基于改进RT-DETR的交通场景图像长尾目标检测方法

邓明君1,2, 聂沛华1, 郭军华2,3, 高宁波4   

  1. 1.华东交通大学 交通运输工程学院,江西 南昌 330013
    2.华东交通大学 综合立体交通信息感知与融合江西省重点实验室,江西 南昌 330013
    3.江西飞行学院,江西 南昌 330088
    4.江苏智城慧宁交通科技有限公司,江苏 南京 210014
  • 收稿日期:2026-03-09 修回日期:2026-05-11 出版日期:2026-09-30 发布日期:2026-10-02
  • 第一作者简介:邓明君(1978-),男,副教授,博士,研究方向为智能交通。
  • 基金资助:
    国家自然科学基金(52262048);江西省自然科学基金(20252BAC240366)

Long-tail Object Detection Method in Transportation Scene Images Based on improved RT-DETR

Deng Mingjun1,2, Nie Peihua1, Guo Junhua2,3, Gao Ningbo4   

  1. 1.School of Transportation Engineering, East China Jiaotong University, Nanchang 330013, China
    2.Jiangxi Provincial Key Laboratory of Comprehensive Stereoscopic Traffic Information Perception and Fusion, East China Jiaotong University, Nanchang 330013, China
    3.Jiangxi Flight University, Nanchang 330088, China
    4.Jiangsu Zhicheng Huining Transportation Technology Co. , Ltd. , Nanjing 210014, China
  • Received:2026-03-09 Revised:2026-05-11 Online:2026-09-30 Published:2026-10-02

摘要:

检测交通环境中的稀有目标对于自动驾驶系统的安全性和可靠性至关重要。针对现有方法对于训练中仅出现极少次数的稀有目标极端情况检测能力有限的问题,提出基于RT-DETR的交通场景图像中长尾目标检测方法TF-DETR (tail free DETR)。构建轻量化的CSP-EDA (CSP-efficient depthwise aggregation)用以重构骨干网络,其核心单元高效深度聚合模块EDA结合池化-转置注意力与卷积门控单元,实现了对尾部样本特征的增强。在编码器中设计空间自适应动态认知交互模块,通过两阶段之间的特征复用,建立针对尾部目标从特征提取到多尺度认知增强的高效信息交互机制。引入风车卷积对特征图进行下采样,通过非对称填充与多方向卷积核的协同,在轻量化设计的基础上增强对尾部目标的检测能力。在CODA2022数据集上的实验结果表明,TF-DETR相较于基线模型AP50实现了4.9%的提升,AR50提升了5.4%,计算量减少了22.9%,参数量减少了25.7%。

关键词: 交通场景图像, 目标检测, 深度学习, 长尾分布, RT-DETR

Abstract:

Detecting rare objects in transportation environments is crucial for the safety and reliability of autonomous driving systems. To address the challenges of limited detection capability on rare objects that appear only a few times during training, we propose an improved long-tailed object detection method for traffic scene images named TF-DETR (tail free DETR) based on RT-DETR model. First, the lightweight CSP-EDA (CSP-efficient depthwise aggregation) was constructed toreconstruct the backbone network. Its core unit, the efficient depthwise aggregation, integrated the pooled-transpose attention and convolutional gated linear unit to enhance feature representations for tail-class samples. Second,the spatially-adaptive dynamic cognitive interaction was deployed in the encoder. Through feature reuse between the two-stage, the interaction module established an efficient information interaction mechanism. The mechanism spanned from feature extraction to multi-scale cognitive enhancement for tail-class objects. Finally, the pinwheel-shaped convolution wasintroduced for feature map downsampling. Through its collaborative design of asymmetric padding and multi-directional kernels, the model enhanced detection capability for tail-class objects while maintaining a lightweight architecture. Experimental results on the CODA2022 dataset show that TF-DETR achieves a 4.9% increase in AP50 and a 5.4% increase in AR50 compared with the baseline. Meanwhile, a 22.9% reduction in FLOPs and a 25.7% decrease in Parameters are realized.

Key words: transportation scene images, object detection, deep learning, long-tail distribution, RT-DETR

中图分类号: