广西师范大学学报(自然科学版) ›› 2026, Vol. 44 ›› Issue (5): 75-86.doi: 10.16088/j.issn.1001-6600.2026012501

• 智能信息处理 • 上一篇    下一篇

基于跨模态语义提示的多目标跟踪

张灿龙1,2*, 徐柏帆1, 黄洪锦1, 卢小春1, 韦春荣3   

  1. 1.教育区块链与智能技术教育部重点实验室(广西师范大学), 广西 桂林 541004;
    2.广西多源信息挖掘与安全重点实验室(广西师范大学), 广西 桂林 541004;
    3.广西师范大学 职业技术师范学院, 广西 桂林 541004
  • 收稿日期:2026-01-25 修回日期:2026-03-24 出版日期:2026-09-05 发布日期:2026-07-24
  • 通讯作者: 张灿龙(1975—), 男, 湖南双峰人, 广西师范大学教授, 博士, 博导。E-mail: zcltyp@163.com
  • 基金资助:
    国家自然科学基金(62266009);广西重点研发计划(2024AB26006)

Multi-target tracking based on cross-modal semantic prompts

Zhang Canlong1,2*, Xu Bofan1, Huang Hongjin1, Lu Xiaochun1, Wei Chunrong3   

  1. 1. Key Lab of Education Blockchain and Intelligent Technology, Ministry of Education (Guangxi Normal University), Guilin Guangxi 541004, China;
    2. Guangxi Key Lab of Multi-source Information Mining & Security (Guangxi Normal University), Guilin Guangxi 541004, China;
    3. Teachers College for Vocational and Technical Education, Guangxi Normal University, Guilin Guangxi 541004, China
  • Received:2026-01-25 Revised:2026-03-24 Online:2026-09-05 Published:2026-07-24

摘要: 语义提示多目标跟踪旨在通过语义提示将高层语义信息引入视觉跟踪过程,实现对目标的语义感知与持续关联。然而,在复杂动态场景中,该任务仍易受到目标身份频繁切换、严重遮挡以及时间一致性不足等问题的影响,其根源在于现有方法监督信号多样性不足以及跨帧空间关系建模能力有限。为此,本文提出一种端到端语义提示多目标跟踪框架 enLPMOT(enhanced language-prompt multi-object tracking)。该方法通过分组查询监督机制引入一对多分配策略,以提升监督多样性;设计增强双路径解码结构,通过原始查询与噪声增强查询的并行解码与自适应融合,提高模型在遮挡与外观变化场景下的鲁棒性;同时引入显式相对几何编码对跨帧空间依赖关系进行建模,以增强轨迹的时间一致性。实验结果表明,本文方法在多个评测指标上均取得具有竞争力的性能。其中,高阶跟踪准确率(higher order tracking accuracy, HOTA)达到40.30%,检测准确率(detection accuracy, DetA)为30.60%,关联召回率(association recall, AssR)与关联精确率(association precision, AssP)分别达到 55.77%和83.53%。

关键词: 多目标跟踪, 语义提示跟踪, 跨模态融合, 多组查询监督, 空间位置编码

Abstract: An end-to-end language-prompt multi-object tracking framework, termed enLPMOT (enhanced language-prompt multi-object tracking), is proposed to introduce high-level semantic information into the visual tracking process and to enable semantic-aware target perception and continuous association. In complex dynamic scenes, language-prompt multi-object tracking is still severely affected by frequent identity switches, heavy occlusion, and insufficient temporal consistency,which are mainly caused by insufficient diversity of supervision signals and inadequate modeling of cross-frame spatial relationships in existing methods. To address these issues, a grouped-query supervision mechanism is introduced, in which a one-to-many assignment strategy is adopted to enhance supervision diversity. An enhanced dual-path decoding structure is designed, where original queries and noise-enhanced queries are decoded in parallel and adaptively fused, so that the robustness of the model under occlusion and appearance variation is improved. Meanwhile, explicit relative geometric encoding is incorporated to model cross-frame spatial dependencies, thereby strengthening the temporal consistency of trajectories. Experimental results demonstrate that competitive performance is achieved by the proposed method on multiple evaluation metrics. Specifically, the HOTA (higher order tracking accuracy) reaches 40.30%, the DetA (detection accuracy) reaches 30.60%, and the AssR (association recall) and AssP (association precision) reach 55.77% and 83.53%, respectively.

Key words: multi-object tracking, semantic cue tracking, cross-modal fusion, multi-set query supervision, spatial location encoding

中图分类号:  TP391.41

[1] Xiao C C, Cao Q, Luo Z G, et al. MambaTrack: a simple baseline for multiple object tracking with state space model[C]//MM’24 Proceedings of the 32nd ACM International Conference on Multimedia. New York, NY: ACM, 2024: 4082-4091. DOI: 10.1145/3664647.3680944.
[2] 李威澎, 王正友, 张颖, 等. 联合检测和特征嵌入的多目标跟踪[J/OL]. 南京邮电大学学报(自然科学版): 1-10[2026-01-25].https://link.cnki.net/urlid/32.1772.tn.20251113.1002.002.
[3] 于莹莹, 刘博. 多目标跟踪方法综述[J]. 科技与创新, 2025(21): 27-29. DOI: 10.15913/j.cnki.kjycx.2025.21.008.
[4] 朱进, 张向荣, 刘希龙, 等. 基于时空特征融合与注意力机制的多目标航迹预测算法[J]. 信号处理, 2025, 41(11): 1800-1813. DOI: 10.12466/xhcl.2025.11.006.
[5] Carion N, Massa F, Synnaeve G, et al. End-to-end object detection with transformers[C]//Computer Vision-ECCV 2020. Cham: Springer, 2020: 213-229. DOI: 10.1007/978-3-030-58452-8_13.
[6] 张龙. 基于Transformer的多目标跟踪算法研究[D]. 南宁: 广西大学, 2025.
[7] Radford A, Kim J W, Hallacy C, et al. Learning transferable visual models from natural language supervision[C]//Proceedings of the 38th International Conference on Machine Learning:PMLR 139.Cambridge, MA: JMLR, 2021: 8748-8763.
[8] Li J, Selvaraju R, Gotmare A, et al. Align before fuse: vision and language representation learning with momentum distillation[J]. Advances in Neural Information Processing Systems, 2021, 34: 9694-9705.
[9] 孙树正, 潘树国, 胡鹏, 等. 基于视觉-空间特征自适应融合的复杂场景多目标跟踪方法[J]. 激光与光电子学进展, 2026, 63(6): 222-234.DOI:10.3788/LOP251780.
[10] Guan Z Y, Wang Z F, Zhang G, et al. Multi-object tracking review: retrospective and emerging trend[J]. Artificial Intelligence Review, 2025, 58(8): 235. DOI: 10.1007/s10462-025-11212-y.
[11] Liang S F, Guan R W, Lian W W, et al. Cognitive disentanglement for referring multi-object tracking[J]. Information Fusion, 2025, 124: 103349. DOI: 10.1016/j.inffus.2025.103349.
[12] Hou X Q, Liu M Q, Zhang S L, et al. Relation DETR: exploring explicit position relation prior for object detection[C]//Computer Vision-ECCV 2024. Cham:Springer Nature Switzerland,2024: 89-105. DOI: 10.1007/978-3-031-72973-7_6.
[13] 马天, 石炜璐, 苟娜娜, 等. 基于自注意力与匹配优化的煤矿井下人员跟踪方法[J]. 煤炭科学技术, 2026, 54(S1): 386-397.
[14] 苏蕾, 郝斌, 张飞, 等. 融合自适应特征和轨迹预测补偿的鸟类跟踪算法[J/OL]. 光电子·激光: 1-12.[2026-01-25]. https://link.cnki.net/urlid/12.1182.O4.20250716.1715.002.
[15] 陈浩, 刘国强, 胡川, 等. 复杂交通环境下基于特征融合的密集行人检测与跟踪方法[J/OL]. 计算机工程与应用:1-16[2026-01-25].https://link.cnki.net/urlid/11.2127.TP.20251231.1529.011.
[16] Javanmardi M, Qi X J. Appearance variation adaptation tracker using adversarial network[J]. Neural Networks, 2020, 129: 334-343. DOI: 10.1016/j.neunet.2020.06.011.
[17] 张灿龙, 李燕茹, 李志欣, 等. 基于核相关滤波与特征融合的分块跟踪算法[J]. 广西师范大学学报(自然科学版), 2020, 38(5): 12-23. DOI: 10.16088/j.issn.1001-6600.2020.05.002.
[18] Bewley A, Ge Z Y, Ott L, et al. Simple online and realtime tracking[C]//2016 IEEE International Conference on Image Processing (ICIP).Piscataway, NJ: IEEE, 2016: 3464-3468. DOI: 10.1109/ICIP.2016.7533003.
[19] Wojke N, Bewley A, Paulus D. Simple online and realtime tracking with a deep association metric[C]//2017 IEEE International Conference on Image Processing (ICIP). Piscataway,NJ:IEEE, 2017: 3645-3649. DOI: 10.1109/ICIP.2017.8296962.
[20] Wang Z D, Zheng L, Liu Y X, et al. Towards real-time multi-object tracking[C]//Computer Vision-ECCV 2020. Cham: Springer, 2020: 107-122. DOI: 10.1007/978-3-030-58621-8_7.
[21] Zhang Y F, Wang C Y, Wang X G, et al. FairMOT: on the fairness of detection and re-identification in multiple object tracking[J]. International Journal of Computer Vision, 2021, 129(11): 3069-3087. DOI: 10.1007/s11263-021-01513-4.
[22] Zhang Y F, Sun P Z, Jiang Y, et al. ByteTrack: multi-object tracking by associating every detection box[C]//Computer Vision-ECCV 2022. Cham: Springer, 2022: 1-21. DOI: 10.1007/978-3-031-20047-2_1.
[23] Li Y H, Liu X Q, Liu L K, et al. LaMOT: language-guided multi-object tracking[C]//2025 IEEE International Conference on Robotics and Automation (ICRA).Piscataway, NJ: IEEE, 2025: 6816-6822. DOI: 10.1109/ICRA55743.2025.11128553.
[24] Wang J, Lai C W, Wang Y Y, et al. EMAT: Efficient feature fusion network for visual tracking via optimized multi-head attention[J]. Neural Networks, 2024, 172: 106110. DOI: 10.1016/j.neunet.2024.106110.
[25] Meinhardt T, Kirillov A, Leal-Taixé L, et al. TrackFormer: multi-object tracking with transformers[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Los Alamitos,CA: IEEE Compater Society, 2022: 8834-8844. DOI: 10.1109/CVPR52688.2022.00864.
[26] Sun P Z, Cao J K, Jiang Y, et al. TransTrack: multiple object tracking with transformer[PP/OL]. V2. arXiv (2021-05-04)[2026-01-25].https://doi.org/10.48550/arXiv.2012.15460.
[27] Zeng F G, Dong B, Zhang Y A, et al. MOTR: end-to-end multiple-object tracking with Transformer[C]//Computer Vision-ECCV 2022. Cham: Springer, 2022: 659-675. DOI: 10.1007/978-3-031-19812-0_38.
[28] 林家丞, 陈嘉俊, 李智勇, 等. 基于语义概念关联的参考多目标跟踪方法[J]. 自动化学报, 2025, 51(12): 2664-2678. DOI: 10.16383/j.aas.c250118.
[29] Zheng Y Z, Zhong B N, Liang Q H, et al. Toward unified token learning for vision-language tracking[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(4): 2125-2135. DOI: 10.1109/TCSVT.2023.3301933.
[30] Zhang H L, Wang J C, Zhang J W, et al. One-stream vision-language memory network for object tracking[J]. IEEE Transactions on Multimedia, 2024, 26: 1720-1730. DOI: 10.1109/TMM.2023.3285441.
[31] Zhao H J, Wang X, Wang D, et al. Transformer vision-language tracking via proxy token guided cross-modal fusion[J]. Pattern Recognition Letters, 2023, 168: 10-16. DOI: 10.1016/j.patrec.2023.02.023.
[32] Botach A, Zheltonozhskii E, Baskin C. End-to-end referring video object segmentation with multimodal transformers[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Los Alamitos, CA: IEEE Computer Society, 2022: 4975-4985. DOI: 10.1109/CVPR52688.2022.00493.
[33] Wu J N, Jiang Y, Sun P Z, et al. Language as queries for referring video object segmentation[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Los Alamitos, CA: IEEE Computer Society, 2022: 4964-4974. DOI: 10.1109/CVPR52688.2022.00492.
[34] Du Y H, Lei C, Zhao Z C, et al. iKUN: speak to trackers without retraining[C]//2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Los Alamitos, CA: IEEE Computer Society,2024: 19135-19144. DOI: 10.1109/CVPR52733.2024.01810.
[35] Wu D M, Han W C, Wang T C, et al. Referring multi-object tracking[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Los Alamitos, CA: IEEE Computer Society, 2023: 14633-14642. DOI: 10.1109/CVPR52729.2023.01406.
[36] Yang A, Miech A, Sivic J, et al. TubeDETR: spatio-temporal video grounding with transformers[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Los Alamitos, CA: IEEE Computer Society,2022: 16421-16432. DOI: 10.1109/CVPR52688.2022.01595.
[37] Koonce B. ResNet 50[M]//Convolutional Neural Networks with Swift for Tensorflow: Image Recognition and Dataset Categorization. Berkeley, CA: Apress, 2021: 63-72. DOI: 10.1007/978-1-4842-6168-2_6.
[38] Geiger A, Lenz P, Urtasun R. Are we ready for autonomous driving? The KITTI vision benchmark suite[C]//2012 IEEE Conference on Computer Vision and Pattern Recognition.Piscataway, NJ: IEEE, 2012: 3354-3361. DOI: 10.1109/CVPR.2012.6248074.
[39] Sun P Z, Cao J K, Jiang Y, et al. DanceTrack: multi-object tracking in uniform appearance and diverse motion[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Los Alamitos, CA: IEEE Computer Society,2022: 20961-20970. DOI: 10.1109/CVPR52688.2022.02032.
[40] Chen Y, Ding S Y, Guo J H, et al. CSTrack: a comprehensive and concise vision transformer tracker[C]//Pattern Recognition and Computer Vision. Singapore: Springer, 2024: 120-132. DOI: 10.1007/978-981-99-8555-5_10.
[41] Lin J C, Chen J J, Peng K Y, et al. EchoTrack: auditory referring multi-object tracking for autonomous driving[J]. IEEE Transactions on Intelligent Transportation Systems, 2024, 25(11): 18964-18977. DOI: 10.1109/TITS.2024.3437645.
[42] He W Y, Jian Y J, Lu Y, et al. Visual-linguistic representation learning with deep cross-modality fusion for referring multi-object tracking[C]//ICASSP 2024: 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).Piscataway, NJ: IEEE, 2024: 6310-6314. DOI: 10.1109/ICASSP48485.2024.10447535.
[1] 田晟, 冯帅涛, 李嘉. 一种基于复合框架的城市道路场景车辆轨迹提取方法[J]. 广西师范大学学报(自然科学版), 2026, 44(2): 31-51.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
[1] 唐程华, 易见兵, 吴欣, 熊文武, 王敬永. 跨域少样本图像语义分割方法综述[J]. 广西师范大学学报(自然科学版), 2026, 44(4): 1 -27 .
[2] 田晟, 谢华林, 陈东. 基于改进深度强化学习的燃料电池汽车能量管理策略[J]. 广西师范大学学报(自然科学版), 2026, 44(4): 28 -45 .
[3] 张旭, 刘迪迪. 基于TD3算法的电动汽车智能充/放电调度策略[J]. 广西师范大学学报(自然科学版), 2026, 44(4): 46 -55 .
[4] 闫远洋, 谢丽蓉, 张龙军, 任娟, 黄晨晨, 胡超. 基于多目标优化的超短期风电功率预测模型[J]. 广西师范大学学报(自然科学版), 2026, 44(4): 56 -70 .
[5] 吕辉, 苏静, 熊枫, 张端宇, 常文涵, 王灿, 马辉. 基于改进SAC算法的微网群双层协同优化调度方法[J]. 广西师范大学学报(自然科学版), 2026, 44(5): 1 -15 .
[6] 杨真, 唐悦, 耿兆杰, 殷旭, 黄永. 复合非晶丝GMI生物传感器对cTnI的灵敏检测[J]. 广西师范大学学报(自然科学版), 2026, 44(5): 16 -26 .
[7] 田培一, 蒋品群, 宋树祥, 夏海英, 蔡超波. 多相位时钟控制的高效率快速稳定升压电荷泵[J]. 广西师范大学学报(自然科学版), 2026, 44(5): 27 -37 .
[8] 陈庚, 宋树祥, 蒋品群, 蔡超波. 12 bit 100 MS/s 逐次逼近型模数转换器设计[J]. 广西师范大学学报(自然科学版), 2026, 44(5): 38 -48 .
[9] 索贵东, 陆志敏, 李自立. EMD-YOLO:一种基于改进YOLO11n的PCB缺陷检测模型[J]. 广西师范大学学报(自然科学版), 2026, 44(5): 49 -62 .
[10] 胡志强, 吕晓琪, 谷宇. 基于Mamba增强局部特征提取的皮肤病变分割模型[J]. 广西师范大学学报(自然科学版), 2026, 44(5): 63 -74 .
版权所有 © 广西师范大学学报(自然科学版)编辑部
地址:广西桂林市三里店育才路15号 邮编:541004
电话:0773-5857325 E-mail: gxsdzkb@mailbox.gxnu.edu.cn
本系统由北京玛格泰克科技发展有限公司设计开发