Journal of Guangxi Normal University(Natural Science Edition) ›› 2026, Vol. 44 ›› Issue (5): 75-86.doi: 10.16088/j.issn.1001-6600.2026012501

• Intelligence Information Processing • Previous Articles     Next Articles

Multi-target tracking based on cross-modal semantic prompts

Zhang Canlong1,2*, Xu Bofan1, Huang Hongjin1, Lu Xiaochun1, Wei Chunrong3   

  1. 1. Key Lab of Education Blockchain and Intelligent Technology, Ministry of Education (Guangxi Normal University), Guilin Guangxi 541004, China;
    2. Guangxi Key Lab of Multi-source Information Mining & Security (Guangxi Normal University), Guilin Guangxi 541004, China;
    3. Teachers College for Vocational and Technical Education, Guangxi Normal University, Guilin Guangxi 541004, China
  • Received:2026-01-25 Revised:2026-03-24 Online:2026-09-05 Published:2026-07-24

Abstract: An end-to-end language-prompt multi-object tracking framework, termed enLPMOT (enhanced language-prompt multi-object tracking), is proposed to introduce high-level semantic information into the visual tracking process and to enable semantic-aware target perception and continuous association. In complex dynamic scenes, language-prompt multi-object tracking is still severely affected by frequent identity switches, heavy occlusion, and insufficient temporal consistency,which are mainly caused by insufficient diversity of supervision signals and inadequate modeling of cross-frame spatial relationships in existing methods. To address these issues, a grouped-query supervision mechanism is introduced, in which a one-to-many assignment strategy is adopted to enhance supervision diversity. An enhanced dual-path decoding structure is designed, where original queries and noise-enhanced queries are decoded in parallel and adaptively fused, so that the robustness of the model under occlusion and appearance variation is improved. Meanwhile, explicit relative geometric encoding is incorporated to model cross-frame spatial dependencies, thereby strengthening the temporal consistency of trajectories. Experimental results demonstrate that competitive performance is achieved by the proposed method on multiple evaluation metrics. Specifically, the HOTA (higher order tracking accuracy) reaches 40.30%, the DetA (detection accuracy) reaches 30.60%, and the AssR (association recall) and AssP (association precision) reach 55.77% and 83.53%, respectively.

Key words: multi-object tracking, semantic cue tracking, cross-modal fusion, multi-set query supervision, spatial location encoding

CLC Number:  TP391.41
[1] Xiao C C, Cao Q, Luo Z G, et al. MambaTrack: a simple baseline for multiple object tracking with state space model[C]//MM’24 Proceedings of the 32nd ACM International Conference on Multimedia. New York, NY: ACM, 2024: 4082-4091. DOI: 10.1145/3664647.3680944.
[2] 李威澎, 王正友, 张颖, 等. 联合检测和特征嵌入的多目标跟踪[J/OL]. 南京邮电大学学报(自然科学版): 1-10[2026-01-25].https://link.cnki.net/urlid/32.1772.tn.20251113.1002.002.
[3] 于莹莹, 刘博. 多目标跟踪方法综述[J]. 科技与创新, 2025(21): 27-29. DOI: 10.15913/j.cnki.kjycx.2025.21.008.
[4] 朱进, 张向荣, 刘希龙, 等. 基于时空特征融合与注意力机制的多目标航迹预测算法[J]. 信号处理, 2025, 41(11): 1800-1813. DOI: 10.12466/xhcl.2025.11.006.
[5] Carion N, Massa F, Synnaeve G, et al. End-to-end object detection with transformers[C]//Computer Vision-ECCV 2020. Cham: Springer, 2020: 213-229. DOI: 10.1007/978-3-030-58452-8_13.
[6] 张龙. 基于Transformer的多目标跟踪算法研究[D]. 南宁: 广西大学, 2025.
[7] Radford A, Kim J W, Hallacy C, et al. Learning transferable visual models from natural language supervision[C]//Proceedings of the 38th International Conference on Machine Learning:PMLR 139.Cambridge, MA: JMLR, 2021: 8748-8763.
[8] Li J, Selvaraju R, Gotmare A, et al. Align before fuse: vision and language representation learning with momentum distillation[J]. Advances in Neural Information Processing Systems, 2021, 34: 9694-9705.
[9] 孙树正, 潘树国, 胡鹏, 等. 基于视觉-空间特征自适应融合的复杂场景多目标跟踪方法[J]. 激光与光电子学进展, 2026, 63(6): 222-234.DOI:10.3788/LOP251780.
[10] Guan Z Y, Wang Z F, Zhang G, et al. Multi-object tracking review: retrospective and emerging trend[J]. Artificial Intelligence Review, 2025, 58(8): 235. DOI: 10.1007/s10462-025-11212-y.
[11] Liang S F, Guan R W, Lian W W, et al. Cognitive disentanglement for referring multi-object tracking[J]. Information Fusion, 2025, 124: 103349. DOI: 10.1016/j.inffus.2025.103349.
[12] Hou X Q, Liu M Q, Zhang S L, et al. Relation DETR: exploring explicit position relation prior for object detection[C]//Computer Vision-ECCV 2024. Cham:Springer Nature Switzerland,2024: 89-105. DOI: 10.1007/978-3-031-72973-7_6.
[13] 马天, 石炜璐, 苟娜娜, 等. 基于自注意力与匹配优化的煤矿井下人员跟踪方法[J]. 煤炭科学技术, 2026, 54(S1): 386-397.
[14] 苏蕾, 郝斌, 张飞, 等. 融合自适应特征和轨迹预测补偿的鸟类跟踪算法[J/OL]. 光电子·激光: 1-12.[2026-01-25]. https://link.cnki.net/urlid/12.1182.O4.20250716.1715.002.
[15] 陈浩, 刘国强, 胡川, 等. 复杂交通环境下基于特征融合的密集行人检测与跟踪方法[J/OL]. 计算机工程与应用:1-16[2026-01-25].https://link.cnki.net/urlid/11.2127.TP.20251231.1529.011.
[16] Javanmardi M, Qi X J. Appearance variation adaptation tracker using adversarial network[J]. Neural Networks, 2020, 129: 334-343. DOI: 10.1016/j.neunet.2020.06.011.
[17] 张灿龙, 李燕茹, 李志欣, 等. 基于核相关滤波与特征融合的分块跟踪算法[J]. 广西师范大学学报(自然科学版), 2020, 38(5): 12-23. DOI: 10.16088/j.issn.1001-6600.2020.05.002.
[18] Bewley A, Ge Z Y, Ott L, et al. Simple online and realtime tracking[C]//2016 IEEE International Conference on Image Processing (ICIP).Piscataway, NJ: IEEE, 2016: 3464-3468. DOI: 10.1109/ICIP.2016.7533003.
[19] Wojke N, Bewley A, Paulus D. Simple online and realtime tracking with a deep association metric[C]//2017 IEEE International Conference on Image Processing (ICIP). Piscataway,NJ:IEEE, 2017: 3645-3649. DOI: 10.1109/ICIP.2017.8296962.
[20] Wang Z D, Zheng L, Liu Y X, et al. Towards real-time multi-object tracking[C]//Computer Vision-ECCV 2020. Cham: Springer, 2020: 107-122. DOI: 10.1007/978-3-030-58621-8_7.
[21] Zhang Y F, Wang C Y, Wang X G, et al. FairMOT: on the fairness of detection and re-identification in multiple object tracking[J]. International Journal of Computer Vision, 2021, 129(11): 3069-3087. DOI: 10.1007/s11263-021-01513-4.
[22] Zhang Y F, Sun P Z, Jiang Y, et al. ByteTrack: multi-object tracking by associating every detection box[C]//Computer Vision-ECCV 2022. Cham: Springer, 2022: 1-21. DOI: 10.1007/978-3-031-20047-2_1.
[23] Li Y H, Liu X Q, Liu L K, et al. LaMOT: language-guided multi-object tracking[C]//2025 IEEE International Conference on Robotics and Automation (ICRA).Piscataway, NJ: IEEE, 2025: 6816-6822. DOI: 10.1109/ICRA55743.2025.11128553.
[24] Wang J, Lai C W, Wang Y Y, et al. EMAT: Efficient feature fusion network for visual tracking via optimized multi-head attention[J]. Neural Networks, 2024, 172: 106110. DOI: 10.1016/j.neunet.2024.106110.
[25] Meinhardt T, Kirillov A, Leal-Taixé L, et al. TrackFormer: multi-object tracking with transformers[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Los Alamitos,CA: IEEE Compater Society, 2022: 8834-8844. DOI: 10.1109/CVPR52688.2022.00864.
[26] Sun P Z, Cao J K, Jiang Y, et al. TransTrack: multiple object tracking with transformer[PP/OL]. V2. arXiv (2021-05-04)[2026-01-25].https://doi.org/10.48550/arXiv.2012.15460.
[27] Zeng F G, Dong B, Zhang Y A, et al. MOTR: end-to-end multiple-object tracking with Transformer[C]//Computer Vision-ECCV 2022. Cham: Springer, 2022: 659-675. DOI: 10.1007/978-3-031-19812-0_38.
[28] 林家丞, 陈嘉俊, 李智勇, 等. 基于语义概念关联的参考多目标跟踪方法[J]. 自动化学报, 2025, 51(12): 2664-2678. DOI: 10.16383/j.aas.c250118.
[29] Zheng Y Z, Zhong B N, Liang Q H, et al. Toward unified token learning for vision-language tracking[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(4): 2125-2135. DOI: 10.1109/TCSVT.2023.3301933.
[30] Zhang H L, Wang J C, Zhang J W, et al. One-stream vision-language memory network for object tracking[J]. IEEE Transactions on Multimedia, 2024, 26: 1720-1730. DOI: 10.1109/TMM.2023.3285441.
[31] Zhao H J, Wang X, Wang D, et al. Transformer vision-language tracking via proxy token guided cross-modal fusion[J]. Pattern Recognition Letters, 2023, 168: 10-16. DOI: 10.1016/j.patrec.2023.02.023.
[32] Botach A, Zheltonozhskii E, Baskin C. End-to-end referring video object segmentation with multimodal transformers[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Los Alamitos, CA: IEEE Computer Society, 2022: 4975-4985. DOI: 10.1109/CVPR52688.2022.00493.
[33] Wu J N, Jiang Y, Sun P Z, et al. Language as queries for referring video object segmentation[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Los Alamitos, CA: IEEE Computer Society, 2022: 4964-4974. DOI: 10.1109/CVPR52688.2022.00492.
[34] Du Y H, Lei C, Zhao Z C, et al. iKUN: speak to trackers without retraining[C]//2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Los Alamitos, CA: IEEE Computer Society,2024: 19135-19144. DOI: 10.1109/CVPR52733.2024.01810.
[35] Wu D M, Han W C, Wang T C, et al. Referring multi-object tracking[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Los Alamitos, CA: IEEE Computer Society, 2023: 14633-14642. DOI: 10.1109/CVPR52729.2023.01406.
[36] Yang A, Miech A, Sivic J, et al. TubeDETR: spatio-temporal video grounding with transformers[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Los Alamitos, CA: IEEE Computer Society,2022: 16421-16432. DOI: 10.1109/CVPR52688.2022.01595.
[37] Koonce B. ResNet 50[M]//Convolutional Neural Networks with Swift for Tensorflow: Image Recognition and Dataset Categorization. Berkeley, CA: Apress, 2021: 63-72. DOI: 10.1007/978-1-4842-6168-2_6.
[38] Geiger A, Lenz P, Urtasun R. Are we ready for autonomous driving? The KITTI vision benchmark suite[C]//2012 IEEE Conference on Computer Vision and Pattern Recognition.Piscataway, NJ: IEEE, 2012: 3354-3361. DOI: 10.1109/CVPR.2012.6248074.
[39] Sun P Z, Cao J K, Jiang Y, et al. DanceTrack: multi-object tracking in uniform appearance and diverse motion[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Los Alamitos, CA: IEEE Computer Society,2022: 20961-20970. DOI: 10.1109/CVPR52688.2022.02032.
[40] Chen Y, Ding S Y, Guo J H, et al. CSTrack: a comprehensive and concise vision transformer tracker[C]//Pattern Recognition and Computer Vision. Singapore: Springer, 2024: 120-132. DOI: 10.1007/978-981-99-8555-5_10.
[41] Lin J C, Chen J J, Peng K Y, et al. EchoTrack: auditory referring multi-object tracking for autonomous driving[J]. IEEE Transactions on Intelligent Transportation Systems, 2024, 25(11): 18964-18977. DOI: 10.1109/TITS.2024.3437645.
[42] He W Y, Jian Y J, Lu Y, et al. Visual-linguistic representation learning with deep cross-modality fusion for referring multi-object tracking[C]//ICASSP 2024: 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).Piscataway, NJ: IEEE, 2024: 6310-6314. DOI: 10.1109/ICASSP48485.2024.10447535.
[1] Suo Guidong, Lu Zhimin, Li Zili. EMD-YOLO: a PCB defect detection model based on improved YOLO11n [J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(5): 49-62.
[2] Hu Zhiqiang, Lü Xiaoqi, Gu Yu. Skin lesion segmentation model based on improved Mamba local feature acquisition [J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(5): 63-74.
[3] Tang Chenghua, Yi Jianbing, Wu Xin, Xiong Wenwu, Wang Jingyong. A review of cross-domain few-shot image semantic segmentation methods [J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(4): 1-27.
[4] Wang Chenglong, Song Qiang, Li Wenfeng, Zhang Shimin. PAM-DETR: a small-object defect detection algorithm for medical gloves based on improved RT-DETR [J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(4): 79-95.
[5] YANG Yunbo, NAN Xinyuan, CAI Xin. Photovoltaic Panel Defect Detection Method Based on Improved YOLO11n [J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(3): 47-59.
[6] QIAN Junlei, WANG Xizhi, ZENG Kai, DU Xueqiang, LIU He, ZHU Liguang. Steel Surface Defect Detection Algorithm Based on MHTD-YOLO11n [J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(3): 60-74.
[7] BI Huanan, GAO Bingpeng, CAI Xin. SOP-DETR: An Underwater Garbage Detection Algorithm Based on Improved RT-DETR [J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(3): 75-88.
[8] WANG Yan, XU Jie, NIU Mengyuan. Multi-scale Underwater Image Enhancement Network with Adaptive Normalization [J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(3): 89-106.
[9] TIAN Sheng, ZHAO Kailong, MIAO Jialin. Research on Automatic Driving Road Traffic Detection Algorithm Based on Improved YOLO11n Model [J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(1): 1-9.
[10] HUANG Wenjie, LUO Weiping, CHEN Zhennan, PENG Zhixiang, DING Zihao. Research on Lightweight PCB Defect Detection Algorithm Based on YOLO11 [J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(1): 56-67.
[11] LI Fengwei, TAN Yumei, SONG Shuxiang, XIA Haiying. Occlusion-Aware Facial Expression Recognition Based on Attention Guidance [J]. Journal of Guangxi Normal University(Natural Science Edition), 2025, 43(5): 104-113.
[12] LIU Tinghan, LIANG Yan, HUANG Pengsheng, BI Jinjie, HUANG Shoulin, LI Tinghui. Facial Acne Detection for Small Object Based on Improved YOLOv8s [J]. Journal of Guangxi Normal University(Natural Science Edition), 2025, 43(5): 114-129.
[13] YI Jianbing, ZHANG Yuxian, CAO Feng, LI Jun, PENG Xin, CHEN Xin. Design of 3D Human Pose Estimation Network Based on Spatio-Temporal Attention [J]. Journal of Guangxi Normal University(Natural Science Edition), 2025, 43(5): 130-144.
[14] TIAN Sheng, XIONG Chenyin, LONG Anyang. Point Cloud Classification Method of Urban Roads Based on Improved PointNet++ [J]. Journal of Guangxi Normal University(Natural Science Edition), 2025, 43(4): 1-14.
[15] LI Zhixin, KUANG Wenlan. Fine-grained Image Classification Combining Adaptive Spatial Mutual Attention and Feature Pair Integration Discrimination [J]. Journal of Guangxi Normal University(Natural Science Edition), 2025, 43(4): 69-82.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
[1] Tang Chenghua, Yi Jianbing, Wu Xin, Xiong Wenwu, Wang Jingyong. A review of cross-domain few-shot image semantic segmentation methods[J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(4): 1 -27 .
[2] Tian Sheng, Xie Hualin, Chen Dong. Energy management strategy for fuel cell vehicles based on improved deep reinforcement learning[J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(4): 28 -45 .
[3] Zhang Xu, Liu Didi. Intelligent charging/discharging scheduling strategy for electric vehicles based on TD3 algorithm[J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(4): 46 -55 .
[4] Yan Yuanyang, Xie Lirong, Zhang Longjun, Ren Juan, Huang Chenchen, Hu Chao. Ultra-short-term wind power prediction model based on multi-objective optimization[J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(4): 56 -70 .
[5] Lü Hui, Su Jing, Xiong Feng, Zhang Duanyu, Chang Wenhan, Wang Can, Ma Hui. Bi-level coordinated optimization scheduling method for microgrid clusters based on improved SAC algorithm[J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(5): 1 -15 .
[6] Yang Zhen, Tang Yue, Geng Zhaojie, Yin Xu, Huang Yong. Giant magnetoimpedance biosensor based on composite amorphous wire for sensitive detection of cTnI[J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(5): 16 -26 .
[7] Tian Peiyi, Jiang Pinqun, Song Shuxiang, Xia Haiying, Cai Chaobo. High-efficiency and fast-stabilizing boost charge pump controlled by multi-phase clock[J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(5): 27 -37 .
[8] Chen Geng, Song Shuxiang, Jiang Pinqun, Cai Chaobo. Design of 12 bit 100 MS/s successive approximation analog-to-digital converter[J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(5): 38 -48 .
[9] Suo Guidong, Lu Zhimin, Li Zili. EMD-YOLO: a PCB defect detection model based on improved YOLO11n[J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(5): 49 -62 .
[10] Hu Zhiqiang, Lü Xiaoqi, Gu Yu. Skin lesion segmentation model based on improved Mamba local feature acquisition[J]. Journal of Guangxi Normal University(Natural Science Edition), 2026, 44(5): 63 -74 .