Lecturer | Researcher in Multimodal Data Fusion School of Computing and Artificial Intelligence, Guangdong Polytechnic Normal University
I am currently a Lecturer at the School of Computing and Artificial Intelligence, Guangdong Polytechnic Normal University. I received my Ph.D. degree in Electronics Information from Sun Yat-sen University.
My primary research interests lie at the intersection of Computer Vision and Multimodal Data Fusion, with specific applications deployed in embodied AI, vehicle-infrastructure cooperation (V2X), and remote sensing. I am dedicated to developing robust perception systems for complex and real-world open scenarios.
I maintain active and long‑term scientific collaborations with researchers from Sun Yat-sen University, the University of Cambridge, and Peng Cheng Laboratory.
- Research on Roadside Multi-View 3D Cooperative Perception under Joint Environmental Noise and Communication Constraints, National Natural Science Foundation of China Young Scientists Fund, 2027–2029.
- Research on an Electricity Inspection Health Diagnosis Model Based on a VAE–Transformer Architecture, enterprise-commissioned project, Jul. 2026–Dec. 2026.
- Industrial Dynamic Visual Tracking Algorithm Software, enterprise-commissioned project, Jun. 2025–Oct. 2025.
- Research on Multimodal Heterogeneous Feature Fusion and Key Technologies, special research project, Oct. 2024–Oct. 2025.
- Deep Spatial–Spectral Feature Extraction for Hyperspectral Remote Sensing Images and Dynamic Monitoring of Dongting Lake Waters, Hunan Provincial Key Graduate Student Project, 2018.
Note: My research philosophy strongly supports open science. Datasets and codebases associated with my publications are made publicly available where possible. For a comprehensive and up-to-date list of my publications, please refer to my Google Scholar Profile.
[CVPR '26🏆 Highlight] Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark. Seng Nam Chen, Hao Chen, Chenglam Ho, Xinyu Mao, Jinping Wang, Yu Zhang, Chao Li.
Introduces SceneBench to evaluate scene-level long video understanding in VLMs, and proposes Scene-RAG to effectively mitigate long-context forgetting via dynamic scene memory.🔗 Paper🔗 Dataset & Code
[IEEE TMM '26] Cognidrive: Cognitive Autonomous Driving Understanding with Multistep Multimodal Chain-of-Thought Reasoning. Xiangyi Qin, Xiaofei Zhang, Shuai Wang, Yuzhen Wei, Jinping Wang*, Xiaojun Tan.
Introduces a multistep multimodal chain-of-thought reasoning framework that integrates environmental perception, semantic understanding, and decision-making to enhance the cognitive capabilities and interpretability of autonomous driving systems. 📖 🔗 Paper
[Information Fusion '26] Inscope: A new real-world 3d infrastructure-side collaborative perception dataset for open traffic scenarios, Xiaofei Zhang#, Yining Li#, Jinping Wang#, Xiangyi Qin, Ying Shen, Zhengping Fan, Xiaojun Tan
A large-scale dataset featuring multi-position LiDARs in a real-world setting, addressing occlusion challenges within I2I perception systems.🔗 Paper🔗 Dataset & Code
[ADVEI '26] Confidence-V2X: Confidence-driven sparse communication for efficient V2X cooperative perception. Xiaojun Tan, Rui Wang, Jinping Wang*, Shuai Wang, Xu Wang, Dongsheng Wu.
A confidence‑aware cooperative perception framework designed to jointly optimize object detection performance and communication efficiency in V2X systems.🔗 Paper🔗 Code
[IEEE TGRS '25] FusDreamer: Label-efficient remote sensing world model for multimodal data classification. Jinping Wang, Weiwei Song, Hao Chen, Jinchang Ren, Huimin Zhao.
A label-efficient remote sensing world model for multimodal data fusion, exploring the potential of the world model in the RS field.🔗 Paper🔗 Code
[IEEE GRSL '25] CaPaT: Cross-Aware Paired-Affine Transformation for Multimodal Data Fusion Network. Jinping Wang, Hao Chen, Xiaofei Zhang, Weiwei Song.
Introduces a direct feature interaction paradigm to improve the transfer efficiency of feature fusion while significantly reducing model parameters.🔗 Paper
[ICASSP '24] BEVLOC: End-to-end 6-dof localization via cross-modality correlation under bird’s eye view. Nanjie Chen, Jinping Wang, Hao Chen, Ying Shen, Shuai Wang, Xiaojun Tan
An end-to-end approach for vehicle localization that fuses monocular image and LiDAR map features in the BEV space via optical flow-based cross-modality correlation.🔗 Paper
[IEEE TCSVT '23] Mutually beneficial transformer for multimodal data fusion. Jinping Wang, Xiaojun Tan.
Introduces dynamic region-aware convolution for spatial guide mask generation and elevation salience agent guidance.🔗 Paper
[IEEE TCSVT '22] AM³Net: Adaptive mutual-learning-based multimodal data fusion network. Jinping Wang, Jun Li, Yanli Shi, Jianhuang Lai, Xiaojun Tan.
Collaborative feature transmission focusing on the specificity of HSI spectral channels and the complementarity of HSI and LiDAR spatial information.🔗 Paper🔗 Code
[IEEE ICASSP '22] Spectral-spatial symmetrical aggregation cross-linking multi-modal data fusion network. Jinping Wang, Jun Li, Xiaojun Tan.
Develops a SACLNet utilizing involution operations and pyramid feature fusion for robust multi-modal data classification.🔗 Paper
[Neurocomputing '22] ASPCNet: Deep adaptive spatial pattern capsule network for hyperspectral image classification. Jinping Wang, Xiaojun Tan, Jianhuang Lai, Jun Li.
An adaptive spatial pattern capsule network architecture based on an enlarged, semantically-adaptive receptive field.🔗 Paper🔗 Code
[Electronics '21] A simulated annealing algorithm and grid map-based UAV coverage path planning method for 3D reconstruction. Sichen Xiao, Xiaojun Tan, Jinping Wang.
Proposes a UAV CPP framework considering both image overlapping and energy efficiency, validated through site experiments.🔗 Paper
[IEEE TGRS '19] Spatial density peak clustering for hyperspectral image classification with noisy labels. Bing Tu, Xiaofei Zhang, Xudong Kang, Jinping Wang, Jón Atli Benediktsson.
Proposes a spatial density peak (SDP) clustering-based method to detect and handle mislabeled samples in HSI training sets.🔗 Paper🔗 Code
[IEEE JSTARS '19] Texture pattern separation for hyperspectral image classification. Bing Tu#, Jinping Wang#, Guoyun Zhang, Xiaofei Zhang, Wei He.
Addresses the layer-separation problem in HSI via a novel TPS feature extraction method.🔗 Paper🔗 Code
[IEEE JSTARS '18] KNN-based representation of superpixels for hyperspectral image classification. Bing Tu#, Jinping Wang#, Xudong Kang, Guoyun Zhang, Xianfeng Ou, Longyuan Guo.
Explores optimal representations of superpixels using two k-selection rules to find the most representative samples.🔗 Paper🔗 Code
[IEEE GRSL '18] Hyperspectral image classification via fusing correlation coefficient and joint sparse representation. Bing Tu, Xiaofei Zhang, Xudong Kang, Guoyun Zhang, Jinping Wang, Jianhui Wu.
A hyperspectral image classification method via fusing correlation coefficient and joint sparse representation.🔗 Paper🔗 Code
- Image Classification Method and System Using a Deep Capsule Network Based on Adaptive Spatial Patterns. Patent No. CN112766340B, Jun. 4, 2024. (Granted)
- Few-Shot Remote Sensing World Model for Multimodal Data Fusion. Patent Application No. 202510038010.2, Jan. 10, 2025. (Application Accepted)
- Cross-Domain Masked Autoencoder Method and System for Multimodal Data Fusion. Patent Application No. 202510960900.9, Jul. 12, 2025. (Application Accepted)
- Feature-Adaptive Mutual-Guidance Method and System for Multi-Source Information Fusion and Classification. Patent No. CN114187526B, Apr. 29, 2025. (Granted)
- Multimodal Remote Sensing Image Processing System. Software Copyright Registration No. 2025R11L0356652, Mar. 10, 2025.
- Deep Learning-Based Waste Classification and Recognition System. Software Copyright Registration No. 2025SR1736527, Sep. 9, 2025.
- Reviewer for
IEEE TMM,IEEE TCSVT,IEEE TGRS,IEEE JSTARS,Knowledge-Based Systems (KBS),CVPR, andIEEE TIM.
- Fundamentals of Artificial Intelligence
- Computer Networks
- Artificial Intelligence and the Information Society
| Year | Competition | Award |
|---|---|---|
| 2025 | Lanqiao Cup (Python Group B) | Provincial / Second Prize |
| 2025 | Future Cup Big Data Challenge | National / Third Prize |
| 2025 | MathorCup Mathematical Modeling Challenge (Freight Volume Forecasting) | National / Successful Participant |
| 2025 | Chinese Society for Electrical Engineering Cup | National / Successful Participant |
| 2025 | Shuwei Cup Mathematical Modeling Challenge (Spring) | National / Second Prize |
| 2025 | APMCM Asia-Pacific Mathematical Contest in Modeling (Chinese Contest) | National / Third Prize |
| 2025 | National College Student Simulation Modeling Challenge | National / Second Prize |
| 2025 | Shuwei Cup Mathematical Modeling Challenge (Autumn) | National / First Prize |
| 2024 | Mathematical Contest in Modeling (MCM) | International / Successful Participant |
| 2024 | MathorCup Big Data Competition | National / Successful Participant |
| 2024 | APMCM Asia-Pacific Mathematical Contest in Modeling (Pet Industry) | National / Third Prize |
- Cimy_PPtools: A comprehensive Python toolbox designed for data preprocessing in hyperspectral classification and fusion tasks, supporting model serialization and result visualization.
✉️ Open to scientific cooperation and academic inquiries. Please feel free to reach out.

