Researching multimodal intelligence, robust learning, and human-centered interaction.
I am currently a specially funded postdoctoral researcher at the School of Airspace Science and Engineering, Shandong University (SDU). I received my Ph.D. degree in Engineering from the School of Software, SDU, in June 2024, under the supervision of Prof. Xiangxu Meng, with Prof. Lei Meng as my co-advisor. Before that, I worked on intelligent recommendation algorithms for user interface elements under the supervision of Prof. Chenglei Yang, and I received my Bachelor’s degree from Huazhong University of Science and Technology (HUST) in June 2018.
My research interests lie in multimedia computing and intelligent human-computer interaction. I aim to develop intelligent systems capable of robust perception, understanding, interaction, and execution in complex and open environments. Specifically, my work focuses on multimodal robust representation learning under data-scarce and long-tailed conditions, user intention modeling in complex interactive scenarios, and embodied navigation and interaction in open environments. My research has been applied to scenarios such as in-vehicle interaction, UAV navigation, hematologic tumor early warning, dietary image analysis, and legal document inspection.
My research story centers on building multimodal systems that stay reliable when data are scarce, incomplete, or long-tailed, and that can better understand user intention in complex human-centered scenarios.
Learning stable representations when data are scarce, noisy, incomplete, or long-tailed, so models remain useful in realistic settings.
Inferring user goals and behavior in complex interactive scenarios to make systems more responsive, interpretable, and human-centered.
Building systems that can perceive, navigate, and interact in open environments, especially when the scene and task keep changing.
Translating core methods into practical systems for healthcare, transportation, and document understanding.
I am always happy to swap ideas or start a project together in multimodal learning, human-computer interaction, and embodied intelligence, especially when the problem is messy, real, and a little fun to untangle.
ACM MM
Uses a benchmark plus iterative reasoning so the system can inspect child-oriented videos and explain why a clip may be risky.
EMNLP
Learns when to send a prompt to different expert branches, so vision-language tuning stays balanced when long-tailed examples are rare.
EMNLP
Shows that a shorter dialogue can still capture enough behavioral signal for personality assessment, without making the conversation drag on.
ACM MM
Tracks what users are likely trying to select when the target keeps moving, which is useful for dynamic interaction scenes.
JSCI
Rebuilds missing table information by borrowing signals across modalities and reducing the bias that incomplete columns usually introduce.
UIST
Turns EEG patterns into a simple estimate of when people are entering or leaving flow during interactive tasks.
AAAI
Learns causal links between aligned visual and semantic features so the classifier relies less on spurious correlations.
EPJ Data Science
Balances learning across common and rare legal motives so long-tailed text classification is less biased.
UbiComp/IMWUT
Looks for shared flow moments across participants from EEG in ubiquitous-computing settings.
CVM
Uses privileged cross-modal hints during training to learn stronger image features for long-tailed classification.
JOS
Uses extra cross-modal information during training to guide image features toward harder classes.
CGI'22
Predicts which interface elements each person is likely to prefer, enabling more personalized UI recommendations.
ADVM'21
Compares adversarial training choices for long-tailed recognition and identifies the ones that stay stable in practice.
I have had the pleasure of working with several master's and undergraduate students, either independently or in collaboration with colleagues.
Powered by Jekyll and Minimal Light theme.