Researcher at Tencent · Ph.D. in Computer Science, Tsinghua University
Homepage · Google Scholar · Repositories
I work on AIGC and vision-language model training at Tencent, which I joined through the Tencent Qingyun Talent Program. My research spans computer vision, multimodal generation, and model alignment.
- AIGC — text-to-image generation, image editing, and controllable generation
- Vision-Language Models — visual reasoning, multimodal interaction, and agents
- Post-training & Alignment — reward modeling, preference learning, and efficient adaptation
-
DisCo · CVPR 2026
Subject-driven text-to-image generation through visual-textual disentanglement and reward-guided recoupling.
Paper -
DSH-Bench · ECCV 2026
A difficulty- and scenario-aware benchmark for evaluating subject-driven text-to-image generation.
Paper -
Affinity-Guided Queries (AGQ) · ICLR 2025
Efficient query-based neuron segmentation, with a 2–3× inference speedup on the evaluated EM benchmarks.
Paper · Code -
Dense Contrastive Loss (DCL) · BMVC 2022
Dense contrastive learning for instance segmentation, without added inference overhead.
Paper · Code -
Boundary Patch Refinement (BPR) · CVPR 2021
Boundary patch selection and refinement for high-quality instance segmentation.
Paper · Code -
Robust Logo Detection · ACM Multimedia 2021
Data augmentation for robust logo detection in e-commerce imagery; 5th among 36,489 teams in the associated challenge.
Paper · Code
See my homepage for the full publication list and academic background.


