Undergraduate student in Artificial Intelligence at the School of Future Technology, Harbin Institute of Technology (HIT). Expected graduation: September 2028.
I am interested in multimodal large language models and efficient model inference. My current project work emphasizes reproducible experimentation, resource-aware model design, and clear technical documentation.
- Multimodal large language models
- Efficient and resource-aware inference
- Reproducible machine learning systems
QI-Budget studies sample-wise visual-token allocation for vision-language models.
- Formulated visual-budget selection as a five-action routing problem for
Qwen2.5-VL-3B-Instruct-4bit. - Evaluated the frozen routing pipeline on 2,017 ScienceQA image samples.
- Achieved 80.119% accuracy while reducing visual-token usage by 66.854% relative to the full-budget baseline, with no statistically significant accuracy degradation.
- Built a leakage-safe out-of-fold training pipeline with feature ablations, independent testing, paired statistical evaluation, and reproducibility checks.
- Programming: Python, C++
- Machine learning and data: NumPy, SciPy, scikit-learn, Hugging Face Datasets, MLX-VLM
- Tools: Git, GitHub, Conda, LaTeX
For academic or project-related inquiries, contact me at 2024112342@stu.hit.edu.cn. I welcome discussions about undergraduate research and collaboration in multimodal learning and efficient inference.