Hi, I'm JJJYmmm.
I expect to graduate in June 2027. My research interests lie in vision-language pre-training, multimodal agent harness, and deployable robotic systems. Feel free to reach out! 1650675829 [at] qq [dot] com
Hi, I'm JJJYmmm.
I expect to graduate in June 2027. My research interests lie in vision-language pre-training, multimodal agent harness, and deployable robotic systems. Feel free to reach out! 1650675829 [at] qq [dot] com
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
Make any agent harness multimodal-native.
Official implement of paper "Revisiting Multimodal Positional Encoding in Vision–Language Models", ICLR 2026
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Qwen3.8 is the large language model series developed by Qwen team, Alibaba Group.