Skip to content
View Ironieser's full-sized avatar
🎯
Loop
🎯
Loop
  • Independent Researcher
  • US

Block or report Ironieser

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Ironieser/README.md

👋 Hello! I'm Sixun Dong (Ironieser)

🌐 Homepage💻 GitHub📧 Email🎓 Google Scholar


🎓 About Me

I focus on cutting-edge research in multimodal learning, computer vision, and LLM agents. My work bridges the gap between vision, language, and temporal understanding, with a particular emphasis on weakly supervised learning and efficient model design.

🔬 Research Interests:

  • Multimodal Learning: Vision-Language Models, Cross-modal Understanding
  • Video Understanding: Temporal Analysis, Action Recognition, Weakly Supervised Learning
  • Time Series Analysis: Forecasting with Multimodal Perspectives
  • LLM Agents: Tool Learning, Feature Transformation, Embodied AI
  • Efficient AI: Token Pruning, Model Compression, Fast Inference

🎯 Current Focus: Developing embodied multimodal agents that can see, understand, reason, plan, and execute in open-world scenarios.


🏆 Research Highlights

📚 Selected Publications

🔥 ECCV 2026 - Better Call CineCrew: Consistent Ultra-Long Narrative-to-Film Generation
First Author | Agentic Film Generation Through Film-Oriented Domain Specific Language
[Code]

🔥 ICLR 2026 - MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
First Author | Training-free VLM Inference Speed Up x 1.87
[Paper] [Code]

🔥 ICASSP 2026 - Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
LLM agents for Robust Dysarthric Speech Recognition
[Paper]] [Code]

🔥 NeurIps 2025 - Sculpting Features from Noise: Reward-Guided Hierarchical Diffusion for Task-Optimal Feature Transformation
Auto-regressive and diffusion model for feature engineering
[Paper] [Code]

🔥 WACV 2024 - MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning
Equipping LLMs with multimodal tool-use capabilities
[Paper] [Code]

🔥 3DV 2024 - RoomDesigner: Encoding Anchor-latents for Style-consistent and Shape-compatible Indoor Scene Generation
Indoor scene generation with style and shape consistency
[Paper] [Code]

🔥 CVPR 2023 - Weakly Supervised Video Representation Learning with Unaligned Text for Sequential Videos
First Author | Video-text alignment without frame-level supervision
[Paper] [Code]

🔥 CVPR 2022 Oral🏆 - TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting
First Author | Repetitive action counting with transformer architecture
[Paper] [Code] [Dataset] [YouTube] [Bilibili]

🚀 Current Projects

  • Efficient Vision-Language Models: Token pruning strategies for VLM acceleration
  • Vision-Language Models Evaluation: Better evaluation strategy for VLMs

🛠 Technical Stack

Languages & Frameworks:
PythonPyTorchC++Linux

Research Areas:
Computer VisionMultimodal LearningTime SeriesLLM Agents


💼 Professional Experience

🔬 Current Position

Independent Researcher | US, Tempe | Present
Focus: Multimodal Learning, Computer Vision, LLM Agent*

🏢 Industry Experience

GenAI Research Intern | Zoom Inc. | May 2025 - Aug 2025 Efficient Vision-Language Modeling

Research Intern (Team Leader) | DGene | Nov 2023 - Jan 2024
Co-Speech Gesture & Head Motion Generation

Research Intern (Team Leader) | Transsion Holdings | Apr 2023 - Aug 2023
Audio-Driven Talking Head Video Generation


🎓 Education

🎓 M.S. Computer Science | ShanghaiTech University | 2024
SVIP-Lab, Advisor: Prof. Shenghua Gao

🎓 B.E. Computer Science (Dual Degree) | Dalian University of Technology | 2020
🎓 B.E. Process Equipment & Control Engineering | Dalian University of Technology | 2020


📊 GitHub Stats

Ironieser

🤝 Academic Service

Reviewer for: CVPR 2023+, ICCV 2023+, ECCV 2024+, ACM MM (2023-2025), ACCV (2024), KDD (2024), NeurIPS 2025, ICML 2025, ICLR 2026, TMM, Neural Networks, TKDD


🎵 Currently Listening

Spotify


🐍 Contribution Snake

github-snake

💬 Let's Connect!

"Building the future of multimodal AI, one model at a time."

EmailHomepageGoogle Scholar

Profile Views

Pinned Loading

  1. SvipRepetitionCounting/TransRACSvipRepetitionCounting/TransRACPublic

    (CVPR 2022 Oral) Official implemention: TransRAC

    Python 122 20

  2. svip-lab/WeakSVRsvip-lab/WeakSVRPublic

    (CVPR 2023) Official implemention of the paper "Weakly Supervised Video Representation Learning with Unaligned Text for Sequential Videos"

    Python 32 4

  3. Storyteller_Of_Auto-BarrageStoryteller_Of_Auto-BarragePublic

    自动发送弹幕插件,可编辑多项参数,详细的提示信息。使用本插件请著名出处。

    JavaScript 9 2

  4. Python-Remote-DevelopmentPython-Remote-DevelopmentPublic

    The introduction for configure the remote development with pycharm or vscode.

    5

  5. TimesCLIPTimesCLIPPublic

    The offical repo of "Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives"

    50 2

  6. MMTokMMTokPublic

    [ICLR 2026] The official repo of "MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs"

    Python 48 5