Head of Research, WoRV @ MaumAI · Adjunct Professor @ DGIST
I'm an AI research engineer with strong foundations in both research and practical engineering:
- 🔬 Research background: 10+ publications in top-tier conferences (NeurIPS, CVPR, ICLR, AAAI) with 1.9K+ citations
- 🤖 Embodied AI: Leading research on foundation models that unify language, vision, and action for robotics and autonomous driving
- 💻 Engineering experience: Previously shipped commercial AI products at Wrtn reaching 5M+ MAUs
- 🎓 Teaching: Undergraduate coursework in Physical AI at DGIST, School of Undergraduate Studies
- 🍸 Fun fact: I'm a licensed home bartender in my free time — Korea's national Craftsman Bartender license!
The work I find most interesting happens where a research result meets a running system. My focus right now:
- 🦾 Vision-Language-Action models and robotics foundation models
- 🌏 World models for robotics and autonomous driving
- 🧠 Autonomous agents with memory and personalization
- 🔄 Diffusion models and generative AI
- 🚀 Scaling research prototypes to production-ready systems
I love meeting curious minds and sharing ideas. Whether you're a fellow researcher, student, entrepreneur, or just someone interested in AI - I'd be happy to chat!
📅 Book a virtual coffee chat with me - 30 minutes of free-form conversation about AI, research, career paths, or anything that sparks your interest.
🌐 More about me: alohays.github.io · CV
💼 We're hiring! The WoRV team is looking for people to work on Physical AI and robotics - full-time roles, internships, and 전문연구요원 (alternative military service) positions. → recruitment.worv.maum.ai
⭐ Check out some of my open-source work:
Robotics & Embodied AI - worv-ai
- D2E - Scaling vision-action pretraining on desktop data for transfer to Embodied AI (ICLR 2026)
- CostNav - A navigation benchmark for real-world economic-cost evaluation of Physical AI agents
- vla-evaluation-harness - One framework to evaluate any VLA model on any robot simulation benchmark
Desktop agents - open-world-agents
- open-world-agents - Everything you need to build a state-of-the-art foundation multimodal desktop agent, end-to-end
- desktop-env - A real-time, high-frequency, real-world desktop environment for desktop-based ML development
Earlier work
- awesome-visual-representation-learning-with-transformers - A curated "awesome list" from 2021, when Vision Transformer research was just taking off




