Homepage · Google Scholar · ORCID
Hi 👋, I'm Yanjun. Second-year PhD at Hong Kong PolyU, joint with EIT Ningbo.
In reinforcement learning, models are trainable. The environments that train them are not.
I'm working on closing that gap.
C3 measures exact credit in cooperative LLM agents. Accuracy Paradox(EMNLP 2024) shows that a better reward model does not always train a better policy.
On the side: OmniSeek, a self-hosted deep-retrieval engine for AI agents, and Myco, persistent memory infrastructure. Both in the MCP ecosystem.
Also into board games 🎲, jazz 🎷, karaoke 🎤, and old-school anime that never quite went away.
Best reached by email.

