PARL (Parallel-Agent Reinforcement Learning) is a training paradigm that teaches models to decompose complex tasks into parallel subtasks and coordinate multiple agents simultaneously.
reinforcement-learning ai ml multi-agent rl swarms agents kimi rlhf agentic deepseek grpo grpotrainer subagents moonshotai
-
Updated
Mar 24, 2026 - Python