Skip to content
@AMAP-ML

AMAP-ML

DreamX

AMAP Spatial Intelligence Models and Systems

Understand, predict, simulate, create, plan, and act in the real world.
DreamX is the unified spatial intelligence model and system portfolio of AMAP-ML, the AI team at Alibaba AMAP. We connect research, engineering, and real-world deployment across maps, mobility, local services, digital content, and interactive worlds.

GitHub followersJoin Us

We advance DreamX through production systems, open-source projects, benchmarks, and publications at ICLR, CVPR, ECCV, ACL, AAAI, SIGGRAPH, ICCV, ICML, KDD, EMNLP, ACM MM, and WWW. We release code and evaluation assets to help the community reproduce, compare, and extend our work.

Join us | Research interns, full-time researchers, and AI engineers in spatial intelligence, LLM agents, reinforcement learning, world models, multimodal learning, embodied AI, recommendation, and generative AI are welcome to get in touch.


Highlights

30+ Open-source Projects | 11 ICLR 2026 Papers | 10 CVPR 2026 Papers | 5 ECCV 2026 Papers
7 ACL 2026 Papers | 5 AAAI 2026 Papers | 4 ICML 2026 Papers | 1 KDD 2026 Oral Paper
5 ICCV 2025 Papers | 2 EMNLP 2025 Oral Papers
Focus: Spatial Intelligence · DreamX · LLM Agents · World Models · Multimodal and Generative AI


What We Build

One Mission: Spatial Intelligence

We define spatial intelligence as the ability of AI to understand the real world and its evolution over space and time; predict future states; generate and simulate digital representations; plan toward human goals; and act through products and embodied systems.

For AMAP, this means connecting maps, mobility, urban environments, local services, digital content, and physical action in one learning and deployment loop. Generative intelligence is an essential part of this mission: it enables AI to represent, create, enrich, and simulate the world rather than standing as a separate product anchor.

Our work is organized around three core problems:

Core problemPrimary DreamX familiesWhat it means
Understand and Predict the WorldDreamX-Predictor · DreamX-RECConnect maps, vision, language, mobility, urban scenes, user intent, and product signals to understand spatial context and forecast how the world evolves.
Generate and Simulate the WorldDreamX-World · DreamX-CreatorCreate map-native assets, videos, 3D scenes, digital content, and persistent interactive worlds with controllability, spatial consistency, temporal coherence, and physical plausibility.
Plan and Act in the WorldDreamX-Agent · DreamX-PhiBuild agents and decision systems that reason, use tools, plan, self-reflect, and turn human goals into actions in digital and physical environments.

The DreamX Family

DreamX turns this mission into six model and system families. Each family focuses on a distinct relationship between intelligence and the real world:

FamilyRoleFocus
DreamX-PredictorPredict the worldModel the spatiotemporal evolution of traffic, mobility, demand, supply, and urban conditions.
DreamX-WorldSimulate the worldLearn dynamic world models for persistent, controllable, physically grounded, and interactive simulation.
DreamX-AgentPlan and complete digital tasksUnderstand user goals, reason over spatial context, use tools, and coordinate complex map, mobility, and local-service workflows.
DreamX-PhiAct in the physical worldConnect perception, reasoning, decision-making, and physical action for embodied intelligence, autonomous systems, and spatial devices.
DreamX-RECConnect people, places, and servicesMatch intent with locations, content, routes, and services under spatial, temporal, and contextual constraints.
DreamX-CreatorCreate and enrich digital assetsGenerate and edit map and navigation assets, local-service content, images, videos, 3D assets, and other spatial media.

This table defines the portfolio at the family level. Public same-name releases and related research artifacts are linked below where available.

One Shared Foundation

The six DreamX families build on a common foundation:

Shared capabilityRole
Spatial data and knowledgeGround models in maps, routes, places, mobility, urban environments, local services, and real-world feedback.
Multimodal foundation modelsConnect language, vision, video, maps, GUIs, sensor observations, user intent, and product signals.
Spatiotemporal and world modelingRepresent geometry, dynamics, long-horizon evolution, causality, interaction, and physical constraints.
Agents, reinforcement learning, and decision-makingTrain models to reason, use tools, recommend, plan, self-reflect, and improve through feedback.
Generative modelingCreate and edit controllable, consistent, high-quality spatial assets, media, scenes, and experiences.
Infrastructure and evaluationSupport scalable data, training, inference, deployment, benchmarks, metrics, and reproducible evaluation.

Flagship Releases

These selected public projects show how the spatial intelligence mission translates into research artifacts, open-source systems, benchmarks, and models. The complete project map below organizes all public artifacts by the three core problems and the shared foundation.

Selected Releases

ProjectContributionWhy it matters
DreamX-World 1.0A general-purpose interactive world model for controllable, long-horizon world simulation.Anchors the DreamX family with an open-source model, technical report, and live interactive system.
SkillClawAgentic skill evolution from real interaction traces.Demonstrates reusable, self-evolving agent capabilities.
FluxTextScene-text editing for controllable visual asset generation.Connects generative modeling with practical content creation.
MobilityBenchRoute-planning agent evaluation in real-world mobility scenarios.Grounds spatial reasoning in an AMAP-native benchmark.
Tree-GRPOTree-search rollouts for LLM agent reinforcement learning.Advances exploration and reasoning in agent training.
GPGMinimalist group policy gradient for model reasoning.Provides a simple and reusable reinforcement-learning foundation.

Recent Updates

  • 2026.06.18 AMAP-ML has 5 papers accepted to ECCV 2026, expanding the team's work across spatial intelligence, generative modeling, and multimodal AI.
  • 2026.06.15DreamX-World releases its 1.0 technical report and open-sources DreamX-World-5B for long-horizon interactive world generation with 1-minute video support.
  • 2026.05.18MobilityBench provides a scalable benchmark for evaluating route-planning agents in real-world mobility scenarios (KDD 2026 Oral).
  • 2026.05.12CoEvolve trains LLM agents through agent-data mutual evolution, using failure signals to synthesize harder tasks as the agent improves (ACL 2026).
  • 2026.05.12Thinking-with-Map strengthens geolocalization with a reinforced parallel map-augmented reasoning agent (ACL 2026 Findings).
  • 2026.05.11DreamX-World releases the 5B-Cam model and inference code for general-purpose interactive world simulation.
  • 2026.05.01UniMRG shows that multi-representation generation strengthens understanding in unified multimodal models, not just generation (ICML 2026).
  • 2026.05.01Train-Free Infinite-Frame Generation extends pretrained video diffusion to arbitrarily long, temporally consistent videos without any retraining (ICML 2026).
  • 2026.05.01D-Evo improves data efficiency in RL with dual difficulty-aware self-evolution that adaptively reshapes both task and sample difficulty (ICML 2026).
  • 2026.05.01EEPO introduces embedding-perturbed exploration for preference optimization in flow models, addressing exploration collapse in continuous generative policies (ICML 2026).
  • 2026.04.22DCW mitigates SNR-t bias and improves diffusion generation quality across model families (CVPR 2026).
  • 2026.04.22EMF extends efficient one-step generation from class-conditioned synthesis to text-conditioned image generation (CVPR 2026).
  • 2026.04.10SkillClaw turns real interaction traces into reusable, evolving skill libraries.
  • 2026.04.01MACE-Dance decouples motion generation and appearance synthesis for high-quality music-driven dance video (SIGGRAPH 2026).
  • 2026.03.23Omni-WorldBench evaluates world models in dynamic 4D interactive settings.
  • 2026.03.20AutoDrive-R2 improves VLA models with reasoning and self-reflection for autonomous driving scenarios (ICLR 2026).
  • 2026.03.18Video-STAR uses tool-augmented reinforcement learning for open-vocabulary action recognition in video (ICLR 2026).
  • 2026.03.11RL3DEdit uses geometry-guided reinforcement learning to make 3D scene edits more multi-view consistent (CVPR 2026).
Earlier Updates
  • 2026.03.01FE2E transfers image-editing priors into dense depth and normal estimation (CVPR 2026).
  • 2026.02.28FASA improves sparse decoding with frequency-aware attention (ICLR 2026).
  • 2026.02.27Eevee provides high-resolution data and evaluation for video-based virtual try-on (CVPR 2026 Findings).
  • 2026.02.06MobilityBench evaluates route-planning agents in real-world mobility scenarios (KDD 2026 Oral).
  • 2026.02.06SpatialGenEval benchmarks spatial intelligence in text-to-image models (ICLR 2026).
  • 2026.02.06Tree-GRPO replaces independent chain rollouts with tree-search rollouts for LLM agent reinforcement learning (ICLR 2026).
  • 2026.02.04Code2World predicts GUI transitions through renderable code generation.
  • 2026.02.04GPG provides a simple group policy gradient baseline for model reasoning (ICLR 2026).
  • 2025.10.22Taming-Hallucinations reduces MLLM video hallucinations with counterfactual video generation.
  • 2025.06.20FluxText provides a diffusion transformer baseline for scene-text editing.

Project Map

The project map complements the family-level view with the research and open-source foundations behind DreamX. Each artifact is organized by the primary role it plays in the spatial intelligence stack.

Understand and Predict the World

RepositoryContributionVenue
MobilityBenchRoute-planning agent evaluation in real-world mobility scenarios.KDD 2026 Oral
Thinking-with-MapMap-augmented geolocalization agent trained with reinforcement learning.ACL 2026 Findings
SocioReasonerVision-language reasoning for urban socio-semantic segmentation.ICLR 2026
DSFNetMulti-scenario route ranking with a public industrial driving-route dataset and AMAP deployment.WWW 2025
IntTravelReal-world dataset and generative framework for integrated multi-task travel recommendation.arXiv 2026
FE2EImage-editing priors for dense geometry estimation.CVPR 2026
UniVG-R1Reasoning-guided universal visual grounding with reinforcement learning.CVPR 2026
Taming-HallucinationsCounterfactual video generation for reducing MLLM video hallucinations.-

Generate and Simulate the World

RepositoryContributionVenue
DreamX-World 1.0General-purpose world model for interactive world simulation.-
Code2WorldGUI world model via renderable code generation.-
FluxTextDiffusion transformer baseline for scene-text editing.-
RL3DEditGeometry-guided reinforcement learning for multi-view consistent 3D scene editing.CVPR 2026
MACE-DanceMotion-appearance cascaded generation for music-driven dance video.SIGGRAPH 2026
Omni-EffectsPrompt-guided and spatially controllable composite visual effects generation.AAAI 2026
S2-GuidanceTraining-free stochastic self-guidance for diffusion models.ICLR 2026
EPGPixel-space generative modeling via self-supervised pre-training.ICLR 2026
USPUnified self-supervised pretraining in VAE space for diffusion models.ICCV 2025
EMFText-conditioned one-step image generation.CVPR 2026
DCWDifferential correction for SNR-t bias in diffusion probabilistic models.CVPR 2026
NarrLVNarrative-centric evaluation for long video generation models.ICLR 2026
ImagerySearchAdaptive test-time search for video generation.AAAI 2026
EeveeHigh-resolution benchmark for video-based virtual try-on.CVPR 2026 Findings
VMBenchPerception-aligned benchmark for video motion generation.ICCV 2025

Plan and Act in the World

RepositoryContributionVenue
SkillClawAgentic evolver for collective skill library improvement.-
AutoDrive-R2Reasoning and self-reflection for VLA models in autonomous driving.ICLR 2026
Tree-GRPOTree-search rollouts for LLM agent reinforcement learning.ICLR 2026
GPGSimple and strong group policy gradient baseline for model reasoning.ICLR 2026
CoEvolveAgent-data mutual evolution for training LLM agents.ACL 2026
MathForgeDifficulty-aware GRPO and multi-aspect reformulation for math reasoning.ICLR 2026
Video-STARTool-using reinforcement learning for open-vocabulary action recognition.ICLR 2026

Shared Foundations and Evaluation

RepositoryContributionVenue
SpatialGenEvalSpatial intelligence evaluation for text-to-image models.ICLR 2026
Omni-WorldBenchBenchmark for interactive response capabilities of world models.arXiv 2026
RealQARealistic image quality and aesthetic scoring with multimodal LLMs.-
FASAFrequency-aware sparse attention for efficient sparse decoding.ICLR 2026

For Collaborators and Applicants

We are looking for people who want to build the next generation of spatial intelligence systems through clean code, reproducible experiments, rigorous evaluation, ambitious problem selection, and real-world product impact.

If you are interested in research internships, full-time roles, or academic collaboration, please email cxxgtxy@gmail.com (homepage) with your CV, representative projects, and research interests.

Pinned Loading

  1. Tree-GRPOTree-GRPOPublic

    [ICLR 2026] Tree Search for LLM Agent Reinforcement Learning

    Python 391 39

  2. Code2WorldCode2WorldPublic

    Code2World: A GUI World Model via Renderable Code Generation

    Python 323 19

  3. GPGGPGPublic

    [ICLR26]GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning

    Python 180 5

  4. DreamX-WorldDreamX-WorldPublic

    DreamX-World: A General-Purpose Interactive World Model

    Python 746 48

  5. SkillClawSkillClawPublic

    Let Skills Evolve Collectively with Agentic Evolver

    Python 2.4k 234

  6. FluxTextFluxTextPublic

    Implementation of "FLUX-Text: A Simple and Advanced Diffusion Transformer Baseline for Scene Text Editing"

    Python 932 31

Repositories

Showing 10 of 50 repositories

Top languages

Loading…

Most used topics

Loading…