Repository files navigation

CUA Skill — Computer Use Agent with Skills

CUA Skill is a skill-based autonomous GUI agent framework for Windows desktop applications. Instead of generating every low-level action from scratch, CUA Skill retrieves and executes pre-recorded action sequences (skills) from an indexed library, enabling reliable and efficient task completion across 17+ Windows applications.


Table of Contents


Overview

CUA Skill provides two agent modes for automating desktop tasks:

ModeDescription
Replay AgentExecutes pre-defined task graphs step-by-step with vision-based grounding
RAG AgentUses an LLM planner with Retrieval-Augmented Generation (RAG) to dynamically retrieve, select, and configure skills at runtime

The core insight is that many desktop tasks share common action patterns (e.g., opening apps, clicking menus, typing text). By indexing these patterns as reusable skills, the agent can accomplish complex tasks through composition rather than trial-and-error.


Architecture

┌───────────────────────────────────────────────────────┐
│ CUA Skill Agent │
├──────────────┬───────────────┬────────────────────────┤
│ Planner │ Retriever │ Mixture Grounding │
│ (LLM-based) │ (BM25 + │ (UI-TARS / UIA Tree) │
│ │ Semantic) │ │
├──────────────┴───────────────┴────────────────────────┤
│ Action System │
│ ┌────────────┐ ┌────────────────┐ ┌─────────────┐ │
│ │ Base │ │ Common │ │ App-Specific│ │
│ │ Actions │ │ Actions │ │ Skills │ │
│ └────────────┘ └────────────────┘ └─────────────┘ │
├───────────────────────────────────────────────────────┤
│ Desktop Environment │
│ (Screenshots, A11y Tree, pyautogui) │
└───────────────────────────────────────────────────────┘

Key Components

  • Planner — LLM-driven (GPT or Qwen) component that generates search queries, selects the best action from retrieved candidates, configures action parameters, and maintains an action memory.
  • Retriever — Hybrid search engine combining BM25+ keyword ranking and semantic embeddings for skill retrieval from the indexed library.
  • Mixture Grounding — Refines click coordinates using vision grounding models (UI-TARS v1 endpoint) in a mixture-of-experts pattern.
  • Action System — 20+ base actions (click, type, scroll, hotkey, etc.) with a registry pattern, plus composable action graphs (DAGs) for multi-step skills.
  • Desktop Environment — Wraps pyautogui and pywinauto for screenshots, accessibility tree extraction, and action execution.

Supported Applications

ApplicationSkill Module
Windows Start / Search / Runcommon_action.py
Bing Searchbing_search_action.py
Google Chromechrome_actions.py
Microsoft Edgemicrosoft_edge_action.py
Microsoft Excelexcel_action.py
Microsoft Wordword_action.py
Microsoft PowerPointpowerpoint_action.py
Notepadnotepad_action.py
Paintpaint_action.py
Calculatorcalculator_action.py
Clockclock_action.py
File Explorerfile_explorer_action.py
VLC Media Playervlc_action.py
VS Codevs_code_action.py
Amazonamazon_action.py
YouTubeyoutube_action.py
Windows Settingswindows_settings_action.py

Getting Started

Prerequisites

  • Python 3.10+
  • Windows OS (agent interacts with the Windows desktop)
  • Azure OpenAI access (for LLM planner and grounding models)
  • (Optional) Local Qwen model for offline planner

Installation

  1. Clone the repository:

    git clone <repo-url>cd cua_skill
  2. Install dependencies:

    pip install -r agent/requirements.txt
  3. For Windows Agent Arena evaluation, install additional dependencies:

    pip install -r agent/requirements_waa.txt

Configuration

  1. Create a .env file in the agent/ directory:

    UITARS_V1_BEARER_KEY="your_uitars_key"
    AZURE_AD_TOKEN=""
    
  2. Configure the agent via JSON config files:

    • agent/config.json — Replay agent settings (environment, grounding, logging)
    • agent/config_rag.json — RAG agent settings (planner model, retrieval parameters, RAG index paths)

Key configuration options:

SettingFileDescription
mixture_grounding.expertisesconfig.jsonVision grounding model endpoints and weights
planner.model_classconfig_rag.jsonLLM planner backend: "gpt" or "qwen"
rag.semantic_weightconfig_rag.jsonBlending weight for hybrid search (default: 0.7)
max_stepsBothMaximum actions per task (default: 50)
max_wall_timeBothTimeout in seconds (default: 300)

Agent Modes

Replay Agent

The Replay Agent (CUAKnowledgeGraphAgent) executes pre-defined task graphs:

  1. Loads a task definition (JSON) as a directed graph of actions
  2. Pops steps one-by-one in topological order
  3. For each step, applies mixture grounding to refine coordinates
  4. Executes the action via the desktop environment
fromagentimportCUAKnowledgeGraphAgentagent=CUAKnowledgeGraphAgent(config="agent/config.json")
agent.proceed(instruction="Open Notepad and type Hello", example=task_json)

RAG Agent

The RAG Agent (CUARAGAgent) dynamically plans and executes tasks:

  1. Feasibility check — LLM evaluates if the task is achievable from the current screen
  2. Loop until completion or timeout:
    • Capture screenshot observation
    • Generate search queries from task description + action history + screenshot
    • Retrieve skills via hybrid BM25 + semantic search over the indexed skill library
    • Select the best action from retrieved candidates (with base action fallbacks)
    • Configure action parameters using the LLM + current screenshot
    • Ground coordinates via the mixture grounding model (UI-TARS)
    • Execute the action
    • Update memory — LLM summarizes the outcome for future context
fromagentimportCUARAGAgentagent=CUARAGAgent(config="agent/config_rag.json")
agent.proceed(instruction="Search for 'weather today' on Bing")

User Task Generation

The user_task_generation/ module provides a pipeline for synthesizing diverse user tasks from primitive operations and compositions:

  1. Define primitive operations per application (e.g., BingSearchLaunch, InsertImage) with argument generators
  2. Create compositions combining primitives into multi-step user tasks
  3. Generate tasks with automatic argument filling, instruction drop-off for diversity, and LLM-based rephrasing
python user_task_generation/user_task_generator.py \
--primitive-operation ./asset/primitive_operation/bingsearch_primitive_operation.json \
--composition ./asset/primitive_operation_composition/bingsearch_primitive_operation_composition.json \
--app-name bingsearch \
--out-dir ./asset/user_task \
--num-tasks 10000

See user_task_generation/README.md for full documentation.


Evaluation (Windows Agent Arena)

CUA Skill can be evaluated in the Windows Agent Arena environment:

  1. Set up the Windows Agent Arena Docker image (see evaluation/WindowsAgentArena/README.md)
  2. Place test JSON files in evaluation/WindowsAgentArena/test_jsons/
  3. Run evaluation:
    cd evaluation/WindowsAgentArena
    sudo bash ./run_cua_rag.sh <test_json_filename> [options]

Available options:

  • --use_gold_image — Use the clean backup storage image
  • --clean_mode — Reset environment between test cases (recommended)
  • --reset_image — Regenerate storage from setup ISO

Project Structure

cua_skill/
├── agent/ # Core agent framework
│ ├── agent.py # Replay agent (CUAKnowledgeGraphAgent)
│ ├── agent_rag.py # RAG agent (CUARAGAgent)
│ ├── agent_waa.py # WAA adapter for replay agent
│ ├── agent_rag_waa.py # WAA adapter for RAG agent
│ ├── planner.py # LLM-based planning (query gen, action selection, config)
│ ├── retrieval.py # Hybrid BM25 + semantic retrieval engine
│ ├── mixture_grounding.py # Vision-based coordinate grounding (UI-TARS)
│ ├── llms.py # LLM clients (Azure OpenAI GPT, local Qwen)
│ ├── desktop_env.py # Desktop environment wrapper
│ ├── replay_task.py # Task graph parser and executor
│ ├── config.json # Replay agent configuration
│ ├── config_rag.json # RAG agent configuration
│ ├── action/ # Action system
│ │ ├── base_action.py # Base actions with registry pattern
│ │ ├── compose_action.py # Composable action DAGs
│ │ ├── <app>_action.py # Application-specific skill modules
│ │ └── argument.py # Action argument definitions
│ └── utils/ # Utilities (logging, UIA, config, etc.)
├── user_task_generation/ # Task synthesis pipeline
│ ├── user_task_generator.py # Main generator script
│ ├── argument_value_generator/ # Realistic argument value generators
│ └── README.md # Task generation documentation
├── evaluation/ # Evaluation tooling
│ └── WindowsAgentArena/ # WAA integration and analysis notebooks
├── docs/ # Project documentation website
├── LICENSE # MIT License
└── README.md # This file

License

This project is licensed under the MIT License. See LICENSE for details.

Copyright (c) 2026 Microsoft Corporation.

About

No description, website, or topics provided.

Resources

Code of conduct

Security policy

Stars

48 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

CUA Skill — Computer Use Agent with Skills

CUA Skill is a skill-based autonomous GUI agent framework for Windows desktop applications. Instead of generating every low-level action from scratch, CUA Skill retrieves and executes pre-recorded action sequences (skills) from an indexed library, enabling reliable and efficient task completion across 17+ Windows applications.


Table of Contents


Overview

CUA Skill provides two agent modes for automating desktop tasks:

ModeDescription
Replay AgentExecutes pre-defined task graphs step-by-step with vision-based grounding
RAG AgentUses an LLM planner with Retrieval-Augmented Generation (RAG) to dynamically retrieve, select, and configure skills at runtime

The core insight is that many desktop tasks share common action patterns (e.g., opening apps, clicking menus, typing text). By indexing these patterns as reusable skills, the agent can accomplish complex tasks through composition rather than trial-and-error.


Architecture

┌───────────────────────────────────────────────────────┐
│ CUA Skill Agent │
├──────────────┬───────────────┬────────────────────────┤
│ Planner │ Retriever │ Mixture Grounding │
│ (LLM-based) │ (BM25 + │ (UI-TARS / UIA Tree) │
│ │ Semantic) │ │
├──────────────┴───────────────┴────────────────────────┤
│ Action System │
│ ┌────────────┐ ┌────────────────┐ ┌─────────────┐ │
│ │ Base │ │ Common │ │ App-Specific│ │
│ │ Actions │ │ Actions │ │ Skills │ │
│ └────────────┘ └────────────────┘ └─────────────┘ │
├───────────────────────────────────────────────────────┤
│ Desktop Environment │
│ (Screenshots, A11y Tree, pyautogui) │
└───────────────────────────────────────────────────────┘

Key Components

  • Planner — LLM-driven (GPT or Qwen) component that generates search queries, selects the best action from retrieved candidates, configures action parameters, and maintains an action memory.
  • Retriever — Hybrid search engine combining BM25+ keyword ranking and semantic embeddings for skill retrieval from the indexed library.
  • Mixture Grounding — Refines click coordinates using vision grounding models (UI-TARS v1 endpoint) in a mixture-of-experts pattern.
  • Action System — 20+ base actions (click, type, scroll, hotkey, etc.) with a registry pattern, plus composable action graphs (DAGs) for multi-step skills.
  • Desktop Environment — Wraps pyautogui and pywinauto for screenshots, accessibility tree extraction, and action execution.

Supported Applications

ApplicationSkill Module
Windows Start / Search / Runcommon_action.py
Bing Searchbing_search_action.py
Google Chromechrome_actions.py
Microsoft Edgemicrosoft_edge_action.py
Microsoft Excelexcel_action.py
Microsoft Wordword_action.py
Microsoft PowerPointpowerpoint_action.py
Notepadnotepad_action.py
Paintpaint_action.py
Calculatorcalculator_action.py
Clockclock_action.py
File Explorerfile_explorer_action.py
VLC Media Playervlc_action.py
VS Codevs_code_action.py
Amazonamazon_action.py
YouTubeyoutube_action.py
Windows Settingswindows_settings_action.py

Getting Started

Prerequisites

  • Python 3.10+
  • Windows OS (agent interacts with the Windows desktop)
  • Azure OpenAI access (for LLM planner and grounding models)
  • (Optional) Local Qwen model for offline planner

Installation

  1. Clone the repository:

    git clone <repo-url>cd cua_skill
  2. Install dependencies:

    pip install -r agent/requirements.txt
  3. For Windows Agent Arena evaluation, install additional dependencies:

    pip install -r agent/requirements_waa.txt

Configuration

  1. Create a .env file in the agent/ directory:

    UITARS_V1_BEARER_KEY="your_uitars_key"
    AZURE_AD_TOKEN=""
    
  2. Configure the agent via JSON config files:

    • agent/config.json — Replay agent settings (environment, grounding, logging)
    • agent/config_rag.json — RAG agent settings (planner model, retrieval parameters, RAG index paths)

Key configuration options:

SettingFileDescription
mixture_grounding.expertisesconfig.jsonVision grounding model endpoints and weights
planner.model_classconfig_rag.jsonLLM planner backend: "gpt" or "qwen"
rag.semantic_weightconfig_rag.jsonBlending weight for hybrid search (default: 0.7)
max_stepsBothMaximum actions per task (default: 50)
max_wall_timeBothTimeout in seconds (default: 300)

Agent Modes

Replay Agent

The Replay Agent (CUAKnowledgeGraphAgent) executes pre-defined task graphs:

  1. Loads a task definition (JSON) as a directed graph of actions
  2. Pops steps one-by-one in topological order
  3. For each step, applies mixture grounding to refine coordinates
  4. Executes the action via the desktop environment
fromagentimportCUAKnowledgeGraphAgentagent=CUAKnowledgeGraphAgent(config="agent/config.json")
agent.proceed(instruction="Open Notepad and type Hello", example=task_json)

RAG Agent

The RAG Agent (CUARAGAgent) dynamically plans and executes tasks:

  1. Feasibility check — LLM evaluates if the task is achievable from the current screen
  2. Loop until completion or timeout:
    • Capture screenshot observation
    • Generate search queries from task description + action history + screenshot
    • Retrieve skills via hybrid BM25 + semantic search over the indexed skill library
    • Select the best action from retrieved candidates (with base action fallbacks)
    • Configure action parameters using the LLM + current screenshot
    • Ground coordinates via the mixture grounding model (UI-TARS)
    • Execute the action
    • Update memory — LLM summarizes the outcome for future context
fromagentimportCUARAGAgentagent=CUARAGAgent(config="agent/config_rag.json")
agent.proceed(instruction="Search for 'weather today' on Bing")

User Task Generation

The user_task_generation/ module provides a pipeline for synthesizing diverse user tasks from primitive operations and compositions:

  1. Define primitive operations per application (e.g., BingSearchLaunch, InsertImage) with argument generators
  2. Create compositions combining primitives into multi-step user tasks
  3. Generate tasks with automatic argument filling, instruction drop-off for diversity, and LLM-based rephrasing
python user_task_generation/user_task_generator.py \
--primitive-operation ./asset/primitive_operation/bingsearch_primitive_operation.json \
--composition ./asset/primitive_operation_composition/bingsearch_primitive_operation_composition.json \
--app-name bingsearch \
--out-dir ./asset/user_task \
--num-tasks 10000

See user_task_generation/README.md for full documentation.


Evaluation (Windows Agent Arena)

CUA Skill can be evaluated in the Windows Agent Arena environment:

  1. Set up the Windows Agent Arena Docker image (see evaluation/WindowsAgentArena/README.md)
  2. Place test JSON files in evaluation/WindowsAgentArena/test_jsons/
  3. Run evaluation:
    cd evaluation/WindowsAgentArena
    sudo bash ./run_cua_rag.sh <test_json_filename> [options]

Available options:

  • --use_gold_image — Use the clean backup storage image
  • --clean_mode — Reset environment between test cases (recommended)
  • --reset_image — Regenerate storage from setup ISO

Project Structure

cua_skill/
├── agent/ # Core agent framework
│ ├── agent.py # Replay agent (CUAKnowledgeGraphAgent)
│ ├── agent_rag.py # RAG agent (CUARAGAgent)
│ ├── agent_waa.py # WAA adapter for replay agent
│ ├── agent_rag_waa.py # WAA adapter for RAG agent
│ ├── planner.py # LLM-based planning (query gen, action selection, config)
│ ├── retrieval.py # Hybrid BM25 + semantic retrieval engine
│ ├── mixture_grounding.py # Vision-based coordinate grounding (UI-TARS)
│ ├── llms.py # LLM clients (Azure OpenAI GPT, local Qwen)
│ ├── desktop_env.py # Desktop environment wrapper
│ ├── replay_task.py # Task graph parser and executor
│ ├── config.json # Replay agent configuration
│ ├── config_rag.json # RAG agent configuration
│ ├── action/ # Action system
│ │ ├── base_action.py # Base actions with registry pattern
│ │ ├── compose_action.py # Composable action DAGs
│ │ ├── <app>_action.py # Application-specific skill modules
│ │ └── argument.py # Action argument definitions
│ └── utils/ # Utilities (logging, UIA, config, etc.)
├── user_task_generation/ # Task synthesis pipeline
│ ├── user_task_generator.py # Main generator script
│ ├── argument_value_generator/ # Realistic argument value generators
│ └── README.md # Task generation documentation
├── evaluation/ # Evaluation tooling
│ └── WindowsAgentArena/ # WAA integration and analysis notebooks
├── docs/ # Project documentation website
├── LICENSE # MIT License
└── README.md # This file

License

This project is licensed under the MIT License. See LICENSE for details.

Copyright (c) 2026 Microsoft Corporation.

About

No description, website, or topics provided.

Resources

Code of conduct

Security policy

Stars

48 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CUA Skill — Computer Use Agent with Skills

CUA Skill is a skill-based autonomous GUI agent framework for Windows desktop applications. Instead of generating every low-level action from scratch, CUA Skill retrieves and executes pre-recorded action sequences (skills) from an indexed library, enabling reliable and efficient task completion across 17+ Windows applications.


Table of Contents


Overview

CUA Skill provides two agent modes for automating desktop tasks:

ModeDescription
Replay AgentExecutes pre-defined task graphs step-by-step with vision-based grounding
RAG AgentUses an LLM planner with Retrieval-Augmented Generation (RAG) to dynamically retrieve, select, and configure skills at runtime

The core insight is that many desktop tasks share common action patterns (e.g., opening apps, clicking menus, typing text). By indexing these patterns as reusable skills, the agent can accomplish complex tasks through composition rather than trial-and-error.


Architecture

┌───────────────────────────────────────────────────────┐
│ CUA Skill Agent │
├──────────────┬───────────────┬────────────────────────┤
│ Planner │ Retriever │ Mixture Grounding │
│ (LLM-based) │ (BM25 + │ (UI-TARS / UIA Tree) │
│ │ Semantic) │ │
├──────────────┴───────────────┴────────────────────────┤
│ Action System │
│ ┌────────────┐ ┌────────────────┐ ┌─────────────┐ │
│ │ Base │ │ Common │ │ App-Specific│ │
│ │ Actions │ │ Actions │ │ Skills │ │
│ └────────────┘ └────────────────┘ └─────────────┘ │
├───────────────────────────────────────────────────────┤
│ Desktop Environment │
│ (Screenshots, A11y Tree, pyautogui) │
└───────────────────────────────────────────────────────┘

Key Components

  • Planner — LLM-driven (GPT or Qwen) component that generates search queries, selects the best action from retrieved candidates, configures action parameters, and maintains an action memory.
  • Retriever — Hybrid search engine combining BM25+ keyword ranking and semantic embeddings for skill retrieval from the indexed library.
  • Mixture Grounding — Refines click coordinates using vision grounding models (UI-TARS v1 endpoint) in a mixture-of-experts pattern.
  • Action System — 20+ base actions (click, type, scroll, hotkey, etc.) with a registry pattern, plus composable action graphs (DAGs) for multi-step skills.
  • Desktop Environment — Wraps pyautogui and pywinauto for screenshots, accessibility tree extraction, and action execution.

Supported Applications

ApplicationSkill Module
Windows Start / Search / Runcommon_action.py
Bing Searchbing_search_action.py
Google Chromechrome_actions.py
Microsoft Edgemicrosoft_edge_action.py
Microsoft Excelexcel_action.py
Microsoft Wordword_action.py
Microsoft PowerPointpowerpoint_action.py
Notepadnotepad_action.py
Paintpaint_action.py
Calculatorcalculator_action.py
Clockclock_action.py
File Explorerfile_explorer_action.py
VLC Media Playervlc_action.py
VS Codevs_code_action.py
Amazonamazon_action.py
YouTubeyoutube_action.py
Windows Settingswindows_settings_action.py

Getting Started

Prerequisites

  • Python 3.10+
  • Windows OS (agent interacts with the Windows desktop)
  • Azure OpenAI access (for LLM planner and grounding models)
  • (Optional) Local Qwen model for offline planner

Installation

  1. Clone the repository:

    git clone <repo-url>cd cua_skill
  2. Install dependencies:

    pip install -r agent/requirements.txt
  3. For Windows Agent Arena evaluation, install additional dependencies:

    pip install -r agent/requirements_waa.txt

Configuration

  1. Create a .env file in the agent/ directory:

    UITARS_V1_BEARER_KEY="your_uitars_key"
    AZURE_AD_TOKEN=""
    
  2. Configure the agent via JSON config files:

    • agent/config.json — Replay agent settings (environment, grounding, logging)
    • agent/config_rag.json — RAG agent settings (planner model, retrieval parameters, RAG index paths)

Key configuration options:

SettingFileDescription
mixture_grounding.expertisesconfig.jsonVision grounding model endpoints and weights
planner.model_classconfig_rag.jsonLLM planner backend: "gpt" or "qwen"
rag.semantic_weightconfig_rag.jsonBlending weight for hybrid search (default: 0.7)
max_stepsBothMaximum actions per task (default: 50)
max_wall_timeBothTimeout in seconds (default: 300)

Agent Modes

Replay Agent

The Replay Agent (CUAKnowledgeGraphAgent) executes pre-defined task graphs:

  1. Loads a task definition (JSON) as a directed graph of actions
  2. Pops steps one-by-one in topological order
  3. For each step, applies mixture grounding to refine coordinates
  4. Executes the action via the desktop environment
fromagentimportCUAKnowledgeGraphAgentagent=CUAKnowledgeGraphAgent(config="agent/config.json")
agent.proceed(instruction="Open Notepad and type Hello", example=task_json)

RAG Agent

The RAG Agent (CUARAGAgent) dynamically plans and executes tasks:

  1. Feasibility check — LLM evaluates if the task is achievable from the current screen
  2. Loop until completion or timeout:
    • Capture screenshot observation
    • Generate search queries from task description + action history + screenshot
    • Retrieve skills via hybrid BM25 + semantic search over the indexed skill library
    • Select the best action from retrieved candidates (with base action fallbacks)
    • Configure action parameters using the LLM + current screenshot
    • Ground coordinates via the mixture grounding model (UI-TARS)
    • Execute the action
    • Update memory — LLM summarizes the outcome for future context
fromagentimportCUARAGAgentagent=CUARAGAgent(config="agent/config_rag.json")
agent.proceed(instruction="Search for 'weather today' on Bing")

User Task Generation

The user_task_generation/ module provides a pipeline for synthesizing diverse user tasks from primitive operations and compositions:

  1. Define primitive operations per application (e.g., BingSearchLaunch, InsertImage) with argument generators
  2. Create compositions combining primitives into multi-step user tasks
  3. Generate tasks with automatic argument filling, instruction drop-off for diversity, and LLM-based rephrasing
python user_task_generation/user_task_generator.py \
--primitive-operation ./asset/primitive_operation/bingsearch_primitive_operation.json \
--composition ./asset/primitive_operation_composition/bingsearch_primitive_operation_composition.json \
--app-name bingsearch \
--out-dir ./asset/user_task \
--num-tasks 10000

See user_task_generation/README.md for full documentation.


Evaluation (Windows Agent Arena)

CUA Skill can be evaluated in the Windows Agent Arena environment:

  1. Set up the Windows Agent Arena Docker image (see evaluation/WindowsAgentArena/README.md)
  2. Place test JSON files in evaluation/WindowsAgentArena/test_jsons/
  3. Run evaluation:
    cd evaluation/WindowsAgentArena
    sudo bash ./run_cua_rag.sh <test_json_filename> [options]

Available options:

  • --use_gold_image — Use the clean backup storage image
  • --clean_mode — Reset environment between test cases (recommended)
  • --reset_image — Regenerate storage from setup ISO

Project Structure

cua_skill/
├── agent/ # Core agent framework
│ ├── agent.py # Replay agent (CUAKnowledgeGraphAgent)
│ ├── agent_rag.py # RAG agent (CUARAGAgent)
│ ├── agent_waa.py # WAA adapter for replay agent
│ ├── agent_rag_waa.py # WAA adapter for RAG agent
│ ├── planner.py # LLM-based planning (query gen, action selection, config)
│ ├── retrieval.py # Hybrid BM25 + semantic retrieval engine
│ ├── mixture_grounding.py # Vision-based coordinate grounding (UI-TARS)
│ ├── llms.py # LLM clients (Azure OpenAI GPT, local Qwen)
│ ├── desktop_env.py # Desktop environment wrapper
│ ├── replay_task.py # Task graph parser and executor
│ ├── config.json # Replay agent configuration
│ ├── config_rag.json # RAG agent configuration
│ ├── action/ # Action system
│ │ ├── base_action.py # Base actions with registry pattern
│ │ ├── compose_action.py # Composable action DAGs
│ │ ├── <app>_action.py # Application-specific skill modules
│ │ └── argument.py # Action argument definitions
│ └── utils/ # Utilities (logging, UIA, config, etc.)
├── user_task_generation/ # Task synthesis pipeline
│ ├── user_task_generator.py # Main generator script
│ ├── argument_value_generator/ # Realistic argument value generators
│ └── README.md # Task generation documentation
├── evaluation/ # Evaluation tooling
│ └── WindowsAgentArena/ # WAA integration and analysis notebooks
├── docs/ # Project documentation website
├── LICENSE # MIT License
└── README.md # This file

License

This project is licensed under the MIT License. See LICENSE for details.

Copyright (c) 2026 Microsoft Corporation.

About

No description, website, or topics provided.

Resources

Code of conduct

Security policy

Stars

48 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CUA Skill — Computer Use Agent with Skills

CUA Skill is a skill-based autonomous GUI agent framework for Windows desktop applications. Instead of generating every low-level action from scratch, CUA Skill retrieves and executes pre-recorded action sequences (skills) from an indexed library, enabling reliable and efficient task completion across 17+ Windows applications.


Table of Contents


Overview

CUA Skill provides two agent modes for automating desktop tasks:

ModeDescription
Replay AgentExecutes pre-defined task graphs step-by-step with vision-based grounding
RAG AgentUses an LLM planner with Retrieval-Augmented Generation (RAG) to dynamically retrieve, select, and configure skills at runtime

The core insight is that many desktop tasks share common action patterns (e.g., opening apps, clicking menus, typing text). By indexing these patterns as reusable skills, the agent can accomplish complex tasks through composition rather than trial-and-error.


Architecture

┌───────────────────────────────────────────────────────┐
│ CUA Skill Agent │
├──────────────┬───────────────┬────────────────────────┤
│ Planner │ Retriever │ Mixture Grounding │
│ (LLM-based) │ (BM25 + │ (UI-TARS / UIA Tree) │
│ │ Semantic) │ │
├──────────────┴───────────────┴────────────────────────┤
│ Action System │
│ ┌────────────┐ ┌────────────────┐ ┌─────────────┐ │
│ │ Base │ │ Common │ │ App-Specific│ │
│ │ Actions │ │ Actions │ │ Skills │ │
│ └────────────┘ └────────────────┘ └─────────────┘ │
├───────────────────────────────────────────────────────┤
│ Desktop Environment │
│ (Screenshots, A11y Tree, pyautogui) │
└───────────────────────────────────────────────────────┘

Key Components

  • Planner — LLM-driven (GPT or Qwen) component that generates search queries, selects the best action from retrieved candidates, configures action parameters, and maintains an action memory.
  • Retriever — Hybrid search engine combining BM25+ keyword ranking and semantic embeddings for skill retrieval from the indexed library.
  • Mixture Grounding — Refines click coordinates using vision grounding models (UI-TARS v1 endpoint) in a mixture-of-experts pattern.
  • Action System — 20+ base actions (click, type, scroll, hotkey, etc.) with a registry pattern, plus composable action graphs (DAGs) for multi-step skills.
  • Desktop Environment — Wraps pyautogui and pywinauto for screenshots, accessibility tree extraction, and action execution.

Supported Applications

ApplicationSkill Module
Windows Start / Search / Runcommon_action.py
Bing Searchbing_search_action.py
Google Chromechrome_actions.py
Microsoft Edgemicrosoft_edge_action.py
Microsoft Excelexcel_action.py
Microsoft Wordword_action.py
Microsoft PowerPointpowerpoint_action.py
Notepadnotepad_action.py
Paintpaint_action.py
Calculatorcalculator_action.py
Clockclock_action.py
File Explorerfile_explorer_action.py
VLC Media Playervlc_action.py
VS Codevs_code_action.py
Amazonamazon_action.py
YouTubeyoutube_action.py
Windows Settingswindows_settings_action.py

Getting Started

Prerequisites

  • Python 3.10+
  • Windows OS (agent interacts with the Windows desktop)
  • Azure OpenAI access (for LLM planner and grounding models)
  • (Optional) Local Qwen model for offline planner

Installation

  1. Clone the repository:

    git clone <repo-url>cd cua_skill
  2. Install dependencies:

    pip install -r agent/requirements.txt
  3. For Windows Agent Arena evaluation, install additional dependencies:

    pip install -r agent/requirements_waa.txt

Configuration

  1. Create a .env file in the agent/ directory:

    UITARS_V1_BEARER_KEY="your_uitars_key"
    AZURE_AD_TOKEN=""
    
  2. Configure the agent via JSON config files:

    • agent/config.json — Replay agent settings (environment, grounding, logging)
    • agent/config_rag.json — RAG agent settings (planner model, retrieval parameters, RAG index paths)

Key configuration options:

SettingFileDescription
mixture_grounding.expertisesconfig.jsonVision grounding model endpoints and weights
planner.model_classconfig_rag.jsonLLM planner backend: "gpt" or "qwen"
rag.semantic_weightconfig_rag.jsonBlending weight for hybrid search (default: 0.7)
max_stepsBothMaximum actions per task (default: 50)
max_wall_timeBothTimeout in seconds (default: 300)

Agent Modes

Replay Agent

The Replay Agent (CUAKnowledgeGraphAgent) executes pre-defined task graphs:

  1. Loads a task definition (JSON) as a directed graph of actions
  2. Pops steps one-by-one in topological order
  3. For each step, applies mixture grounding to refine coordinates
  4. Executes the action via the desktop environment
fromagentimportCUAKnowledgeGraphAgentagent=CUAKnowledgeGraphAgent(config="agent/config.json")
agent.proceed(instruction="Open Notepad and type Hello", example=task_json)

RAG Agent

The RAG Agent (CUARAGAgent) dynamically plans and executes tasks:

  1. Feasibility check — LLM evaluates if the task is achievable from the current screen
  2. Loop until completion or timeout:
    • Capture screenshot observation
    • Generate search queries from task description + action history + screenshot
    • Retrieve skills via hybrid BM25 + semantic search over the indexed skill library
    • Select the best action from retrieved candidates (with base action fallbacks)
    • Configure action parameters using the LLM + current screenshot
    • Ground coordinates via the mixture grounding model (UI-TARS)
    • Execute the action
    • Update memory — LLM summarizes the outcome for future context
fromagentimportCUARAGAgentagent=CUARAGAgent(config="agent/config_rag.json")
agent.proceed(instruction="Search for 'weather today' on Bing")

User Task Generation

The user_task_generation/ module provides a pipeline for synthesizing diverse user tasks from primitive operations and compositions:

  1. Define primitive operations per application (e.g., BingSearchLaunch, InsertImage) with argument generators
  2. Create compositions combining primitives into multi-step user tasks
  3. Generate tasks with automatic argument filling, instruction drop-off for diversity, and LLM-based rephrasing
python user_task_generation/user_task_generator.py \
--primitive-operation ./asset/primitive_operation/bingsearch_primitive_operation.json \
--composition ./asset/primitive_operation_composition/bingsearch_primitive_operation_composition.json \
--app-name bingsearch \
--out-dir ./asset/user_task \
--num-tasks 10000

See user_task_generation/README.md for full documentation.


Evaluation (Windows Agent Arena)

CUA Skill can be evaluated in the Windows Agent Arena environment:

  1. Set up the Windows Agent Arena Docker image (see evaluation/WindowsAgentArena/README.md)
  2. Place test JSON files in evaluation/WindowsAgentArena/test_jsons/
  3. Run evaluation:
    cd evaluation/WindowsAgentArena
    sudo bash ./run_cua_rag.sh <test_json_filename> [options]

Available options:

  • --use_gold_image — Use the clean backup storage image
  • --clean_mode — Reset environment between test cases (recommended)
  • --reset_image — Regenerate storage from setup ISO

Project Structure

cua_skill/
├── agent/ # Core agent framework
│ ├── agent.py # Replay agent (CUAKnowledgeGraphAgent)
│ ├── agent_rag.py # RAG agent (CUARAGAgent)
│ ├── agent_waa.py # WAA adapter for replay agent
│ ├── agent_rag_waa.py # WAA adapter for RAG agent
│ ├── planner.py # LLM-based planning (query gen, action selection, config)
│ ├── retrieval.py # Hybrid BM25 + semantic retrieval engine
│ ├── mixture_grounding.py # Vision-based coordinate grounding (UI-TARS)
│ ├── llms.py # LLM clients (Azure OpenAI GPT, local Qwen)
│ ├── desktop_env.py # Desktop environment wrapper
│ ├── replay_task.py # Task graph parser and executor
│ ├── config.json # Replay agent configuration
│ ├── config_rag.json # RAG agent configuration
│ ├── action/ # Action system
│ │ ├── base_action.py # Base actions with registry pattern
│ │ ├── compose_action.py # Composable action DAGs
│ │ ├── <app>_action.py # Application-specific skill modules
│ │ └── argument.py # Action argument definitions
│ └── utils/ # Utilities (logging, UIA, config, etc.)
├── user_task_generation/ # Task synthesis pipeline
│ ├── user_task_generator.py # Main generator script
│ ├── argument_value_generator/ # Realistic argument value generators
│ └── README.md # Task generation documentation
├── evaluation/ # Evaluation tooling
│ └── WindowsAgentArena/ # WAA integration and analysis notebooks
├── docs/ # Project documentation website
├── LICENSE # MIT License
└── README.md # This file

License

This project is licensed under the MIT License. See LICENSE for details.

Copyright (c) 2026 Microsoft Corporation.

About

No description, website, or topics provided.

Resources

Code of conduct

Security policy

Stars

48 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

CUA Skill — Computer Use Agent with Skills

CUA Skill is a skill-based autonomous GUI agent framework for Windows desktop applications. Instead of generating every low-level action from scratch, CUA Skill retrieves and executes pre-recorded action sequences (skills) from an indexed library, enabling reliable and efficient task completion across 17+ Windows applications.


Table of Contents


Overview

CUA Skill provides two agent modes for automating desktop tasks:

ModeDescription
Replay AgentExecutes pre-defined task graphs step-by-step with vision-based grounding
RAG AgentUses an LLM planner with Retrieval-Augmented Generation (RAG) to dynamically retrieve, select, and configure skills at runtime

The core insight is that many desktop tasks share common action patterns (e.g., opening apps, clicking menus, typing text). By indexing these patterns as reusable skills, the agent can accomplish complex tasks through composition rather than trial-and-error.


Architecture

┌───────────────────────────────────────────────────────┐
│ CUA Skill Agent │
├──────────────┬───────────────┬────────────────────────┤
│ Planner │ Retriever │ Mixture Grounding │
│ (LLM-based) │ (BM25 + │ (UI-TARS / UIA Tree) │
│ │ Semantic) │ │
├──────────────┴───────────────┴────────────────────────┤
│ Action System │
│ ┌────────────┐ ┌────────────────┐ ┌─────────────┐ │
│ │ Base │ │ Common │ │ App-Specific│ │
│ │ Actions │ │ Actions │ │ Skills │ │
│ └────────────┘ └────────────────┘ └─────────────┘ │
├───────────────────────────────────────────────────────┤
│ Desktop Environment │
│ (Screenshots, A11y Tree, pyautogui) │
└───────────────────────────────────────────────────────┘

Key Components

  • Planner — LLM-driven (GPT or Qwen) component that generates search queries, selects the best action from retrieved candidates, configures action parameters, and maintains an action memory.
  • Retriever — Hybrid search engine combining BM25+ keyword ranking and semantic embeddings for skill retrieval from the indexed library.
  • Mixture Grounding — Refines click coordinates using vision grounding models (UI-TARS v1 endpoint) in a mixture-of-experts pattern.
  • Action System — 20+ base actions (click, type, scroll, hotkey, etc.) with a registry pattern, plus composable action graphs (DAGs) for multi-step skills.
  • Desktop Environment — Wraps pyautogui and pywinauto for screenshots, accessibility tree extraction, and action execution.

Supported Applications

ApplicationSkill Module
Windows Start / Search / Runcommon_action.py
Bing Searchbing_search_action.py
Google Chromechrome_actions.py
Microsoft Edgemicrosoft_edge_action.py
Microsoft Excelexcel_action.py
Microsoft Wordword_action.py
Microsoft PowerPointpowerpoint_action.py
Notepadnotepad_action.py
Paintpaint_action.py
Calculatorcalculator_action.py
Clockclock_action.py
File Explorerfile_explorer_action.py
VLC Media Playervlc_action.py
VS Codevs_code_action.py
Amazonamazon_action.py
YouTubeyoutube_action.py
Windows Settingswindows_settings_action.py

Getting Started

Prerequisites

  • Python 3.10+
  • Windows OS (agent interacts with the Windows desktop)
  • Azure OpenAI access (for LLM planner and grounding models)
  • (Optional) Local Qwen model for offline planner

Installation

  1. Clone the repository:

    git clone <repo-url>cd cua_skill
  2. Install dependencies:

    pip install -r agent/requirements.txt
  3. For Windows Agent Arena evaluation, install additional dependencies:

    pip install -r agent/requirements_waa.txt

Configuration

  1. Create a .env file in the agent/ directory:

    UITARS_V1_BEARER_KEY="your_uitars_key"
    AZURE_AD_TOKEN=""
    
  2. Configure the agent via JSON config files:

    • agent/config.json — Replay agent settings (environment, grounding, logging)
    • agent/config_rag.json — RAG agent settings (planner model, retrieval parameters, RAG index paths)

Key configuration options:

SettingFileDescription
mixture_grounding.expertisesconfig.jsonVision grounding model endpoints and weights
planner.model_classconfig_rag.jsonLLM planner backend: "gpt" or "qwen"
rag.semantic_weightconfig_rag.jsonBlending weight for hybrid search (default: 0.7)
max_stepsBothMaximum actions per task (default: 50)
max_wall_timeBothTimeout in seconds (default: 300)

Agent Modes

Replay Agent

The Replay Agent (CUAKnowledgeGraphAgent) executes pre-defined task graphs:

  1. Loads a task definition (JSON) as a directed graph of actions
  2. Pops steps one-by-one in topological order
  3. For each step, applies mixture grounding to refine coordinates
  4. Executes the action via the desktop environment
fromagentimportCUAKnowledgeGraphAgentagent=CUAKnowledgeGraphAgent(config="agent/config.json")
agent.proceed(instruction="Open Notepad and type Hello", example=task_json)

RAG Agent

The RAG Agent (CUARAGAgent) dynamically plans and executes tasks:

  1. Feasibility check — LLM evaluates if the task is achievable from the current screen
  2. Loop until completion or timeout:
    • Capture screenshot observation
    • Generate search queries from task description + action history + screenshot
    • Retrieve skills via hybrid BM25 + semantic search over the indexed skill library
    • Select the best action from retrieved candidates (with base action fallbacks)
    • Configure action parameters using the LLM + current screenshot
    • Ground coordinates via the mixture grounding model (UI-TARS)
    • Execute the action
    • Update memory — LLM summarizes the outcome for future context
fromagentimportCUARAGAgentagent=CUARAGAgent(config="agent/config_rag.json")
agent.proceed(instruction="Search for 'weather today' on Bing")

User Task Generation

The user_task_generation/ module provides a pipeline for synthesizing diverse user tasks from primitive operations and compositions:

  1. Define primitive operations per application (e.g., BingSearchLaunch, InsertImage) with argument generators
  2. Create compositions combining primitives into multi-step user tasks
  3. Generate tasks with automatic argument filling, instruction drop-off for diversity, and LLM-based rephrasing
python user_task_generation/user_task_generator.py \
--primitive-operation ./asset/primitive_operation/bingsearch_primitive_operation.json \
--composition ./asset/primitive_operation_composition/bingsearch_primitive_operation_composition.json \
--app-name bingsearch \
--out-dir ./asset/user_task \
--num-tasks 10000

See user_task_generation/README.md for full documentation.


Evaluation (Windows Agent Arena)

CUA Skill can be evaluated in the Windows Agent Arena environment:

  1. Set up the Windows Agent Arena Docker image (see evaluation/WindowsAgentArena/README.md)
  2. Place test JSON files in evaluation/WindowsAgentArena/test_jsons/
  3. Run evaluation:
    cd evaluation/WindowsAgentArena
    sudo bash ./run_cua_rag.sh <test_json_filename> [options]

Available options:

  • --use_gold_image — Use the clean backup storage image
  • --clean_mode — Reset environment between test cases (recommended)
  • --reset_image — Regenerate storage from setup ISO

Project Structure

cua_skill/
├── agent/ # Core agent framework
│ ├── agent.py # Replay agent (CUAKnowledgeGraphAgent)
│ ├── agent_rag.py # RAG agent (CUARAGAgent)
│ ├── agent_waa.py # WAA adapter for replay agent
│ ├── agent_rag_waa.py # WAA adapter for RAG agent
│ ├── planner.py # LLM-based planning (query gen, action selection, config)
│ ├── retrieval.py # Hybrid BM25 + semantic retrieval engine
│ ├── mixture_grounding.py # Vision-based coordinate grounding (UI-TARS)
│ ├── llms.py # LLM clients (Azure OpenAI GPT, local Qwen)
│ ├── desktop_env.py # Desktop environment wrapper
│ ├── replay_task.py # Task graph parser and executor
│ ├── config.json # Replay agent configuration
│ ├── config_rag.json # RAG agent configuration
│ ├── action/ # Action system
│ │ ├── base_action.py # Base actions with registry pattern
│ │ ├── compose_action.py # Composable action DAGs
│ │ ├── <app>_action.py # Application-specific skill modules
│ │ └── argument.py # Action argument definitions
│ └── utils/ # Utilities (logging, UIA, config, etc.)
├── user_task_generation/ # Task synthesis pipeline
│ ├── user_task_generator.py # Main generator script
│ ├── argument_value_generator/ # Realistic argument value generators
│ └── README.md # Task generation documentation
├── evaluation/ # Evaluation tooling
│ └── WindowsAgentArena/ # WAA integration and analysis notebooks
├── docs/ # Project documentation website
├── LICENSE # MIT License
└── README.md # This file

License

This project is licensed under the MIT License. See LICENSE for details.

Copyright (c) 2026 Microsoft Corporation.

About

No description, website, or topics provided.

Resources

Code of conduct

Security policy

Stars

48 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CUA Skill — Computer Use Agent with Skills

CUA Skill is a skill-based autonomous GUI agent framework for Windows desktop applications. Instead of generating every low-level action from scratch, CUA Skill retrieves and executes pre-recorded action sequences (skills) from an indexed library, enabling reliable and efficient task completion across 17+ Windows applications.


Table of Contents


Overview

CUA Skill provides two agent modes for automating desktop tasks:

ModeDescription
Replay AgentExecutes pre-defined task graphs step-by-step with vision-based grounding
RAG AgentUses an LLM planner with Retrieval-Augmented Generation (RAG) to dynamically retrieve, select, and configure skills at runtime

The core insight is that many desktop tasks share common action patterns (e.g., opening apps, clicking menus, typing text). By indexing these patterns as reusable skills, the agent can accomplish complex tasks through composition rather than trial-and-error.


Architecture

┌───────────────────────────────────────────────────────┐
│ CUA Skill Agent │
├──────────────┬───────────────┬────────────────────────┤
│ Planner │ Retriever │ Mixture Grounding │
│ (LLM-based) │ (BM25 + │ (UI-TARS / UIA Tree) │
│ │ Semantic) │ │
├──────────────┴───────────────┴────────────────────────┤
│ Action System │
│ ┌────────────┐ ┌────────────────┐ ┌─────────────┐ │
│ │ Base │ │ Common │ │ App-Specific│ │
│ │ Actions │ │ Actions │ │ Skills │ │
│ └────────────┘ └────────────────┘ └─────────────┘ │
├───────────────────────────────────────────────────────┤
│ Desktop Environment │
│ (Screenshots, A11y Tree, pyautogui) │
└───────────────────────────────────────────────────────┘

Key Components

  • Planner — LLM-driven (GPT or Qwen) component that generates search queries, selects the best action from retrieved candidates, configures action parameters, and maintains an action memory.
  • Retriever — Hybrid search engine combining BM25+ keyword ranking and semantic embeddings for skill retrieval from the indexed library.
  • Mixture Grounding — Refines click coordinates using vision grounding models (UI-TARS v1 endpoint) in a mixture-of-experts pattern.
  • Action System — 20+ base actions (click, type, scroll, hotkey, etc.) with a registry pattern, plus composable action graphs (DAGs) for multi-step skills.
  • Desktop Environment — Wraps pyautogui and pywinauto for screenshots, accessibility tree extraction, and action execution.

Supported Applications

ApplicationSkill Module
Windows Start / Search / Runcommon_action.py
Bing Searchbing_search_action.py
Google Chromechrome_actions.py
Microsoft Edgemicrosoft_edge_action.py
Microsoft Excelexcel_action.py
Microsoft Wordword_action.py
Microsoft PowerPointpowerpoint_action.py
Notepadnotepad_action.py
Paintpaint_action.py
Calculatorcalculator_action.py
Clockclock_action.py
File Explorerfile_explorer_action.py
VLC Media Playervlc_action.py
VS Codevs_code_action.py
Amazonamazon_action.py
YouTubeyoutube_action.py
Windows Settingswindows_settings_action.py

Getting Started

Prerequisites

  • Python 3.10+
  • Windows OS (agent interacts with the Windows desktop)
  • Azure OpenAI access (for LLM planner and grounding models)
  • (Optional) Local Qwen model for offline planner

Installation

  1. Clone the repository:

    git clone <repo-url>cd cua_skill
  2. Install dependencies:

    pip install -r agent/requirements.txt
  3. For Windows Agent Arena evaluation, install additional dependencies:

    pip install -r agent/requirements_waa.txt

Configuration

  1. Create a .env file in the agent/ directory:

    UITARS_V1_BEARER_KEY="your_uitars_key"
    AZURE_AD_TOKEN=""
    
  2. Configure the agent via JSON config files:

    • agent/config.json — Replay agent settings (environment, grounding, logging)
    • agent/config_rag.json — RAG agent settings (planner model, retrieval parameters, RAG index paths)

Key configuration options:

SettingFileDescription
mixture_grounding.expertisesconfig.jsonVision grounding model endpoints and weights
planner.model_classconfig_rag.jsonLLM planner backend: "gpt" or "qwen"
rag.semantic_weightconfig_rag.jsonBlending weight for hybrid search (default: 0.7)
max_stepsBothMaximum actions per task (default: 50)
max_wall_timeBothTimeout in seconds (default: 300)

Agent Modes

Replay Agent

The Replay Agent (CUAKnowledgeGraphAgent) executes pre-defined task graphs:

  1. Loads a task definition (JSON) as a directed graph of actions
  2. Pops steps one-by-one in topological order
  3. For each step, applies mixture grounding to refine coordinates
  4. Executes the action via the desktop environment
fromagentimportCUAKnowledgeGraphAgentagent=CUAKnowledgeGraphAgent(config="agent/config.json")
agent.proceed(instruction="Open Notepad and type Hello", example=task_json)

RAG Agent

The RAG Agent (CUARAGAgent) dynamically plans and executes tasks:

  1. Feasibility check — LLM evaluates if the task is achievable from the current screen
  2. Loop until completion or timeout:
    • Capture screenshot observation
    • Generate search queries from task description + action history + screenshot
    • Retrieve skills via hybrid BM25 + semantic search over the indexed skill library
    • Select the best action from retrieved candidates (with base action fallbacks)
    • Configure action parameters using the LLM + current screenshot
    • Ground coordinates via the mixture grounding model (UI-TARS)
    • Execute the action
    • Update memory — LLM summarizes the outcome for future context
fromagentimportCUARAGAgentagent=CUARAGAgent(config="agent/config_rag.json")
agent.proceed(instruction="Search for 'weather today' on Bing")

User Task Generation

The user_task_generation/ module provides a pipeline for synthesizing diverse user tasks from primitive operations and compositions:

  1. Define primitive operations per application (e.g., BingSearchLaunch, InsertImage) with argument generators
  2. Create compositions combining primitives into multi-step user tasks
  3. Generate tasks with automatic argument filling, instruction drop-off for diversity, and LLM-based rephrasing
python user_task_generation/user_task_generator.py \
--primitive-operation ./asset/primitive_operation/bingsearch_primitive_operation.json \
--composition ./asset/primitive_operation_composition/bingsearch_primitive_operation_composition.json \
--app-name bingsearch \
--out-dir ./asset/user_task \
--num-tasks 10000

See user_task_generation/README.md for full documentation.


Evaluation (Windows Agent Arena)

CUA Skill can be evaluated in the Windows Agent Arena environment:

  1. Set up the Windows Agent Arena Docker image (see evaluation/WindowsAgentArena/README.md)
  2. Place test JSON files in evaluation/WindowsAgentArena/test_jsons/
  3. Run evaluation:
    cd evaluation/WindowsAgentArena
    sudo bash ./run_cua_rag.sh <test_json_filename> [options]

Available options:

  • --use_gold_image — Use the clean backup storage image
  • --clean_mode — Reset environment between test cases (recommended)
  • --reset_image — Regenerate storage from setup ISO

Project Structure

cua_skill/
├── agent/ # Core agent framework
│ ├── agent.py # Replay agent (CUAKnowledgeGraphAgent)
│ ├── agent_rag.py # RAG agent (CUARAGAgent)
│ ├── agent_waa.py # WAA adapter for replay agent
│ ├── agent_rag_waa.py # WAA adapter for RAG agent
│ ├── planner.py # LLM-based planning (query gen, action selection, config)
│ ├── retrieval.py # Hybrid BM25 + semantic retrieval engine
│ ├── mixture_grounding.py # Vision-based coordinate grounding (UI-TARS)
│ ├── llms.py # LLM clients (Azure OpenAI GPT, local Qwen)
│ ├── desktop_env.py # Desktop environment wrapper
│ ├── replay_task.py # Task graph parser and executor
│ ├── config.json # Replay agent configuration
│ ├── config_rag.json # RAG agent configuration
│ ├── action/ # Action system
│ │ ├── base_action.py # Base actions with registry pattern
│ │ ├── compose_action.py # Composable action DAGs
│ │ ├── <app>_action.py # Application-specific skill modules
│ │ └── argument.py # Action argument definitions
│ └── utils/ # Utilities (logging, UIA, config, etc.)
├── user_task_generation/ # Task synthesis pipeline
│ ├── user_task_generator.py # Main generator script
│ ├── argument_value_generator/ # Realistic argument value generators
│ └── README.md # Task generation documentation
├── evaluation/ # Evaluation tooling
│ └── WindowsAgentArena/ # WAA integration and analysis notebooks
├── docs/ # Project documentation website
├── LICENSE # MIT License
└── README.md # This file

License

This project is licensed under the MIT License. See LICENSE for details.

Copyright (c) 2026 Microsoft Corporation.

About

No description, website, or topics provided.

Resources

Code of conduct

Security policy

Stars

48 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CUA Skill — Computer Use Agent with Skills

CUA Skill is a skill-based autonomous GUI agent framework for Windows desktop applications. Instead of generating every low-level action from scratch, CUA Skill retrieves and executes pre-recorded action sequences (skills) from an indexed library, enabling reliable and efficient task completion across 17+ Windows applications.


Table of Contents


Overview

CUA Skill provides two agent modes for automating desktop tasks:

ModeDescription
Replay AgentExecutes pre-defined task graphs step-by-step with vision-based grounding
RAG AgentUses an LLM planner with Retrieval-Augmented Generation (RAG) to dynamically retrieve, select, and configure skills at runtime

The core insight is that many desktop tasks share common action patterns (e.g., opening apps, clicking menus, typing text). By indexing these patterns as reusable skills, the agent can accomplish complex tasks through composition rather than trial-and-error.


Architecture

┌───────────────────────────────────────────────────────┐
│ CUA Skill Agent │
├──────────────┬───────────────┬────────────────────────┤
│ Planner │ Retriever │ Mixture Grounding │
│ (LLM-based) │ (BM25 + │ (UI-TARS / UIA Tree) │
│ │ Semantic) │ │
├──────────────┴───────────────┴────────────────────────┤
│ Action System │
│ ┌────────────┐ ┌────────────────┐ ┌─────────────┐ │
│ │ Base │ │ Common │ │ App-Specific│ │
│ │ Actions │ │ Actions │ │ Skills │ │
│ └────────────┘ └────────────────┘ └─────────────┘ │
├───────────────────────────────────────────────────────┤
│ Desktop Environment │
│ (Screenshots, A11y Tree, pyautogui) │
└───────────────────────────────────────────────────────┘

Key Components

  • Planner — LLM-driven (GPT or Qwen) component that generates search queries, selects the best action from retrieved candidates, configures action parameters, and maintains an action memory.
  • Retriever — Hybrid search engine combining BM25+ keyword ranking and semantic embeddings for skill retrieval from the indexed library.
  • Mixture Grounding — Refines click coordinates using vision grounding models (UI-TARS v1 endpoint) in a mixture-of-experts pattern.
  • Action System — 20+ base actions (click, type, scroll, hotkey, etc.) with a registry pattern, plus composable action graphs (DAGs) for multi-step skills.
  • Desktop Environment — Wraps pyautogui and pywinauto for screenshots, accessibility tree extraction, and action execution.

Supported Applications

ApplicationSkill Module
Windows Start / Search / Runcommon_action.py
Bing Searchbing_search_action.py
Google Chromechrome_actions.py
Microsoft Edgemicrosoft_edge_action.py
Microsoft Excelexcel_action.py
Microsoft Wordword_action.py
Microsoft PowerPointpowerpoint_action.py
Notepadnotepad_action.py
Paintpaint_action.py
Calculatorcalculator_action.py
Clockclock_action.py
File Explorerfile_explorer_action.py
VLC Media Playervlc_action.py
VS Codevs_code_action.py
Amazonamazon_action.py
YouTubeyoutube_action.py
Windows Settingswindows_settings_action.py

Getting Started

Prerequisites

  • Python 3.10+
  • Windows OS (agent interacts with the Windows desktop)
  • Azure OpenAI access (for LLM planner and grounding models)
  • (Optional) Local Qwen model for offline planner

Installation

  1. Clone the repository:

    git clone <repo-url>cd cua_skill
  2. Install dependencies:

    pip install -r agent/requirements.txt
  3. For Windows Agent Arena evaluation, install additional dependencies:

    pip install -r agent/requirements_waa.txt

Configuration

  1. Create a .env file in the agent/ directory:

    UITARS_V1_BEARER_KEY="your_uitars_key"
    AZURE_AD_TOKEN=""
    
  2. Configure the agent via JSON config files:

    • agent/config.json — Replay agent settings (environment, grounding, logging)
    • agent/config_rag.json — RAG agent settings (planner model, retrieval parameters, RAG index paths)

Key configuration options:

SettingFileDescription
mixture_grounding.expertisesconfig.jsonVision grounding model endpoints and weights
planner.model_classconfig_rag.jsonLLM planner backend: "gpt" or "qwen"
rag.semantic_weightconfig_rag.jsonBlending weight for hybrid search (default: 0.7)
max_stepsBothMaximum actions per task (default: 50)
max_wall_timeBothTimeout in seconds (default: 300)

Agent Modes

Replay Agent

The Replay Agent (CUAKnowledgeGraphAgent) executes pre-defined task graphs:

  1. Loads a task definition (JSON) as a directed graph of actions
  2. Pops steps one-by-one in topological order
  3. For each step, applies mixture grounding to refine coordinates
  4. Executes the action via the desktop environment
fromagentimportCUAKnowledgeGraphAgentagent=CUAKnowledgeGraphAgent(config="agent/config.json")
agent.proceed(instruction="Open Notepad and type Hello", example=task_json)

RAG Agent

The RAG Agent (CUARAGAgent) dynamically plans and executes tasks:

  1. Feasibility check — LLM evaluates if the task is achievable from the current screen
  2. Loop until completion or timeout:
    • Capture screenshot observation
    • Generate search queries from task description + action history + screenshot
    • Retrieve skills via hybrid BM25 + semantic search over the indexed skill library
    • Select the best action from retrieved candidates (with base action fallbacks)
    • Configure action parameters using the LLM + current screenshot
    • Ground coordinates via the mixture grounding model (UI-TARS)
    • Execute the action
    • Update memory — LLM summarizes the outcome for future context
fromagentimportCUARAGAgentagent=CUARAGAgent(config="agent/config_rag.json")
agent.proceed(instruction="Search for 'weather today' on Bing")

User Task Generation

The user_task_generation/ module provides a pipeline for synthesizing diverse user tasks from primitive operations and compositions:

  1. Define primitive operations per application (e.g., BingSearchLaunch, InsertImage) with argument generators
  2. Create compositions combining primitives into multi-step user tasks
  3. Generate tasks with automatic argument filling, instruction drop-off for diversity, and LLM-based rephrasing
python user_task_generation/user_task_generator.py \
--primitive-operation ./asset/primitive_operation/bingsearch_primitive_operation.json \
--composition ./asset/primitive_operation_composition/bingsearch_primitive_operation_composition.json \
--app-name bingsearch \
--out-dir ./asset/user_task \
--num-tasks 10000

See user_task_generation/README.md for full documentation.


Evaluation (Windows Agent Arena)

CUA Skill can be evaluated in the Windows Agent Arena environment:

  1. Set up the Windows Agent Arena Docker image (see evaluation/WindowsAgentArena/README.md)
  2. Place test JSON files in evaluation/WindowsAgentArena/test_jsons/
  3. Run evaluation:
    cd evaluation/WindowsAgentArena
    sudo bash ./run_cua_rag.sh <test_json_filename> [options]

Available options:

  • --use_gold_image — Use the clean backup storage image
  • --clean_mode — Reset environment between test cases (recommended)
  • --reset_image — Regenerate storage from setup ISO

Project Structure

cua_skill/
├── agent/ # Core agent framework
│ ├── agent.py # Replay agent (CUAKnowledgeGraphAgent)
│ ├── agent_rag.py # RAG agent (CUARAGAgent)
│ ├── agent_waa.py # WAA adapter for replay agent
│ ├── agent_rag_waa.py # WAA adapter for RAG agent
│ ├── planner.py # LLM-based planning (query gen, action selection, config)
│ ├── retrieval.py # Hybrid BM25 + semantic retrieval engine
│ ├── mixture_grounding.py # Vision-based coordinate grounding (UI-TARS)
│ ├── llms.py # LLM clients (Azure OpenAI GPT, local Qwen)
│ ├── desktop_env.py # Desktop environment wrapper
│ ├── replay_task.py # Task graph parser and executor
│ ├── config.json # Replay agent configuration
│ ├── config_rag.json # RAG agent configuration
│ ├── action/ # Action system
│ │ ├── base_action.py # Base actions with registry pattern
│ │ ├── compose_action.py # Composable action DAGs
│ │ ├── <app>_action.py # Application-specific skill modules
│ │ └── argument.py # Action argument definitions
│ └── utils/ # Utilities (logging, UIA, config, etc.)
├── user_task_generation/ # Task synthesis pipeline
│ ├── user_task_generator.py # Main generator script
│ ├── argument_value_generator/ # Realistic argument value generators
│ └── README.md # Task generation documentation
├── evaluation/ # Evaluation tooling
│ └── WindowsAgentArena/ # WAA integration and analysis notebooks
├── docs/ # Project documentation website
├── LICENSE # MIT License
└── README.md # This file

License

This project is licensed under the MIT License. See LICENSE for details.

Copyright (c) 2026 Microsoft Corporation.

About

No description, website, or topics provided.

Resources

Code of conduct

Security policy

Stars

48 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

CUA Skill — Computer Use Agent with Skills

CUA Skill is a skill-based autonomous GUI agent framework for Windows desktop applications. Instead of generating every low-level action from scratch, CUA Skill retrieves and executes pre-recorded action sequences (skills) from an indexed library, enabling reliable and efficient task completion across 17+ Windows applications.


Table of Contents


Overview

CUA Skill provides two agent modes for automating desktop tasks:

ModeDescription
Replay AgentExecutes pre-defined task graphs step-by-step with vision-based grounding
RAG AgentUses an LLM planner with Retrieval-Augmented Generation (RAG) to dynamically retrieve, select, and configure skills at runtime

The core insight is that many desktop tasks share common action patterns (e.g., opening apps, clicking menus, typing text). By indexing these patterns as reusable skills, the agent can accomplish complex tasks through composition rather than trial-and-error.


Architecture

┌───────────────────────────────────────────────────────┐
│ CUA Skill Agent │
├──────────────┬───────────────┬────────────────────────┤
│ Planner │ Retriever │ Mixture Grounding │
│ (LLM-based) │ (BM25 + │ (UI-TARS / UIA Tree) │
│ │ Semantic) │ │
├──────────────┴───────────────┴────────────────────────┤
│ Action System │
│ ┌────────────┐ ┌────────────────┐ ┌─────────────┐ │
│ │ Base │ │ Common │ │ App-Specific│ │
│ │ Actions │ │ Actions │ │ Skills │ │
│ └────────────┘ └────────────────┘ └─────────────┘ │
├───────────────────────────────────────────────────────┤
│ Desktop Environment │
│ (Screenshots, A11y Tree, pyautogui) │
└───────────────────────────────────────────────────────┘

Key Components

  • Planner — LLM-driven (GPT or Qwen) component that generates search queries, selects the best action from retrieved candidates, configures action parameters, and maintains an action memory.
  • Retriever — Hybrid search engine combining BM25+ keyword ranking and semantic embeddings for skill retrieval from the indexed library.
  • Mixture Grounding — Refines click coordinates using vision grounding models (UI-TARS v1 endpoint) in a mixture-of-experts pattern.
  • Action System — 20+ base actions (click, type, scroll, hotkey, etc.) with a registry pattern, plus composable action graphs (DAGs) for multi-step skills.
  • Desktop Environment — Wraps pyautogui and pywinauto for screenshots, accessibility tree extraction, and action execution.

Supported Applications

ApplicationSkill Module
Windows Start / Search / Runcommon_action.py
Bing Searchbing_search_action.py
Google Chromechrome_actions.py
Microsoft Edgemicrosoft_edge_action.py
Microsoft Excelexcel_action.py
Microsoft Wordword_action.py
Microsoft PowerPointpowerpoint_action.py
Notepadnotepad_action.py
Paintpaint_action.py
Calculatorcalculator_action.py
Clockclock_action.py
File Explorerfile_explorer_action.py
VLC Media Playervlc_action.py
VS Codevs_code_action.py
Amazonamazon_action.py
YouTubeyoutube_action.py
Windows Settingswindows_settings_action.py

Getting Started

Prerequisites

  • Python 3.10+
  • Windows OS (agent interacts with the Windows desktop)
  • Azure OpenAI access (for LLM planner and grounding models)
  • (Optional) Local Qwen model for offline planner

Installation

  1. Clone the repository:

    git clone <repo-url>cd cua_skill
  2. Install dependencies:

    pip install -r agent/requirements.txt
  3. For Windows Agent Arena evaluation, install additional dependencies:

    pip install -r agent/requirements_waa.txt

Configuration

  1. Create a .env file in the agent/ directory:

    UITARS_V1_BEARER_KEY="your_uitars_key"
    AZURE_AD_TOKEN=""
    
  2. Configure the agent via JSON config files:

    • agent/config.json — Replay agent settings (environment, grounding, logging)
    • agent/config_rag.json — RAG agent settings (planner model, retrieval parameters, RAG index paths)

Key configuration options:

SettingFileDescription
mixture_grounding.expertisesconfig.jsonVision grounding model endpoints and weights
planner.model_classconfig_rag.jsonLLM planner backend: "gpt" or "qwen"
rag.semantic_weightconfig_rag.jsonBlending weight for hybrid search (default: 0.7)
max_stepsBothMaximum actions per task (default: 50)
max_wall_timeBothTimeout in seconds (default: 300)

Agent Modes

Replay Agent

The Replay Agent (CUAKnowledgeGraphAgent) executes pre-defined task graphs:

  1. Loads a task definition (JSON) as a directed graph of actions
  2. Pops steps one-by-one in topological order
  3. For each step, applies mixture grounding to refine coordinates
  4. Executes the action via the desktop environment
fromagentimportCUAKnowledgeGraphAgentagent=CUAKnowledgeGraphAgent(config="agent/config.json")
agent.proceed(instruction="Open Notepad and type Hello", example=task_json)

RAG Agent

The RAG Agent (CUARAGAgent) dynamically plans and executes tasks:

  1. Feasibility check — LLM evaluates if the task is achievable from the current screen
  2. Loop until completion or timeout:
    • Capture screenshot observation
    • Generate search queries from task description + action history + screenshot
    • Retrieve skills via hybrid BM25 + semantic search over the indexed skill library
    • Select the best action from retrieved candidates (with base action fallbacks)
    • Configure action parameters using the LLM + current screenshot
    • Ground coordinates via the mixture grounding model (UI-TARS)
    • Execute the action
    • Update memory — LLM summarizes the outcome for future context
fromagentimportCUARAGAgentagent=CUARAGAgent(config="agent/config_rag.json")
agent.proceed(instruction="Search for 'weather today' on Bing")

User Task Generation

The user_task_generation/ module provides a pipeline for synthesizing diverse user tasks from primitive operations and compositions:

  1. Define primitive operations per application (e.g., BingSearchLaunch, InsertImage) with argument generators
  2. Create compositions combining primitives into multi-step user tasks
  3. Generate tasks with automatic argument filling, instruction drop-off for diversity, and LLM-based rephrasing
python user_task_generation/user_task_generator.py \
--primitive-operation ./asset/primitive_operation/bingsearch_primitive_operation.json \
--composition ./asset/primitive_operation_composition/bingsearch_primitive_operation_composition.json \
--app-name bingsearch \
--out-dir ./asset/user_task \
--num-tasks 10000

See user_task_generation/README.md for full documentation.


Evaluation (Windows Agent Arena)

CUA Skill can be evaluated in the Windows Agent Arena environment:

  1. Set up the Windows Agent Arena Docker image (see evaluation/WindowsAgentArena/README.md)
  2. Place test JSON files in evaluation/WindowsAgentArena/test_jsons/
  3. Run evaluation:
    cd evaluation/WindowsAgentArena
    sudo bash ./run_cua_rag.sh <test_json_filename> [options]

Available options:

  • --use_gold_image — Use the clean backup storage image
  • --clean_mode — Reset environment between test cases (recommended)
  • --reset_image — Regenerate storage from setup ISO

Project Structure

cua_skill/
├── agent/ # Core agent framework
│ ├── agent.py # Replay agent (CUAKnowledgeGraphAgent)
│ ├── agent_rag.py # RAG agent (CUARAGAgent)
│ ├── agent_waa.py # WAA adapter for replay agent
│ ├── agent_rag_waa.py # WAA adapter for RAG agent
│ ├── planner.py # LLM-based planning (query gen, action selection, config)
│ ├── retrieval.py # Hybrid BM25 + semantic retrieval engine
│ ├── mixture_grounding.py # Vision-based coordinate grounding (UI-TARS)
│ ├── llms.py # LLM clients (Azure OpenAI GPT, local Qwen)
│ ├── desktop_env.py # Desktop environment wrapper
│ ├── replay_task.py # Task graph parser and executor
│ ├── config.json # Replay agent configuration
│ ├── config_rag.json # RAG agent configuration
│ ├── action/ # Action system
│ │ ├── base_action.py # Base actions with registry pattern
│ │ ├── compose_action.py # Composable action DAGs
│ │ ├── <app>_action.py # Application-specific skill modules
│ │ └── argument.py # Action argument definitions
│ └── utils/ # Utilities (logging, UIA, config, etc.)
├── user_task_generation/ # Task synthesis pipeline
│ ├── user_task_generator.py # Main generator script
│ ├── argument_value_generator/ # Realistic argument value generators
│ └── README.md # Task generation documentation
├── evaluation/ # Evaluation tooling
│ └── WindowsAgentArena/ # WAA integration and analysis notebooks
├── docs/ # Project documentation website
├── LICENSE # MIT License
└── README.md # This file

License

This project is licensed under the MIT License. See LICENSE for details.

Copyright (c) 2026 Microsoft Corporation.

About

No description, website, or topics provided.

Resources

Code of conduct

Security policy

Stars

48 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages