Repository files navigation

SAfactory

中文 | English

SAfactory is a scalable infrastructure for agent evaluation, trajectory collection, and reinforcement learning training. It schedules agents and benchmark environments as external runtimes, routes model calls through a session-aware OpenAI-compatible Gateway, records trajectories, and feeds completed rollouts into training systems such as Slime.

Why SAfactoryDemoAgent SkillQuick StartDocumentationCitation

PythonLicenseExecutionLLM


✨ Why SAfactory

SAfactory provides one unified workflow for agent onboarding, benchmark onboarding, evaluation, rollout data generation, and RL training.

The core runtime contract is:

  • one dataset row becomes one scheduled episode;
  • each episode gets its own session_id and Gateway session;
  • the runtime calls the target model through the Gateway;
  • the runtime returns one JSON result;
  • rule_evaluator.py converts runtime output and trajectory data into the SAfactory format;
  • completed trainable trajectories can be consumed by the RL Buffer Server or persisted as reusable data assets.

🎬 Demo

demo.1.mp4

Click to watch the full demo

🧩 Agent Skill Quick Start

This repository includes a lightweight Agent skill that helps agents use SAfactory through the standard workflows:

skills/safactory-workflows/SKILL.md

It covers three common requests:

  • onboard a new benchmark or custom environment into SAfactory;
  • run Docker-mode evaluation for a selected environment;
  • start GRPO / RL training for a selected environment.

When working with an Agent, use prompts such as:

Use skills/safactory-workflows to help me onboard this benchmark into SAfactory.
Use the safactory-workflows skill to run geo3k evaluation in Docker mode.
Use the safactory-workflows skill to start GRPO training for my_env.

The skill does not replace the docs. It guides the Agent to read docs/guides/, docs/reference/, and the root README as needed, while using the standard env/geo3k/ environment as the reference implementation. If your Agent supports local skill discovery, add skills/safactory-workflows/ to its skill search path; otherwise mention this path explicitly in the request.

🚀 Quick Start

1. SAfactory Installation And Gateway Configuration

Before running Geo3K, prepare the runtime image and dataset as described in Standard Environment: Geo3K.

Install SAfactory:

git clone https://github.com/AI45Lab/SAfactory.git
cd SAfactory
pip install -U -r requirements.txt

Docker must be available locally, and the current user must be able to run docker build, docker run, and docker exec.

Create a local Gateway config:

cp gateway/config.example.yaml gateway/config.local.yaml

Edit gateway/config.local.yaml and explicitly set the storage path and model route:

listen_host: 0.0.0.0listen_port: 8000base_session_path: /v1/sessionsmax_steps: -1storage_type: sqlitestorage_config:
db_url: sqlite://env_trajs.dbllm_routes:
geo3k_model:
base_url: http://YOUR_LLM_HOST/v1api_key: YOUR_API_KEYsupports_stream: truemax_concurrency: 64

Start the Gateway:

python -m gateway --config gateway/config.local.yaml

Check readiness from another terminal:

curl http://127.0.0.1:8000/readyz

When using Cloud storage, the recommended configuration is to omit landing_table and env_config_table from the Gateway YAML and set only one profile in .env:

WT_SDK_PROFILE=test

or:

WT_SDK_PROFILE=production

The test profile selects landing_test and env_config_test; the production/prod profile selects wind_tunnel_landing and evaluation_env_config. This keeps trajectory and environment-config writes in the same environment. Explicit table names in storage_config override the profile and should be used only for intentional special cases. SAfactory does not access the serving table.

2. Minimal Evaluation With Geo3K

Run a minimal Geo3K evaluation in Docker mode:

python launcher.py \
--mode docker \
--agent-config env/geo3k/geo3k_config.yaml \
--agent-start-config env/geo3k/geo3k_start.yaml \
--gateway-base-url http://127.0.0.1:8000/v1/sessions \
--llm-model geo3k_model \
--enable-evaluation \
--db-path sqlite://env_trajs.db \
--job-id geo3k-docker-smoke \
--pool-size 1 \
--max-workers 1 \
--max-steps 10

--llm-model must match a key under llm_routes. With --enable-evaluation, SAfactory calls env/geo3k/rule_evaluator.py and writes the final reward.

3. Minimal Training With Geo3K

Geo3K training uses the RL bridge under rl/ and the example config at rl/examples/geo3k_vl/env.sh.

Before starting, edit rl/examples/geo3k_vl/env.sh, or export the same variables in each terminal. Set the required local paths and scale the first run down:

export AIEVOBOX_ROOT=$(pwd)export AIEVOBOX_MODE=docker
export AIEVOBOX_AGENT_CONFIG=${AIEVOBOX_ROOT}/env/geo3k/geo3k_config.yaml
export AIEVOBOX_AGENT_START_CONFIG=${AIEVOBOX_ROOT}/env/geo3k/geo3k_start.yaml
export AIEVOBOX_GATEWAY_HOST=127.0.0.1
export AIEVOBOX_GATEWAY_PORT=8000
export RL_MODEL=geo3k_model
export AIEVOBOX_POOL_SIZE=2
export RL_GROUP_SIZE=2
export RL_EPOCH=1
export HF_CKPT_DIR=/path/to/hf-checkpoint
export SLIME_HOME=/path/to/slime
export MEGATRON_HOME=/path/to/Megatron-LM

Then start two processes from the repository root:

RL_ENV_SH=rl/examples/geo3k_vl/env.sh bash rl/run_slime_generator.sh
RL_ENV_SH=rl/examples/geo3k_vl/env.sh bash rl/run_buffer_server.sh

Rewards for RL training come from completed trajectories in the database: rule_evaluator.py writes the reward after evaluation, and Buffer Server fetches trajectory rows with rewards from the same database used by Launcher / Gateway before serving them to the training process.

Buffer Server can automatically start a Gateway and route RL_MODEL to the Slime-hosted LLM proxy. If a manually started Gateway is already using the same port, stop it first. Set AIEVOBOX_GATEWAY_AUTOSTART=0 only when the external Gateway already has the correct route and storage configuration.

📚 Documentation

Guides

GuideWhat it covers
Custom EnvironmentsHow to onboard a new external runtime adapter.
EvaluationRule evaluator discovery, interfaces, reward writing, and Geo3K evaluation.
RL TrainingBuffer Server, Slime generator, Geo3K training path, and key variables.
Data ManagerSQLite storage behavior, table schema, row types, and query examples.
S3 + LanceDB StorageHow to switch trajectory storage from local SQLite to S3 + LanceDB.

Internal

Internal DocWhat it covers
RJob ModeRemote RJob runtime config, authentication, mounts, Gateway reachability, and Geo3K examples.
Sandbox ModeBrainbox Sandbox Environment config, volumes, lifecycle, and launch flow.

Reference

ReferenceWhat it covers
CLI And ConfigurationLauncher flags, Gateway config, agent config, start config, RJob, and Sandbox settings.
Supported EnvironmentsChecked-in adapters, the standard Geo3K path, and the runtime matrix.
Gateway ReferenceOpenAI-compatible routes, session endpoints, telemetry, request logs, and storage consistency.
ReportSAfactory report.

📖 Citation

If SAfactory or SAfactory-generated datasets are useful in your work, please cite this repository and the specific dataset or report you used.

@misc{chen2026safactoryscalableagenticinfrastructure,
title={SAfactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence},
author={Shanghai AI Lab},
year={2026},
eprint={2605.06230},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2605.06230},
}

About

SAfactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence

Resources

Stars

197 stars

Watchers

13 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

SAfactory

中文 | English

SAfactory is a scalable infrastructure for agent evaluation, trajectory collection, and reinforcement learning training. It schedules agents and benchmark environments as external runtimes, routes model calls through a session-aware OpenAI-compatible Gateway, records trajectories, and feeds completed rollouts into training systems such as Slime.

Why SAfactoryDemoAgent SkillQuick StartDocumentationCitation

PythonLicenseExecutionLLM


✨ Why SAfactory

SAfactory provides one unified workflow for agent onboarding, benchmark onboarding, evaluation, rollout data generation, and RL training.

The core runtime contract is:

  • one dataset row becomes one scheduled episode;
  • each episode gets its own session_id and Gateway session;
  • the runtime calls the target model through the Gateway;
  • the runtime returns one JSON result;
  • rule_evaluator.py converts runtime output and trajectory data into the SAfactory format;
  • completed trainable trajectories can be consumed by the RL Buffer Server or persisted as reusable data assets.

🎬 Demo

demo.1.mp4

Click to watch the full demo

🧩 Agent Skill Quick Start

This repository includes a lightweight Agent skill that helps agents use SAfactory through the standard workflows:

skills/safactory-workflows/SKILL.md

It covers three common requests:

  • onboard a new benchmark or custom environment into SAfactory;
  • run Docker-mode evaluation for a selected environment;
  • start GRPO / RL training for a selected environment.

When working with an Agent, use prompts such as:

Use skills/safactory-workflows to help me onboard this benchmark into SAfactory.
Use the safactory-workflows skill to run geo3k evaluation in Docker mode.
Use the safactory-workflows skill to start GRPO training for my_env.

The skill does not replace the docs. It guides the Agent to read docs/guides/, docs/reference/, and the root README as needed, while using the standard env/geo3k/ environment as the reference implementation. If your Agent supports local skill discovery, add skills/safactory-workflows/ to its skill search path; otherwise mention this path explicitly in the request.

🚀 Quick Start

1. SAfactory Installation And Gateway Configuration

Before running Geo3K, prepare the runtime image and dataset as described in Standard Environment: Geo3K.

Install SAfactory:

git clone https://github.com/AI45Lab/SAfactory.git
cd SAfactory
pip install -U -r requirements.txt

Docker must be available locally, and the current user must be able to run docker build, docker run, and docker exec.

Create a local Gateway config:

cp gateway/config.example.yaml gateway/config.local.yaml

Edit gateway/config.local.yaml and explicitly set the storage path and model route:

listen_host: 0.0.0.0listen_port: 8000base_session_path: /v1/sessionsmax_steps: -1storage_type: sqlitestorage_config:
db_url: sqlite://env_trajs.dbllm_routes:
geo3k_model:
base_url: http://YOUR_LLM_HOST/v1api_key: YOUR_API_KEYsupports_stream: truemax_concurrency: 64

Start the Gateway:

python -m gateway --config gateway/config.local.yaml

Check readiness from another terminal:

curl http://127.0.0.1:8000/readyz

When using Cloud storage, the recommended configuration is to omit landing_table and env_config_table from the Gateway YAML and set only one profile in .env:

WT_SDK_PROFILE=test

or:

WT_SDK_PROFILE=production

The test profile selects landing_test and env_config_test; the production/prod profile selects wind_tunnel_landing and evaluation_env_config. This keeps trajectory and environment-config writes in the same environment. Explicit table names in storage_config override the profile and should be used only for intentional special cases. SAfactory does not access the serving table.

2. Minimal Evaluation With Geo3K

Run a minimal Geo3K evaluation in Docker mode:

python launcher.py \
--mode docker \
--agent-config env/geo3k/geo3k_config.yaml \
--agent-start-config env/geo3k/geo3k_start.yaml \
--gateway-base-url http://127.0.0.1:8000/v1/sessions \
--llm-model geo3k_model \
--enable-evaluation \
--db-path sqlite://env_trajs.db \
--job-id geo3k-docker-smoke \
--pool-size 1 \
--max-workers 1 \
--max-steps 10

--llm-model must match a key under llm_routes. With --enable-evaluation, SAfactory calls env/geo3k/rule_evaluator.py and writes the final reward.

3. Minimal Training With Geo3K

Geo3K training uses the RL bridge under rl/ and the example config at rl/examples/geo3k_vl/env.sh.

Before starting, edit rl/examples/geo3k_vl/env.sh, or export the same variables in each terminal. Set the required local paths and scale the first run down:

export AIEVOBOX_ROOT=$(pwd)export AIEVOBOX_MODE=docker
export AIEVOBOX_AGENT_CONFIG=${AIEVOBOX_ROOT}/env/geo3k/geo3k_config.yaml
export AIEVOBOX_AGENT_START_CONFIG=${AIEVOBOX_ROOT}/env/geo3k/geo3k_start.yaml
export AIEVOBOX_GATEWAY_HOST=127.0.0.1
export AIEVOBOX_GATEWAY_PORT=8000
export RL_MODEL=geo3k_model
export AIEVOBOX_POOL_SIZE=2
export RL_GROUP_SIZE=2
export RL_EPOCH=1
export HF_CKPT_DIR=/path/to/hf-checkpoint
export SLIME_HOME=/path/to/slime
export MEGATRON_HOME=/path/to/Megatron-LM

Then start two processes from the repository root:

RL_ENV_SH=rl/examples/geo3k_vl/env.sh bash rl/run_slime_generator.sh
RL_ENV_SH=rl/examples/geo3k_vl/env.sh bash rl/run_buffer_server.sh

Rewards for RL training come from completed trajectories in the database: rule_evaluator.py writes the reward after evaluation, and Buffer Server fetches trajectory rows with rewards from the same database used by Launcher / Gateway before serving them to the training process.

Buffer Server can automatically start a Gateway and route RL_MODEL to the Slime-hosted LLM proxy. If a manually started Gateway is already using the same port, stop it first. Set AIEVOBOX_GATEWAY_AUTOSTART=0 only when the external Gateway already has the correct route and storage configuration.

📚 Documentation

Guides

GuideWhat it covers
Custom EnvironmentsHow to onboard a new external runtime adapter.
EvaluationRule evaluator discovery, interfaces, reward writing, and Geo3K evaluation.
RL TrainingBuffer Server, Slime generator, Geo3K training path, and key variables.
Data ManagerSQLite storage behavior, table schema, row types, and query examples.
S3 + LanceDB StorageHow to switch trajectory storage from local SQLite to S3 + LanceDB.

Internal

Internal DocWhat it covers
RJob ModeRemote RJob runtime config, authentication, mounts, Gateway reachability, and Geo3K examples.
Sandbox ModeBrainbox Sandbox Environment config, volumes, lifecycle, and launch flow.

Reference

ReferenceWhat it covers
CLI And ConfigurationLauncher flags, Gateway config, agent config, start config, RJob, and Sandbox settings.
Supported EnvironmentsChecked-in adapters, the standard Geo3K path, and the runtime matrix.
Gateway ReferenceOpenAI-compatible routes, session endpoints, telemetry, request logs, and storage consistency.
ReportSAfactory report.

📖 Citation

If SAfactory or SAfactory-generated datasets are useful in your work, please cite this repository and the specific dataset or report you used.

@misc{chen2026safactoryscalableagenticinfrastructure,
title={SAfactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence},
author={Shanghai AI Lab},
year={2026},
eprint={2605.06230},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2605.06230},
}

About

SAfactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence

Resources

Stars

197 stars

Watchers

13 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SAfactory

中文 | English

SAfactory is a scalable infrastructure for agent evaluation, trajectory collection, and reinforcement learning training. It schedules agents and benchmark environments as external runtimes, routes model calls through a session-aware OpenAI-compatible Gateway, records trajectories, and feeds completed rollouts into training systems such as Slime.

Why SAfactoryDemoAgent SkillQuick StartDocumentationCitation

PythonLicenseExecutionLLM


✨ Why SAfactory

SAfactory provides one unified workflow for agent onboarding, benchmark onboarding, evaluation, rollout data generation, and RL training.

The core runtime contract is:

  • one dataset row becomes one scheduled episode;
  • each episode gets its own session_id and Gateway session;
  • the runtime calls the target model through the Gateway;
  • the runtime returns one JSON result;
  • rule_evaluator.py converts runtime output and trajectory data into the SAfactory format;
  • completed trainable trajectories can be consumed by the RL Buffer Server or persisted as reusable data assets.

🎬 Demo

demo.1.mp4

Click to watch the full demo

🧩 Agent Skill Quick Start

This repository includes a lightweight Agent skill that helps agents use SAfactory through the standard workflows:

skills/safactory-workflows/SKILL.md

It covers three common requests:

  • onboard a new benchmark or custom environment into SAfactory;
  • run Docker-mode evaluation for a selected environment;
  • start GRPO / RL training for a selected environment.

When working with an Agent, use prompts such as:

Use skills/safactory-workflows to help me onboard this benchmark into SAfactory.
Use the safactory-workflows skill to run geo3k evaluation in Docker mode.
Use the safactory-workflows skill to start GRPO training for my_env.

The skill does not replace the docs. It guides the Agent to read docs/guides/, docs/reference/, and the root README as needed, while using the standard env/geo3k/ environment as the reference implementation. If your Agent supports local skill discovery, add skills/safactory-workflows/ to its skill search path; otherwise mention this path explicitly in the request.

🚀 Quick Start

1. SAfactory Installation And Gateway Configuration

Before running Geo3K, prepare the runtime image and dataset as described in Standard Environment: Geo3K.

Install SAfactory:

git clone https://github.com/AI45Lab/SAfactory.git
cd SAfactory
pip install -U -r requirements.txt

Docker must be available locally, and the current user must be able to run docker build, docker run, and docker exec.

Create a local Gateway config:

cp gateway/config.example.yaml gateway/config.local.yaml

Edit gateway/config.local.yaml and explicitly set the storage path and model route:

listen_host: 0.0.0.0listen_port: 8000base_session_path: /v1/sessionsmax_steps: -1storage_type: sqlitestorage_config:
db_url: sqlite://env_trajs.dbllm_routes:
geo3k_model:
base_url: http://YOUR_LLM_HOST/v1api_key: YOUR_API_KEYsupports_stream: truemax_concurrency: 64

Start the Gateway:

python -m gateway --config gateway/config.local.yaml

Check readiness from another terminal:

curl http://127.0.0.1:8000/readyz

When using Cloud storage, the recommended configuration is to omit landing_table and env_config_table from the Gateway YAML and set only one profile in .env:

WT_SDK_PROFILE=test

or:

WT_SDK_PROFILE=production

The test profile selects landing_test and env_config_test; the production/prod profile selects wind_tunnel_landing and evaluation_env_config. This keeps trajectory and environment-config writes in the same environment. Explicit table names in storage_config override the profile and should be used only for intentional special cases. SAfactory does not access the serving table.

2. Minimal Evaluation With Geo3K

Run a minimal Geo3K evaluation in Docker mode:

python launcher.py \
--mode docker \
--agent-config env/geo3k/geo3k_config.yaml \
--agent-start-config env/geo3k/geo3k_start.yaml \
--gateway-base-url http://127.0.0.1:8000/v1/sessions \
--llm-model geo3k_model \
--enable-evaluation \
--db-path sqlite://env_trajs.db \
--job-id geo3k-docker-smoke \
--pool-size 1 \
--max-workers 1 \
--max-steps 10

--llm-model must match a key under llm_routes. With --enable-evaluation, SAfactory calls env/geo3k/rule_evaluator.py and writes the final reward.

3. Minimal Training With Geo3K

Geo3K training uses the RL bridge under rl/ and the example config at rl/examples/geo3k_vl/env.sh.

Before starting, edit rl/examples/geo3k_vl/env.sh, or export the same variables in each terminal. Set the required local paths and scale the first run down:

export AIEVOBOX_ROOT=$(pwd)export AIEVOBOX_MODE=docker
export AIEVOBOX_AGENT_CONFIG=${AIEVOBOX_ROOT}/env/geo3k/geo3k_config.yaml
export AIEVOBOX_AGENT_START_CONFIG=${AIEVOBOX_ROOT}/env/geo3k/geo3k_start.yaml
export AIEVOBOX_GATEWAY_HOST=127.0.0.1
export AIEVOBOX_GATEWAY_PORT=8000
export RL_MODEL=geo3k_model
export AIEVOBOX_POOL_SIZE=2
export RL_GROUP_SIZE=2
export RL_EPOCH=1
export HF_CKPT_DIR=/path/to/hf-checkpoint
export SLIME_HOME=/path/to/slime
export MEGATRON_HOME=/path/to/Megatron-LM

Then start two processes from the repository root:

RL_ENV_SH=rl/examples/geo3k_vl/env.sh bash rl/run_slime_generator.sh
RL_ENV_SH=rl/examples/geo3k_vl/env.sh bash rl/run_buffer_server.sh

Rewards for RL training come from completed trajectories in the database: rule_evaluator.py writes the reward after evaluation, and Buffer Server fetches trajectory rows with rewards from the same database used by Launcher / Gateway before serving them to the training process.

Buffer Server can automatically start a Gateway and route RL_MODEL to the Slime-hosted LLM proxy. If a manually started Gateway is already using the same port, stop it first. Set AIEVOBOX_GATEWAY_AUTOSTART=0 only when the external Gateway already has the correct route and storage configuration.

📚 Documentation

Guides

GuideWhat it covers
Custom EnvironmentsHow to onboard a new external runtime adapter.
EvaluationRule evaluator discovery, interfaces, reward writing, and Geo3K evaluation.
RL TrainingBuffer Server, Slime generator, Geo3K training path, and key variables.
Data ManagerSQLite storage behavior, table schema, row types, and query examples.
S3 + LanceDB StorageHow to switch trajectory storage from local SQLite to S3 + LanceDB.

Internal

Internal DocWhat it covers
RJob ModeRemote RJob runtime config, authentication, mounts, Gateway reachability, and Geo3K examples.
Sandbox ModeBrainbox Sandbox Environment config, volumes, lifecycle, and launch flow.

Reference

ReferenceWhat it covers
CLI And ConfigurationLauncher flags, Gateway config, agent config, start config, RJob, and Sandbox settings.
Supported EnvironmentsChecked-in adapters, the standard Geo3K path, and the runtime matrix.
Gateway ReferenceOpenAI-compatible routes, session endpoints, telemetry, request logs, and storage consistency.
ReportSAfactory report.

📖 Citation

If SAfactory or SAfactory-generated datasets are useful in your work, please cite this repository and the specific dataset or report you used.

@misc{chen2026safactoryscalableagenticinfrastructure,
title={SAfactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence},
author={Shanghai AI Lab},
year={2026},
eprint={2605.06230},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2605.06230},
}

About

SAfactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence

Resources

Stars

197 stars

Watchers

13 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SAfactory

中文 | English

SAfactory is a scalable infrastructure for agent evaluation, trajectory collection, and reinforcement learning training. It schedules agents and benchmark environments as external runtimes, routes model calls through a session-aware OpenAI-compatible Gateway, records trajectories, and feeds completed rollouts into training systems such as Slime.

Why SAfactoryDemoAgent SkillQuick StartDocumentationCitation

PythonLicenseExecutionLLM


✨ Why SAfactory

SAfactory provides one unified workflow for agent onboarding, benchmark onboarding, evaluation, rollout data generation, and RL training.

The core runtime contract is:

  • one dataset row becomes one scheduled episode;
  • each episode gets its own session_id and Gateway session;
  • the runtime calls the target model through the Gateway;
  • the runtime returns one JSON result;
  • rule_evaluator.py converts runtime output and trajectory data into the SAfactory format;
  • completed trainable trajectories can be consumed by the RL Buffer Server or persisted as reusable data assets.

🎬 Demo

demo.1.mp4

Click to watch the full demo

🧩 Agent Skill Quick Start

This repository includes a lightweight Agent skill that helps agents use SAfactory through the standard workflows:

skills/safactory-workflows/SKILL.md

It covers three common requests:

  • onboard a new benchmark or custom environment into SAfactory;
  • run Docker-mode evaluation for a selected environment;
  • start GRPO / RL training for a selected environment.

When working with an Agent, use prompts such as:

Use skills/safactory-workflows to help me onboard this benchmark into SAfactory.
Use the safactory-workflows skill to run geo3k evaluation in Docker mode.
Use the safactory-workflows skill to start GRPO training for my_env.

The skill does not replace the docs. It guides the Agent to read docs/guides/, docs/reference/, and the root README as needed, while using the standard env/geo3k/ environment as the reference implementation. If your Agent supports local skill discovery, add skills/safactory-workflows/ to its skill search path; otherwise mention this path explicitly in the request.

🚀 Quick Start

1. SAfactory Installation And Gateway Configuration

Before running Geo3K, prepare the runtime image and dataset as described in Standard Environment: Geo3K.

Install SAfactory:

git clone https://github.com/AI45Lab/SAfactory.git
cd SAfactory
pip install -U -r requirements.txt

Docker must be available locally, and the current user must be able to run docker build, docker run, and docker exec.

Create a local Gateway config:

cp gateway/config.example.yaml gateway/config.local.yaml

Edit gateway/config.local.yaml and explicitly set the storage path and model route:

listen_host: 0.0.0.0listen_port: 8000base_session_path: /v1/sessionsmax_steps: -1storage_type: sqlitestorage_config:
db_url: sqlite://env_trajs.dbllm_routes:
geo3k_model:
base_url: http://YOUR_LLM_HOST/v1api_key: YOUR_API_KEYsupports_stream: truemax_concurrency: 64

Start the Gateway:

python -m gateway --config gateway/config.local.yaml

Check readiness from another terminal:

curl http://127.0.0.1:8000/readyz

When using Cloud storage, the recommended configuration is to omit landing_table and env_config_table from the Gateway YAML and set only one profile in .env:

WT_SDK_PROFILE=test

or:

WT_SDK_PROFILE=production

The test profile selects landing_test and env_config_test; the production/prod profile selects wind_tunnel_landing and evaluation_env_config. This keeps trajectory and environment-config writes in the same environment. Explicit table names in storage_config override the profile and should be used only for intentional special cases. SAfactory does not access the serving table.

2. Minimal Evaluation With Geo3K

Run a minimal Geo3K evaluation in Docker mode:

python launcher.py \
--mode docker \
--agent-config env/geo3k/geo3k_config.yaml \
--agent-start-config env/geo3k/geo3k_start.yaml \
--gateway-base-url http://127.0.0.1:8000/v1/sessions \
--llm-model geo3k_model \
--enable-evaluation \
--db-path sqlite://env_trajs.db \
--job-id geo3k-docker-smoke \
--pool-size 1 \
--max-workers 1 \
--max-steps 10

--llm-model must match a key under llm_routes. With --enable-evaluation, SAfactory calls env/geo3k/rule_evaluator.py and writes the final reward.

3. Minimal Training With Geo3K

Geo3K training uses the RL bridge under rl/ and the example config at rl/examples/geo3k_vl/env.sh.

Before starting, edit rl/examples/geo3k_vl/env.sh, or export the same variables in each terminal. Set the required local paths and scale the first run down:

export AIEVOBOX_ROOT=$(pwd)export AIEVOBOX_MODE=docker
export AIEVOBOX_AGENT_CONFIG=${AIEVOBOX_ROOT}/env/geo3k/geo3k_config.yaml
export AIEVOBOX_AGENT_START_CONFIG=${AIEVOBOX_ROOT}/env/geo3k/geo3k_start.yaml
export AIEVOBOX_GATEWAY_HOST=127.0.0.1
export AIEVOBOX_GATEWAY_PORT=8000
export RL_MODEL=geo3k_model
export AIEVOBOX_POOL_SIZE=2
export RL_GROUP_SIZE=2
export RL_EPOCH=1
export HF_CKPT_DIR=/path/to/hf-checkpoint
export SLIME_HOME=/path/to/slime
export MEGATRON_HOME=/path/to/Megatron-LM

Then start two processes from the repository root:

RL_ENV_SH=rl/examples/geo3k_vl/env.sh bash rl/run_slime_generator.sh
RL_ENV_SH=rl/examples/geo3k_vl/env.sh bash rl/run_buffer_server.sh

Rewards for RL training come from completed trajectories in the database: rule_evaluator.py writes the reward after evaluation, and Buffer Server fetches trajectory rows with rewards from the same database used by Launcher / Gateway before serving them to the training process.

Buffer Server can automatically start a Gateway and route RL_MODEL to the Slime-hosted LLM proxy. If a manually started Gateway is already using the same port, stop it first. Set AIEVOBOX_GATEWAY_AUTOSTART=0 only when the external Gateway already has the correct route and storage configuration.

📚 Documentation

Guides

GuideWhat it covers
Custom EnvironmentsHow to onboard a new external runtime adapter.
EvaluationRule evaluator discovery, interfaces, reward writing, and Geo3K evaluation.
RL TrainingBuffer Server, Slime generator, Geo3K training path, and key variables.
Data ManagerSQLite storage behavior, table schema, row types, and query examples.
S3 + LanceDB StorageHow to switch trajectory storage from local SQLite to S3 + LanceDB.

Internal

Internal DocWhat it covers
RJob ModeRemote RJob runtime config, authentication, mounts, Gateway reachability, and Geo3K examples.
Sandbox ModeBrainbox Sandbox Environment config, volumes, lifecycle, and launch flow.

Reference

ReferenceWhat it covers
CLI And ConfigurationLauncher flags, Gateway config, agent config, start config, RJob, and Sandbox settings.
Supported EnvironmentsChecked-in adapters, the standard Geo3K path, and the runtime matrix.
Gateway ReferenceOpenAI-compatible routes, session endpoints, telemetry, request logs, and storage consistency.
ReportSAfactory report.

📖 Citation

If SAfactory or SAfactory-generated datasets are useful in your work, please cite this repository and the specific dataset or report you used.

@misc{chen2026safactoryscalableagenticinfrastructure,
title={SAfactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence},
author={Shanghai AI Lab},
year={2026},
eprint={2605.06230},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2605.06230},
}

About

SAfactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence

Resources

Stars

197 stars

Watchers

13 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

SAfactory

中文 | English

SAfactory is a scalable infrastructure for agent evaluation, trajectory collection, and reinforcement learning training. It schedules agents and benchmark environments as external runtimes, routes model calls through a session-aware OpenAI-compatible Gateway, records trajectories, and feeds completed rollouts into training systems such as Slime.

Why SAfactoryDemoAgent SkillQuick StartDocumentationCitation

PythonLicenseExecutionLLM


✨ Why SAfactory

SAfactory provides one unified workflow for agent onboarding, benchmark onboarding, evaluation, rollout data generation, and RL training.

The core runtime contract is:

  • one dataset row becomes one scheduled episode;
  • each episode gets its own session_id and Gateway session;
  • the runtime calls the target model through the Gateway;
  • the runtime returns one JSON result;
  • rule_evaluator.py converts runtime output and trajectory data into the SAfactory format;
  • completed trainable trajectories can be consumed by the RL Buffer Server or persisted as reusable data assets.

🎬 Demo

demo.1.mp4

Click to watch the full demo

🧩 Agent Skill Quick Start

This repository includes a lightweight Agent skill that helps agents use SAfactory through the standard workflows:

skills/safactory-workflows/SKILL.md

It covers three common requests:

  • onboard a new benchmark or custom environment into SAfactory;
  • run Docker-mode evaluation for a selected environment;
  • start GRPO / RL training for a selected environment.

When working with an Agent, use prompts such as:

Use skills/safactory-workflows to help me onboard this benchmark into SAfactory.
Use the safactory-workflows skill to run geo3k evaluation in Docker mode.
Use the safactory-workflows skill to start GRPO training for my_env.

The skill does not replace the docs. It guides the Agent to read docs/guides/, docs/reference/, and the root README as needed, while using the standard env/geo3k/ environment as the reference implementation. If your Agent supports local skill discovery, add skills/safactory-workflows/ to its skill search path; otherwise mention this path explicitly in the request.

🚀 Quick Start

1. SAfactory Installation And Gateway Configuration

Before running Geo3K, prepare the runtime image and dataset as described in Standard Environment: Geo3K.

Install SAfactory:

git clone https://github.com/AI45Lab/SAfactory.git
cd SAfactory
pip install -U -r requirements.txt

Docker must be available locally, and the current user must be able to run docker build, docker run, and docker exec.

Create a local Gateway config:

cp gateway/config.example.yaml gateway/config.local.yaml

Edit gateway/config.local.yaml and explicitly set the storage path and model route:

listen_host: 0.0.0.0listen_port: 8000base_session_path: /v1/sessionsmax_steps: -1storage_type: sqlitestorage_config:
db_url: sqlite://env_trajs.dbllm_routes:
geo3k_model:
base_url: http://YOUR_LLM_HOST/v1api_key: YOUR_API_KEYsupports_stream: truemax_concurrency: 64

Start the Gateway:

python -m gateway --config gateway/config.local.yaml

Check readiness from another terminal:

curl http://127.0.0.1:8000/readyz

When using Cloud storage, the recommended configuration is to omit landing_table and env_config_table from the Gateway YAML and set only one profile in .env:

WT_SDK_PROFILE=test

or:

WT_SDK_PROFILE=production

The test profile selects landing_test and env_config_test; the production/prod profile selects wind_tunnel_landing and evaluation_env_config. This keeps trajectory and environment-config writes in the same environment. Explicit table names in storage_config override the profile and should be used only for intentional special cases. SAfactory does not access the serving table.

2. Minimal Evaluation With Geo3K

Run a minimal Geo3K evaluation in Docker mode:

python launcher.py \
--mode docker \
--agent-config env/geo3k/geo3k_config.yaml \
--agent-start-config env/geo3k/geo3k_start.yaml \
--gateway-base-url http://127.0.0.1:8000/v1/sessions \
--llm-model geo3k_model \
--enable-evaluation \
--db-path sqlite://env_trajs.db \
--job-id geo3k-docker-smoke \
--pool-size 1 \
--max-workers 1 \
--max-steps 10

--llm-model must match a key under llm_routes. With --enable-evaluation, SAfactory calls env/geo3k/rule_evaluator.py and writes the final reward.

3. Minimal Training With Geo3K

Geo3K training uses the RL bridge under rl/ and the example config at rl/examples/geo3k_vl/env.sh.

Before starting, edit rl/examples/geo3k_vl/env.sh, or export the same variables in each terminal. Set the required local paths and scale the first run down:

export AIEVOBOX_ROOT=$(pwd)export AIEVOBOX_MODE=docker
export AIEVOBOX_AGENT_CONFIG=${AIEVOBOX_ROOT}/env/geo3k/geo3k_config.yaml
export AIEVOBOX_AGENT_START_CONFIG=${AIEVOBOX_ROOT}/env/geo3k/geo3k_start.yaml
export AIEVOBOX_GATEWAY_HOST=127.0.0.1
export AIEVOBOX_GATEWAY_PORT=8000
export RL_MODEL=geo3k_model
export AIEVOBOX_POOL_SIZE=2
export RL_GROUP_SIZE=2
export RL_EPOCH=1
export HF_CKPT_DIR=/path/to/hf-checkpoint
export SLIME_HOME=/path/to/slime
export MEGATRON_HOME=/path/to/Megatron-LM

Then start two processes from the repository root:

RL_ENV_SH=rl/examples/geo3k_vl/env.sh bash rl/run_slime_generator.sh
RL_ENV_SH=rl/examples/geo3k_vl/env.sh bash rl/run_buffer_server.sh

Rewards for RL training come from completed trajectories in the database: rule_evaluator.py writes the reward after evaluation, and Buffer Server fetches trajectory rows with rewards from the same database used by Launcher / Gateway before serving them to the training process.

Buffer Server can automatically start a Gateway and route RL_MODEL to the Slime-hosted LLM proxy. If a manually started Gateway is already using the same port, stop it first. Set AIEVOBOX_GATEWAY_AUTOSTART=0 only when the external Gateway already has the correct route and storage configuration.

📚 Documentation

Guides

GuideWhat it covers
Custom EnvironmentsHow to onboard a new external runtime adapter.
EvaluationRule evaluator discovery, interfaces, reward writing, and Geo3K evaluation.
RL TrainingBuffer Server, Slime generator, Geo3K training path, and key variables.
Data ManagerSQLite storage behavior, table schema, row types, and query examples.
S3 + LanceDB StorageHow to switch trajectory storage from local SQLite to S3 + LanceDB.

Internal

Internal DocWhat it covers
RJob ModeRemote RJob runtime config, authentication, mounts, Gateway reachability, and Geo3K examples.
Sandbox ModeBrainbox Sandbox Environment config, volumes, lifecycle, and launch flow.

Reference

ReferenceWhat it covers
CLI And ConfigurationLauncher flags, Gateway config, agent config, start config, RJob, and Sandbox settings.
Supported EnvironmentsChecked-in adapters, the standard Geo3K path, and the runtime matrix.
Gateway ReferenceOpenAI-compatible routes, session endpoints, telemetry, request logs, and storage consistency.
ReportSAfactory report.

📖 Citation

If SAfactory or SAfactory-generated datasets are useful in your work, please cite this repository and the specific dataset or report you used.

@misc{chen2026safactoryscalableagenticinfrastructure,
title={SAfactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence},
author={Shanghai AI Lab},
year={2026},
eprint={2605.06230},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2605.06230},
}

About

SAfactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence

Resources

Stars

197 stars

Watchers

13 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SAfactory

中文 | English

SAfactory is a scalable infrastructure for agent evaluation, trajectory collection, and reinforcement learning training. It schedules agents and benchmark environments as external runtimes, routes model calls through a session-aware OpenAI-compatible Gateway, records trajectories, and feeds completed rollouts into training systems such as Slime.

Why SAfactoryDemoAgent SkillQuick StartDocumentationCitation

PythonLicenseExecutionLLM


✨ Why SAfactory

SAfactory provides one unified workflow for agent onboarding, benchmark onboarding, evaluation, rollout data generation, and RL training.

The core runtime contract is:

  • one dataset row becomes one scheduled episode;
  • each episode gets its own session_id and Gateway session;
  • the runtime calls the target model through the Gateway;
  • the runtime returns one JSON result;
  • rule_evaluator.py converts runtime output and trajectory data into the SAfactory format;
  • completed trainable trajectories can be consumed by the RL Buffer Server or persisted as reusable data assets.

🎬 Demo

demo.1.mp4

Click to watch the full demo

🧩 Agent Skill Quick Start

This repository includes a lightweight Agent skill that helps agents use SAfactory through the standard workflows:

skills/safactory-workflows/SKILL.md

It covers three common requests:

  • onboard a new benchmark or custom environment into SAfactory;
  • run Docker-mode evaluation for a selected environment;
  • start GRPO / RL training for a selected environment.

When working with an Agent, use prompts such as:

Use skills/safactory-workflows to help me onboard this benchmark into SAfactory.
Use the safactory-workflows skill to run geo3k evaluation in Docker mode.
Use the safactory-workflows skill to start GRPO training for my_env.

The skill does not replace the docs. It guides the Agent to read docs/guides/, docs/reference/, and the root README as needed, while using the standard env/geo3k/ environment as the reference implementation. If your Agent supports local skill discovery, add skills/safactory-workflows/ to its skill search path; otherwise mention this path explicitly in the request.

🚀 Quick Start

1. SAfactory Installation And Gateway Configuration

Before running Geo3K, prepare the runtime image and dataset as described in Standard Environment: Geo3K.

Install SAfactory:

git clone https://github.com/AI45Lab/SAfactory.git
cd SAfactory
pip install -U -r requirements.txt

Docker must be available locally, and the current user must be able to run docker build, docker run, and docker exec.

Create a local Gateway config:

cp gateway/config.example.yaml gateway/config.local.yaml

Edit gateway/config.local.yaml and explicitly set the storage path and model route:

listen_host: 0.0.0.0listen_port: 8000base_session_path: /v1/sessionsmax_steps: -1storage_type: sqlitestorage_config:
db_url: sqlite://env_trajs.dbllm_routes:
geo3k_model:
base_url: http://YOUR_LLM_HOST/v1api_key: YOUR_API_KEYsupports_stream: truemax_concurrency: 64

Start the Gateway:

python -m gateway --config gateway/config.local.yaml

Check readiness from another terminal:

curl http://127.0.0.1:8000/readyz

When using Cloud storage, the recommended configuration is to omit landing_table and env_config_table from the Gateway YAML and set only one profile in .env:

WT_SDK_PROFILE=test

or:

WT_SDK_PROFILE=production

The test profile selects landing_test and env_config_test; the production/prod profile selects wind_tunnel_landing and evaluation_env_config. This keeps trajectory and environment-config writes in the same environment. Explicit table names in storage_config override the profile and should be used only for intentional special cases. SAfactory does not access the serving table.

2. Minimal Evaluation With Geo3K

Run a minimal Geo3K evaluation in Docker mode:

python launcher.py \
--mode docker \
--agent-config env/geo3k/geo3k_config.yaml \
--agent-start-config env/geo3k/geo3k_start.yaml \
--gateway-base-url http://127.0.0.1:8000/v1/sessions \
--llm-model geo3k_model \
--enable-evaluation \
--db-path sqlite://env_trajs.db \
--job-id geo3k-docker-smoke \
--pool-size 1 \
--max-workers 1 \
--max-steps 10

--llm-model must match a key under llm_routes. With --enable-evaluation, SAfactory calls env/geo3k/rule_evaluator.py and writes the final reward.

3. Minimal Training With Geo3K

Geo3K training uses the RL bridge under rl/ and the example config at rl/examples/geo3k_vl/env.sh.

Before starting, edit rl/examples/geo3k_vl/env.sh, or export the same variables in each terminal. Set the required local paths and scale the first run down:

export AIEVOBOX_ROOT=$(pwd)export AIEVOBOX_MODE=docker
export AIEVOBOX_AGENT_CONFIG=${AIEVOBOX_ROOT}/env/geo3k/geo3k_config.yaml
export AIEVOBOX_AGENT_START_CONFIG=${AIEVOBOX_ROOT}/env/geo3k/geo3k_start.yaml
export AIEVOBOX_GATEWAY_HOST=127.0.0.1
export AIEVOBOX_GATEWAY_PORT=8000
export RL_MODEL=geo3k_model
export AIEVOBOX_POOL_SIZE=2
export RL_GROUP_SIZE=2
export RL_EPOCH=1
export HF_CKPT_DIR=/path/to/hf-checkpoint
export SLIME_HOME=/path/to/slime
export MEGATRON_HOME=/path/to/Megatron-LM

Then start two processes from the repository root:

RL_ENV_SH=rl/examples/geo3k_vl/env.sh bash rl/run_slime_generator.sh
RL_ENV_SH=rl/examples/geo3k_vl/env.sh bash rl/run_buffer_server.sh

Rewards for RL training come from completed trajectories in the database: rule_evaluator.py writes the reward after evaluation, and Buffer Server fetches trajectory rows with rewards from the same database used by Launcher / Gateway before serving them to the training process.

Buffer Server can automatically start a Gateway and route RL_MODEL to the Slime-hosted LLM proxy. If a manually started Gateway is already using the same port, stop it first. Set AIEVOBOX_GATEWAY_AUTOSTART=0 only when the external Gateway already has the correct route and storage configuration.

📚 Documentation

Guides

GuideWhat it covers
Custom EnvironmentsHow to onboard a new external runtime adapter.
EvaluationRule evaluator discovery, interfaces, reward writing, and Geo3K evaluation.
RL TrainingBuffer Server, Slime generator, Geo3K training path, and key variables.
Data ManagerSQLite storage behavior, table schema, row types, and query examples.
S3 + LanceDB StorageHow to switch trajectory storage from local SQLite to S3 + LanceDB.

Internal

Internal DocWhat it covers
RJob ModeRemote RJob runtime config, authentication, mounts, Gateway reachability, and Geo3K examples.
Sandbox ModeBrainbox Sandbox Environment config, volumes, lifecycle, and launch flow.

Reference

ReferenceWhat it covers
CLI And ConfigurationLauncher flags, Gateway config, agent config, start config, RJob, and Sandbox settings.
Supported EnvironmentsChecked-in adapters, the standard Geo3K path, and the runtime matrix.
Gateway ReferenceOpenAI-compatible routes, session endpoints, telemetry, request logs, and storage consistency.
ReportSAfactory report.

📖 Citation

If SAfactory or SAfactory-generated datasets are useful in your work, please cite this repository and the specific dataset or report you used.

@misc{chen2026safactoryscalableagenticinfrastructure,
title={SAfactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence},
author={Shanghai AI Lab},
year={2026},
eprint={2605.06230},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2605.06230},
}

About

SAfactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence

Resources

Stars

197 stars

Watchers

13 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SAfactory

中文 | English

SAfactory is a scalable infrastructure for agent evaluation, trajectory collection, and reinforcement learning training. It schedules agents and benchmark environments as external runtimes, routes model calls through a session-aware OpenAI-compatible Gateway, records trajectories, and feeds completed rollouts into training systems such as Slime.

Why SAfactoryDemoAgent SkillQuick StartDocumentationCitation

PythonLicenseExecutionLLM


✨ Why SAfactory

SAfactory provides one unified workflow for agent onboarding, benchmark onboarding, evaluation, rollout data generation, and RL training.

The core runtime contract is:

  • one dataset row becomes one scheduled episode;
  • each episode gets its own session_id and Gateway session;
  • the runtime calls the target model through the Gateway;
  • the runtime returns one JSON result;
  • rule_evaluator.py converts runtime output and trajectory data into the SAfactory format;
  • completed trainable trajectories can be consumed by the RL Buffer Server or persisted as reusable data assets.

🎬 Demo

demo.1.mp4

Click to watch the full demo

🧩 Agent Skill Quick Start

This repository includes a lightweight Agent skill that helps agents use SAfactory through the standard workflows:

skills/safactory-workflows/SKILL.md

It covers three common requests:

  • onboard a new benchmark or custom environment into SAfactory;
  • run Docker-mode evaluation for a selected environment;
  • start GRPO / RL training for a selected environment.

When working with an Agent, use prompts such as:

Use skills/safactory-workflows to help me onboard this benchmark into SAfactory.
Use the safactory-workflows skill to run geo3k evaluation in Docker mode.
Use the safactory-workflows skill to start GRPO training for my_env.

The skill does not replace the docs. It guides the Agent to read docs/guides/, docs/reference/, and the root README as needed, while using the standard env/geo3k/ environment as the reference implementation. If your Agent supports local skill discovery, add skills/safactory-workflows/ to its skill search path; otherwise mention this path explicitly in the request.

🚀 Quick Start

1. SAfactory Installation And Gateway Configuration

Before running Geo3K, prepare the runtime image and dataset as described in Standard Environment: Geo3K.

Install SAfactory:

git clone https://github.com/AI45Lab/SAfactory.git
cd SAfactory
pip install -U -r requirements.txt

Docker must be available locally, and the current user must be able to run docker build, docker run, and docker exec.

Create a local Gateway config:

cp gateway/config.example.yaml gateway/config.local.yaml

Edit gateway/config.local.yaml and explicitly set the storage path and model route:

listen_host: 0.0.0.0listen_port: 8000base_session_path: /v1/sessionsmax_steps: -1storage_type: sqlitestorage_config:
db_url: sqlite://env_trajs.dbllm_routes:
geo3k_model:
base_url: http://YOUR_LLM_HOST/v1api_key: YOUR_API_KEYsupports_stream: truemax_concurrency: 64

Start the Gateway:

python -m gateway --config gateway/config.local.yaml

Check readiness from another terminal:

curl http://127.0.0.1:8000/readyz

When using Cloud storage, the recommended configuration is to omit landing_table and env_config_table from the Gateway YAML and set only one profile in .env:

WT_SDK_PROFILE=test

or:

WT_SDK_PROFILE=production

The test profile selects landing_test and env_config_test; the production/prod profile selects wind_tunnel_landing and evaluation_env_config. This keeps trajectory and environment-config writes in the same environment. Explicit table names in storage_config override the profile and should be used only for intentional special cases. SAfactory does not access the serving table.

2. Minimal Evaluation With Geo3K

Run a minimal Geo3K evaluation in Docker mode:

python launcher.py \
--mode docker \
--agent-config env/geo3k/geo3k_config.yaml \
--agent-start-config env/geo3k/geo3k_start.yaml \
--gateway-base-url http://127.0.0.1:8000/v1/sessions \
--llm-model geo3k_model \
--enable-evaluation \
--db-path sqlite://env_trajs.db \
--job-id geo3k-docker-smoke \
--pool-size 1 \
--max-workers 1 \
--max-steps 10

--llm-model must match a key under llm_routes. With --enable-evaluation, SAfactory calls env/geo3k/rule_evaluator.py and writes the final reward.

3. Minimal Training With Geo3K

Geo3K training uses the RL bridge under rl/ and the example config at rl/examples/geo3k_vl/env.sh.

Before starting, edit rl/examples/geo3k_vl/env.sh, or export the same variables in each terminal. Set the required local paths and scale the first run down:

export AIEVOBOX_ROOT=$(pwd)export AIEVOBOX_MODE=docker
export AIEVOBOX_AGENT_CONFIG=${AIEVOBOX_ROOT}/env/geo3k/geo3k_config.yaml
export AIEVOBOX_AGENT_START_CONFIG=${AIEVOBOX_ROOT}/env/geo3k/geo3k_start.yaml
export AIEVOBOX_GATEWAY_HOST=127.0.0.1
export AIEVOBOX_GATEWAY_PORT=8000
export RL_MODEL=geo3k_model
export AIEVOBOX_POOL_SIZE=2
export RL_GROUP_SIZE=2
export RL_EPOCH=1
export HF_CKPT_DIR=/path/to/hf-checkpoint
export SLIME_HOME=/path/to/slime
export MEGATRON_HOME=/path/to/Megatron-LM

Then start two processes from the repository root:

RL_ENV_SH=rl/examples/geo3k_vl/env.sh bash rl/run_slime_generator.sh
RL_ENV_SH=rl/examples/geo3k_vl/env.sh bash rl/run_buffer_server.sh

Rewards for RL training come from completed trajectories in the database: rule_evaluator.py writes the reward after evaluation, and Buffer Server fetches trajectory rows with rewards from the same database used by Launcher / Gateway before serving them to the training process.

Buffer Server can automatically start a Gateway and route RL_MODEL to the Slime-hosted LLM proxy. If a manually started Gateway is already using the same port, stop it first. Set AIEVOBOX_GATEWAY_AUTOSTART=0 only when the external Gateway already has the correct route and storage configuration.

📚 Documentation

Guides

GuideWhat it covers
Custom EnvironmentsHow to onboard a new external runtime adapter.
EvaluationRule evaluator discovery, interfaces, reward writing, and Geo3K evaluation.
RL TrainingBuffer Server, Slime generator, Geo3K training path, and key variables.
Data ManagerSQLite storage behavior, table schema, row types, and query examples.
S3 + LanceDB StorageHow to switch trajectory storage from local SQLite to S3 + LanceDB.

Internal

Internal DocWhat it covers
RJob ModeRemote RJob runtime config, authentication, mounts, Gateway reachability, and Geo3K examples.
Sandbox ModeBrainbox Sandbox Environment config, volumes, lifecycle, and launch flow.

Reference

ReferenceWhat it covers
CLI And ConfigurationLauncher flags, Gateway config, agent config, start config, RJob, and Sandbox settings.
Supported EnvironmentsChecked-in adapters, the standard Geo3K path, and the runtime matrix.
Gateway ReferenceOpenAI-compatible routes, session endpoints, telemetry, request logs, and storage consistency.
ReportSAfactory report.

📖 Citation

If SAfactory or SAfactory-generated datasets are useful in your work, please cite this repository and the specific dataset or report you used.

@misc{chen2026safactoryscalableagenticinfrastructure,
title={SAfactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence},
author={Shanghai AI Lab},
year={2026},
eprint={2605.06230},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2605.06230},
}

About

SAfactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence

Resources

Stars

197 stars

Watchers

13 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

SAfactory

中文 | English

SAfactory is a scalable infrastructure for agent evaluation, trajectory collection, and reinforcement learning training. It schedules agents and benchmark environments as external runtimes, routes model calls through a session-aware OpenAI-compatible Gateway, records trajectories, and feeds completed rollouts into training systems such as Slime.

Why SAfactoryDemoAgent SkillQuick StartDocumentationCitation

PythonLicenseExecutionLLM


✨ Why SAfactory

SAfactory provides one unified workflow for agent onboarding, benchmark onboarding, evaluation, rollout data generation, and RL training.

The core runtime contract is:

  • one dataset row becomes one scheduled episode;
  • each episode gets its own session_id and Gateway session;
  • the runtime calls the target model through the Gateway;
  • the runtime returns one JSON result;
  • rule_evaluator.py converts runtime output and trajectory data into the SAfactory format;
  • completed trainable trajectories can be consumed by the RL Buffer Server or persisted as reusable data assets.

🎬 Demo

demo.1.mp4

Click to watch the full demo

🧩 Agent Skill Quick Start

This repository includes a lightweight Agent skill that helps agents use SAfactory through the standard workflows:

skills/safactory-workflows/SKILL.md

It covers three common requests:

  • onboard a new benchmark or custom environment into SAfactory;
  • run Docker-mode evaluation for a selected environment;
  • start GRPO / RL training for a selected environment.

When working with an Agent, use prompts such as:

Use skills/safactory-workflows to help me onboard this benchmark into SAfactory.
Use the safactory-workflows skill to run geo3k evaluation in Docker mode.
Use the safactory-workflows skill to start GRPO training for my_env.

The skill does not replace the docs. It guides the Agent to read docs/guides/, docs/reference/, and the root README as needed, while using the standard env/geo3k/ environment as the reference implementation. If your Agent supports local skill discovery, add skills/safactory-workflows/ to its skill search path; otherwise mention this path explicitly in the request.

🚀 Quick Start

1. SAfactory Installation And Gateway Configuration

Before running Geo3K, prepare the runtime image and dataset as described in Standard Environment: Geo3K.

Install SAfactory:

git clone https://github.com/AI45Lab/SAfactory.git
cd SAfactory
pip install -U -r requirements.txt

Docker must be available locally, and the current user must be able to run docker build, docker run, and docker exec.

Create a local Gateway config:

cp gateway/config.example.yaml gateway/config.local.yaml

Edit gateway/config.local.yaml and explicitly set the storage path and model route:

listen_host: 0.0.0.0listen_port: 8000base_session_path: /v1/sessionsmax_steps: -1storage_type: sqlitestorage_config:
db_url: sqlite://env_trajs.dbllm_routes:
geo3k_model:
base_url: http://YOUR_LLM_HOST/v1api_key: YOUR_API_KEYsupports_stream: truemax_concurrency: 64

Start the Gateway:

python -m gateway --config gateway/config.local.yaml

Check readiness from another terminal:

curl http://127.0.0.1:8000/readyz

When using Cloud storage, the recommended configuration is to omit landing_table and env_config_table from the Gateway YAML and set only one profile in .env:

WT_SDK_PROFILE=test

or:

WT_SDK_PROFILE=production

The test profile selects landing_test and env_config_test; the production/prod profile selects wind_tunnel_landing and evaluation_env_config. This keeps trajectory and environment-config writes in the same environment. Explicit table names in storage_config override the profile and should be used only for intentional special cases. SAfactory does not access the serving table.

2. Minimal Evaluation With Geo3K

Run a minimal Geo3K evaluation in Docker mode:

python launcher.py \
--mode docker \
--agent-config env/geo3k/geo3k_config.yaml \
--agent-start-config env/geo3k/geo3k_start.yaml \
--gateway-base-url http://127.0.0.1:8000/v1/sessions \
--llm-model geo3k_model \
--enable-evaluation \
--db-path sqlite://env_trajs.db \
--job-id geo3k-docker-smoke \
--pool-size 1 \
--max-workers 1 \
--max-steps 10

--llm-model must match a key under llm_routes. With --enable-evaluation, SAfactory calls env/geo3k/rule_evaluator.py and writes the final reward.

3. Minimal Training With Geo3K

Geo3K training uses the RL bridge under rl/ and the example config at rl/examples/geo3k_vl/env.sh.

Before starting, edit rl/examples/geo3k_vl/env.sh, or export the same variables in each terminal. Set the required local paths and scale the first run down:

export AIEVOBOX_ROOT=$(pwd)export AIEVOBOX_MODE=docker
export AIEVOBOX_AGENT_CONFIG=${AIEVOBOX_ROOT}/env/geo3k/geo3k_config.yaml
export AIEVOBOX_AGENT_START_CONFIG=${AIEVOBOX_ROOT}/env/geo3k/geo3k_start.yaml
export AIEVOBOX_GATEWAY_HOST=127.0.0.1
export AIEVOBOX_GATEWAY_PORT=8000
export RL_MODEL=geo3k_model
export AIEVOBOX_POOL_SIZE=2
export RL_GROUP_SIZE=2
export RL_EPOCH=1
export HF_CKPT_DIR=/path/to/hf-checkpoint
export SLIME_HOME=/path/to/slime
export MEGATRON_HOME=/path/to/Megatron-LM

Then start two processes from the repository root:

RL_ENV_SH=rl/examples/geo3k_vl/env.sh bash rl/run_slime_generator.sh
RL_ENV_SH=rl/examples/geo3k_vl/env.sh bash rl/run_buffer_server.sh

Rewards for RL training come from completed trajectories in the database: rule_evaluator.py writes the reward after evaluation, and Buffer Server fetches trajectory rows with rewards from the same database used by Launcher / Gateway before serving them to the training process.

Buffer Server can automatically start a Gateway and route RL_MODEL to the Slime-hosted LLM proxy. If a manually started Gateway is already using the same port, stop it first. Set AIEVOBOX_GATEWAY_AUTOSTART=0 only when the external Gateway already has the correct route and storage configuration.

📚 Documentation

Guides

GuideWhat it covers
Custom EnvironmentsHow to onboard a new external runtime adapter.
EvaluationRule evaluator discovery, interfaces, reward writing, and Geo3K evaluation.
RL TrainingBuffer Server, Slime generator, Geo3K training path, and key variables.
Data ManagerSQLite storage behavior, table schema, row types, and query examples.
S3 + LanceDB StorageHow to switch trajectory storage from local SQLite to S3 + LanceDB.

Internal

Internal DocWhat it covers
RJob ModeRemote RJob runtime config, authentication, mounts, Gateway reachability, and Geo3K examples.
Sandbox ModeBrainbox Sandbox Environment config, volumes, lifecycle, and launch flow.

Reference

ReferenceWhat it covers
CLI And ConfigurationLauncher flags, Gateway config, agent config, start config, RJob, and Sandbox settings.
Supported EnvironmentsChecked-in adapters, the standard Geo3K path, and the runtime matrix.
Gateway ReferenceOpenAI-compatible routes, session endpoints, telemetry, request logs, and storage consistency.
ReportSAfactory report.

📖 Citation

If SAfactory or SAfactory-generated datasets are useful in your work, please cite this repository and the specific dataset or report you used.

@misc{chen2026safactoryscalableagenticinfrastructure,
title={SAfactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence},
author={Shanghai AI Lab},
year={2026},
eprint={2605.06230},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2605.06230},
}

About

SAfactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence

Resources

Stars

197 stars

Watchers

13 watching

Forks

Releases

Packages

Contributors

Languages