Repository files navigation

RPent-logo
Hugging Face

English简体中文

RPent: Agentic Infrastructure for the Physical World

RPent (Recursive Physical Agent) is an open framework for building embodied agents that continuously evolve through recursive interaction with the physical world. Rather than prescribing a single foundation model, RPent provides a recursive agent framework that harnesses heterogeneous intelligence, including perception, reasoning, memory, execution, and self-evolution, into a unified physical agent. Through continuous interaction, reflection, and adaptation, RPent enables physical agents to acquire new capabilities and evolve beyond their initial design. We build RPent upon a foundation of service-oriented, standardized, and composable design principles, ensuring the framework remains highly extensible.

RPent framework

Who Should Consider Using RPent?

RPent is built for four kinds of users:

  • Embodied intelligence researchers targeting high success rates on embodied tasks and benchmarks — especially long-horizon manipulation. RPent's memory-guided, agentic composition consistently lifts task success beyond what a frozen VLA delivers alone.
  • Online-learning and reinforcement-learning researchers studying self-evolving embodied agents. RPent's recursive interaction, reflection, and memory-distillation loops provide a ready substrate for continual and reinforcement learning in the physical world.
  • Robotics application developers deploying embodied solutions on real robot hardware. RPent's service-oriented, standardized architecture and customizable agentic control logic improve real-world success rates and shorten the path from prototype to production.
  • End users of deployed embodied agents — the customers of the developers above. Install RPent together with the relevant real-robot extensions and run the predefined tasks out of the box, with no ML expertise required.

What's NEW!

Feature Matrix

Agentic PlannerAction PrimitiveSimulatorReal World
  • Franka
  • SO-101

Quick Start

1. Install RPent with a single pip install.

git clone https://github.com/RLinf/RPent rpent &&cd rpent
pip install -e ".[full]"

.[full] is the default end-to-end stack (openpi Pi0.5 VLA + LIBERO-PRO and RoboCasa365 simulators + SAM 3.0 on the RLinf runtime). If you don't need the whole stack, see the installation docs for narrower extras.

2. Download the LIBERO-PRO simulator assets.

liberopro-download-assets --skip-existing

💡 Slow connection to Hugging Face? Download through the mirror: HF_ENDPOINT=https://hf-mirror.com liberopro-download-assets --skip-existing.

See the installation docs for other simulators.

3. Configure keys and checkpoints, then run.

# Anthropic key; no need to export the base url if you use the official endpoint.export ANTHROPIC_BASE_URL=https://xxx
export ANTHROPIC_API_KEY=sk-xxx
# VLA checkpoint — download from# https://huggingface.co/RLinf/RLinf-Pi05-LIBERO-130-fullshot-SFT
hf download RLinf/RLinf-Pi05-LIBERO-130-fullshot-SFT \
--exclude optimizer.pt \
--local-dir ./checkpoints/RLinf-Pi05-LIBERO-130-fullshot-SFT
export PI05_CHECKPOINT_PATH=$PWD/checkpoints/RLinf-Pi05-LIBERO-130-fullshot-SFT
# SAM 3.0 checkpoint — download from# https://modelscope.cn/models/facebook/sam3
pip install -U modelscope
modelscope download facebook/sam3 \
--local-dir ./checkpoints/sam3
export SAM3_CHECKPOINT_PATH=$PWD/checkpoints/sam3/sam3.pt
export LIBERO_TYPE=pro
# Run one task: libero_object_swap, task 2, seed 0, using Claude Code# with Claude Opus 4.8.
rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
--cuda-device 0 --planner claude_code --model claude-opus-4-8

See the planner docs to configure other planners (api, codex) and model providers. For the exploration workflow and local-memory evaluation, see LIBERO exploration mode.

Interactive CLI mode

Add --interactive (-i) to steer the agent live from your terminal. At the you> prompt, the built-in task is pre-filled — press Enter to use it or replace it with your own — then type any message while it runs to steer the agent at the next turn (/help lists commands; /quit or Ctrl-D ends). Requires an interactive terminal (TTY).

rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
--planner claude_code --model claude-opus-4-8 --interactive

Live Dashboard

Add --dashboard to start a local Dashboard and print its URL in the terminal. Open the URL and confirm the configuration; once the services are ready, start a task with /rpent-task <suite> <task> <seed>. The page streams agent reasoning, camera views, and the action timeline, and you can submit another task after the current one finishes. Use --dashboard-language zh-cn for the Chinese UI.

rpent --robot libero --dashboard --dashboard-language zh-cn \
--planner claude_code --model claude-opus-4-8

For a complete list of CLI options, see the Key CLI options table in the Quick Start docs. RoboCasa and RoboTwin use their own entrypoints and CLI — see the RoboCasa and RoboTwin docs.

For more detailed documentation, see the RPent documentation.

Citation and Acknowledgement

If you find RPent or Harness VLA helpful, please cite the paper:

@article{zhang2026harnessvla,
title={Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents},
author={Zhang, Yixian and Zhang, Huanming and Gao, Feng and Li, Xiao and Liu, Zhihao and Zhu, Chunyang and Qiu, Jiaxing and Yan, Yuchen and Liu, Jiyuan and Tang, Wenhao and Fang, Zhengru and Nie, Yi and Wei, Changxu and Wang, Yu and Ding, Wenbo and Yu, Chao},
journal={arXiv preprint arXiv:2607.08448},
year={2026},
url={https://arxiv.org/abs/2607.08448}
}

RPent builds on the simulators, VLA models, and training infrastructure of RLinf, and on the agent SDKs of the broader open-source community — pydantic-ai, the Claude Agent SDK, and the OpenAI Codex SDK. Thanks to the teams behind LIBERO, RoboCasa, robosuite, MuJoCo, and openpi.

About

RPent: Agentic Infrastructure for the Physical World

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

RPent-logo
Hugging Face

English简体中文

RPent: Agentic Infrastructure for the Physical World

RPent (Recursive Physical Agent) is an open framework for building embodied agents that continuously evolve through recursive interaction with the physical world. Rather than prescribing a single foundation model, RPent provides a recursive agent framework that harnesses heterogeneous intelligence, including perception, reasoning, memory, execution, and self-evolution, into a unified physical agent. Through continuous interaction, reflection, and adaptation, RPent enables physical agents to acquire new capabilities and evolve beyond their initial design. We build RPent upon a foundation of service-oriented, standardized, and composable design principles, ensuring the framework remains highly extensible.

RPent framework

Who Should Consider Using RPent?

RPent is built for four kinds of users:

  • Embodied intelligence researchers targeting high success rates on embodied tasks and benchmarks — especially long-horizon manipulation. RPent's memory-guided, agentic composition consistently lifts task success beyond what a frozen VLA delivers alone.
  • Online-learning and reinforcement-learning researchers studying self-evolving embodied agents. RPent's recursive interaction, reflection, and memory-distillation loops provide a ready substrate for continual and reinforcement learning in the physical world.
  • Robotics application developers deploying embodied solutions on real robot hardware. RPent's service-oriented, standardized architecture and customizable agentic control logic improve real-world success rates and shorten the path from prototype to production.
  • End users of deployed embodied agents — the customers of the developers above. Install RPent together with the relevant real-robot extensions and run the predefined tasks out of the box, with no ML expertise required.

What's NEW!

Feature Matrix

Agentic PlannerAction PrimitiveSimulatorReal World
  • Franka
  • SO-101

Quick Start

1. Install RPent with a single pip install.

git clone https://github.com/RLinf/RPent rpent &&cd rpent
pip install -e ".[full]"

.[full] is the default end-to-end stack (openpi Pi0.5 VLA + LIBERO-PRO and RoboCasa365 simulators + SAM 3.0 on the RLinf runtime). If you don't need the whole stack, see the installation docs for narrower extras.

2. Download the LIBERO-PRO simulator assets.

liberopro-download-assets --skip-existing

💡 Slow connection to Hugging Face? Download through the mirror: HF_ENDPOINT=https://hf-mirror.com liberopro-download-assets --skip-existing.

See the installation docs for other simulators.

3. Configure keys and checkpoints, then run.

# Anthropic key; no need to export the base url if you use the official endpoint.export ANTHROPIC_BASE_URL=https://xxx
export ANTHROPIC_API_KEY=sk-xxx
# VLA checkpoint — download from# https://huggingface.co/RLinf/RLinf-Pi05-LIBERO-130-fullshot-SFT
hf download RLinf/RLinf-Pi05-LIBERO-130-fullshot-SFT \
--exclude optimizer.pt \
--local-dir ./checkpoints/RLinf-Pi05-LIBERO-130-fullshot-SFT
export PI05_CHECKPOINT_PATH=$PWD/checkpoints/RLinf-Pi05-LIBERO-130-fullshot-SFT
# SAM 3.0 checkpoint — download from# https://modelscope.cn/models/facebook/sam3
pip install -U modelscope
modelscope download facebook/sam3 \
--local-dir ./checkpoints/sam3
export SAM3_CHECKPOINT_PATH=$PWD/checkpoints/sam3/sam3.pt
export LIBERO_TYPE=pro
# Run one task: libero_object_swap, task 2, seed 0, using Claude Code# with Claude Opus 4.8.
rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
--cuda-device 0 --planner claude_code --model claude-opus-4-8

See the planner docs to configure other planners (api, codex) and model providers. For the exploration workflow and local-memory evaluation, see LIBERO exploration mode.

Interactive CLI mode

Add --interactive (-i) to steer the agent live from your terminal. At the you> prompt, the built-in task is pre-filled — press Enter to use it or replace it with your own — then type any message while it runs to steer the agent at the next turn (/help lists commands; /quit or Ctrl-D ends). Requires an interactive terminal (TTY).

rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
--planner claude_code --model claude-opus-4-8 --interactive

Live Dashboard

Add --dashboard to start a local Dashboard and print its URL in the terminal. Open the URL and confirm the configuration; once the services are ready, start a task with /rpent-task <suite> <task> <seed>. The page streams agent reasoning, camera views, and the action timeline, and you can submit another task after the current one finishes. Use --dashboard-language zh-cn for the Chinese UI.

rpent --robot libero --dashboard --dashboard-language zh-cn \
--planner claude_code --model claude-opus-4-8

For a complete list of CLI options, see the Key CLI options table in the Quick Start docs. RoboCasa and RoboTwin use their own entrypoints and CLI — see the RoboCasa and RoboTwin docs.

For more detailed documentation, see the RPent documentation.

Citation and Acknowledgement

If you find RPent or Harness VLA helpful, please cite the paper:

@article{zhang2026harnessvla,
title={Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents},
author={Zhang, Yixian and Zhang, Huanming and Gao, Feng and Li, Xiao and Liu, Zhihao and Zhu, Chunyang and Qiu, Jiaxing and Yan, Yuchen and Liu, Jiyuan and Tang, Wenhao and Fang, Zhengru and Nie, Yi and Wei, Changxu and Wang, Yu and Ding, Wenbo and Yu, Chao},
journal={arXiv preprint arXiv:2607.08448},
year={2026},
url={https://arxiv.org/abs/2607.08448}
}

RPent builds on the simulators, VLA models, and training infrastructure of RLinf, and on the agent SDKs of the broader open-source community — pydantic-ai, the Claude Agent SDK, and the OpenAI Codex SDK. Thanks to the teams behind LIBERO, RoboCasa, robosuite, MuJoCo, and openpi.

About

RPent: Agentic Infrastructure for the Physical World

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

RPent-logo
Hugging Face

English简体中文

RPent: Agentic Infrastructure for the Physical World

RPent (Recursive Physical Agent) is an open framework for building embodied agents that continuously evolve through recursive interaction with the physical world. Rather than prescribing a single foundation model, RPent provides a recursive agent framework that harnesses heterogeneous intelligence, including perception, reasoning, memory, execution, and self-evolution, into a unified physical agent. Through continuous interaction, reflection, and adaptation, RPent enables physical agents to acquire new capabilities and evolve beyond their initial design. We build RPent upon a foundation of service-oriented, standardized, and composable design principles, ensuring the framework remains highly extensible.

RPent framework

Who Should Consider Using RPent?

RPent is built for four kinds of users:

  • Embodied intelligence researchers targeting high success rates on embodied tasks and benchmarks — especially long-horizon manipulation. RPent's memory-guided, agentic composition consistently lifts task success beyond what a frozen VLA delivers alone.
  • Online-learning and reinforcement-learning researchers studying self-evolving embodied agents. RPent's recursive interaction, reflection, and memory-distillation loops provide a ready substrate for continual and reinforcement learning in the physical world.
  • Robotics application developers deploying embodied solutions on real robot hardware. RPent's service-oriented, standardized architecture and customizable agentic control logic improve real-world success rates and shorten the path from prototype to production.
  • End users of deployed embodied agents — the customers of the developers above. Install RPent together with the relevant real-robot extensions and run the predefined tasks out of the box, with no ML expertise required.

What's NEW!

Feature Matrix

Agentic PlannerAction PrimitiveSimulatorReal World
  • Franka
  • SO-101

Quick Start

1. Install RPent with a single pip install.

git clone https://github.com/RLinf/RPent rpent &&cd rpent
pip install -e ".[full]"

.[full] is the default end-to-end stack (openpi Pi0.5 VLA + LIBERO-PRO and RoboCasa365 simulators + SAM 3.0 on the RLinf runtime). If you don't need the whole stack, see the installation docs for narrower extras.

2. Download the LIBERO-PRO simulator assets.

liberopro-download-assets --skip-existing

💡 Slow connection to Hugging Face? Download through the mirror: HF_ENDPOINT=https://hf-mirror.com liberopro-download-assets --skip-existing.

See the installation docs for other simulators.

3. Configure keys and checkpoints, then run.

# Anthropic key; no need to export the base url if you use the official endpoint.export ANTHROPIC_BASE_URL=https://xxx
export ANTHROPIC_API_KEY=sk-xxx
# VLA checkpoint — download from# https://huggingface.co/RLinf/RLinf-Pi05-LIBERO-130-fullshot-SFT
hf download RLinf/RLinf-Pi05-LIBERO-130-fullshot-SFT \
--exclude optimizer.pt \
--local-dir ./checkpoints/RLinf-Pi05-LIBERO-130-fullshot-SFT
export PI05_CHECKPOINT_PATH=$PWD/checkpoints/RLinf-Pi05-LIBERO-130-fullshot-SFT
# SAM 3.0 checkpoint — download from# https://modelscope.cn/models/facebook/sam3
pip install -U modelscope
modelscope download facebook/sam3 \
--local-dir ./checkpoints/sam3
export SAM3_CHECKPOINT_PATH=$PWD/checkpoints/sam3/sam3.pt
export LIBERO_TYPE=pro
# Run one task: libero_object_swap, task 2, seed 0, using Claude Code# with Claude Opus 4.8.
rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
--cuda-device 0 --planner claude_code --model claude-opus-4-8

See the planner docs to configure other planners (api, codex) and model providers. For the exploration workflow and local-memory evaluation, see LIBERO exploration mode.

Interactive CLI mode

Add --interactive (-i) to steer the agent live from your terminal. At the you> prompt, the built-in task is pre-filled — press Enter to use it or replace it with your own — then type any message while it runs to steer the agent at the next turn (/help lists commands; /quit or Ctrl-D ends). Requires an interactive terminal (TTY).

rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
--planner claude_code --model claude-opus-4-8 --interactive

Live Dashboard

Add --dashboard to start a local Dashboard and print its URL in the terminal. Open the URL and confirm the configuration; once the services are ready, start a task with /rpent-task <suite> <task> <seed>. The page streams agent reasoning, camera views, and the action timeline, and you can submit another task after the current one finishes. Use --dashboard-language zh-cn for the Chinese UI.

rpent --robot libero --dashboard --dashboard-language zh-cn \
--planner claude_code --model claude-opus-4-8

For a complete list of CLI options, see the Key CLI options table in the Quick Start docs. RoboCasa and RoboTwin use their own entrypoints and CLI — see the RoboCasa and RoboTwin docs.

For more detailed documentation, see the RPent documentation.

Citation and Acknowledgement

If you find RPent or Harness VLA helpful, please cite the paper:

@article{zhang2026harnessvla,
title={Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents},
author={Zhang, Yixian and Zhang, Huanming and Gao, Feng and Li, Xiao and Liu, Zhihao and Zhu, Chunyang and Qiu, Jiaxing and Yan, Yuchen and Liu, Jiyuan and Tang, Wenhao and Fang, Zhengru and Nie, Yi and Wei, Changxu and Wang, Yu and Ding, Wenbo and Yu, Chao},
journal={arXiv preprint arXiv:2607.08448},
year={2026},
url={https://arxiv.org/abs/2607.08448}
}

RPent builds on the simulators, VLA models, and training infrastructure of RLinf, and on the agent SDKs of the broader open-source community — pydantic-ai, the Claude Agent SDK, and the OpenAI Codex SDK. Thanks to the teams behind LIBERO, RoboCasa, robosuite, MuJoCo, and openpi.

About

RPent: Agentic Infrastructure for the Physical World

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

RPent-logo
Hugging Face

English简体中文

RPent: Agentic Infrastructure for the Physical World

RPent (Recursive Physical Agent) is an open framework for building embodied agents that continuously evolve through recursive interaction with the physical world. Rather than prescribing a single foundation model, RPent provides a recursive agent framework that harnesses heterogeneous intelligence, including perception, reasoning, memory, execution, and self-evolution, into a unified physical agent. Through continuous interaction, reflection, and adaptation, RPent enables physical agents to acquire new capabilities and evolve beyond their initial design. We build RPent upon a foundation of service-oriented, standardized, and composable design principles, ensuring the framework remains highly extensible.

RPent framework

Who Should Consider Using RPent?

RPent is built for four kinds of users:

  • Embodied intelligence researchers targeting high success rates on embodied tasks and benchmarks — especially long-horizon manipulation. RPent's memory-guided, agentic composition consistently lifts task success beyond what a frozen VLA delivers alone.
  • Online-learning and reinforcement-learning researchers studying self-evolving embodied agents. RPent's recursive interaction, reflection, and memory-distillation loops provide a ready substrate for continual and reinforcement learning in the physical world.
  • Robotics application developers deploying embodied solutions on real robot hardware. RPent's service-oriented, standardized architecture and customizable agentic control logic improve real-world success rates and shorten the path from prototype to production.
  • End users of deployed embodied agents — the customers of the developers above. Install RPent together with the relevant real-robot extensions and run the predefined tasks out of the box, with no ML expertise required.

What's NEW!

Feature Matrix

Agentic PlannerAction PrimitiveSimulatorReal World
  • Franka
  • SO-101

Quick Start

1. Install RPent with a single pip install.

git clone https://github.com/RLinf/RPent rpent &&cd rpent
pip install -e ".[full]"

.[full] is the default end-to-end stack (openpi Pi0.5 VLA + LIBERO-PRO and RoboCasa365 simulators + SAM 3.0 on the RLinf runtime). If you don't need the whole stack, see the installation docs for narrower extras.

2. Download the LIBERO-PRO simulator assets.

liberopro-download-assets --skip-existing

💡 Slow connection to Hugging Face? Download through the mirror: HF_ENDPOINT=https://hf-mirror.com liberopro-download-assets --skip-existing.

See the installation docs for other simulators.

3. Configure keys and checkpoints, then run.

# Anthropic key; no need to export the base url if you use the official endpoint.export ANTHROPIC_BASE_URL=https://xxx
export ANTHROPIC_API_KEY=sk-xxx
# VLA checkpoint — download from# https://huggingface.co/RLinf/RLinf-Pi05-LIBERO-130-fullshot-SFT
hf download RLinf/RLinf-Pi05-LIBERO-130-fullshot-SFT \
--exclude optimizer.pt \
--local-dir ./checkpoints/RLinf-Pi05-LIBERO-130-fullshot-SFT
export PI05_CHECKPOINT_PATH=$PWD/checkpoints/RLinf-Pi05-LIBERO-130-fullshot-SFT
# SAM 3.0 checkpoint — download from# https://modelscope.cn/models/facebook/sam3
pip install -U modelscope
modelscope download facebook/sam3 \
--local-dir ./checkpoints/sam3
export SAM3_CHECKPOINT_PATH=$PWD/checkpoints/sam3/sam3.pt
export LIBERO_TYPE=pro
# Run one task: libero_object_swap, task 2, seed 0, using Claude Code# with Claude Opus 4.8.
rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
--cuda-device 0 --planner claude_code --model claude-opus-4-8

See the planner docs to configure other planners (api, codex) and model providers. For the exploration workflow and local-memory evaluation, see LIBERO exploration mode.

Interactive CLI mode

Add --interactive (-i) to steer the agent live from your terminal. At the you> prompt, the built-in task is pre-filled — press Enter to use it or replace it with your own — then type any message while it runs to steer the agent at the next turn (/help lists commands; /quit or Ctrl-D ends). Requires an interactive terminal (TTY).

rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
--planner claude_code --model claude-opus-4-8 --interactive

Live Dashboard

Add --dashboard to start a local Dashboard and print its URL in the terminal. Open the URL and confirm the configuration; once the services are ready, start a task with /rpent-task <suite> <task> <seed>. The page streams agent reasoning, camera views, and the action timeline, and you can submit another task after the current one finishes. Use --dashboard-language zh-cn for the Chinese UI.

rpent --robot libero --dashboard --dashboard-language zh-cn \
--planner claude_code --model claude-opus-4-8

For a complete list of CLI options, see the Key CLI options table in the Quick Start docs. RoboCasa and RoboTwin use their own entrypoints and CLI — see the RoboCasa and RoboTwin docs.

For more detailed documentation, see the RPent documentation.

Citation and Acknowledgement

If you find RPent or Harness VLA helpful, please cite the paper:

@article{zhang2026harnessvla,
title={Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents},
author={Zhang, Yixian and Zhang, Huanming and Gao, Feng and Li, Xiao and Liu, Zhihao and Zhu, Chunyang and Qiu, Jiaxing and Yan, Yuchen and Liu, Jiyuan and Tang, Wenhao and Fang, Zhengru and Nie, Yi and Wei, Changxu and Wang, Yu and Ding, Wenbo and Yu, Chao},
journal={arXiv preprint arXiv:2607.08448},
year={2026},
url={https://arxiv.org/abs/2607.08448}
}

RPent builds on the simulators, VLA models, and training infrastructure of RLinf, and on the agent SDKs of the broader open-source community — pydantic-ai, the Claude Agent SDK, and the OpenAI Codex SDK. Thanks to the teams behind LIBERO, RoboCasa, robosuite, MuJoCo, and openpi.

About

RPent: Agentic Infrastructure for the Physical World

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

RPent-logo
Hugging Face

English简体中文

RPent: Agentic Infrastructure for the Physical World

RPent (Recursive Physical Agent) is an open framework for building embodied agents that continuously evolve through recursive interaction with the physical world. Rather than prescribing a single foundation model, RPent provides a recursive agent framework that harnesses heterogeneous intelligence, including perception, reasoning, memory, execution, and self-evolution, into a unified physical agent. Through continuous interaction, reflection, and adaptation, RPent enables physical agents to acquire new capabilities and evolve beyond their initial design. We build RPent upon a foundation of service-oriented, standardized, and composable design principles, ensuring the framework remains highly extensible.

RPent framework

Who Should Consider Using RPent?

RPent is built for four kinds of users:

  • Embodied intelligence researchers targeting high success rates on embodied tasks and benchmarks — especially long-horizon manipulation. RPent's memory-guided, agentic composition consistently lifts task success beyond what a frozen VLA delivers alone.
  • Online-learning and reinforcement-learning researchers studying self-evolving embodied agents. RPent's recursive interaction, reflection, and memory-distillation loops provide a ready substrate for continual and reinforcement learning in the physical world.
  • Robotics application developers deploying embodied solutions on real robot hardware. RPent's service-oriented, standardized architecture and customizable agentic control logic improve real-world success rates and shorten the path from prototype to production.
  • End users of deployed embodied agents — the customers of the developers above. Install RPent together with the relevant real-robot extensions and run the predefined tasks out of the box, with no ML expertise required.

What's NEW!

Feature Matrix

Agentic PlannerAction PrimitiveSimulatorReal World
  • Franka
  • SO-101

Quick Start

1. Install RPent with a single pip install.

git clone https://github.com/RLinf/RPent rpent &&cd rpent
pip install -e ".[full]"

.[full] is the default end-to-end stack (openpi Pi0.5 VLA + LIBERO-PRO and RoboCasa365 simulators + SAM 3.0 on the RLinf runtime). If you don't need the whole stack, see the installation docs for narrower extras.

2. Download the LIBERO-PRO simulator assets.

liberopro-download-assets --skip-existing

💡 Slow connection to Hugging Face? Download through the mirror: HF_ENDPOINT=https://hf-mirror.com liberopro-download-assets --skip-existing.

See the installation docs for other simulators.

3. Configure keys and checkpoints, then run.

# Anthropic key; no need to export the base url if you use the official endpoint.export ANTHROPIC_BASE_URL=https://xxx
export ANTHROPIC_API_KEY=sk-xxx
# VLA checkpoint — download from# https://huggingface.co/RLinf/RLinf-Pi05-LIBERO-130-fullshot-SFT
hf download RLinf/RLinf-Pi05-LIBERO-130-fullshot-SFT \
--exclude optimizer.pt \
--local-dir ./checkpoints/RLinf-Pi05-LIBERO-130-fullshot-SFT
export PI05_CHECKPOINT_PATH=$PWD/checkpoints/RLinf-Pi05-LIBERO-130-fullshot-SFT
# SAM 3.0 checkpoint — download from# https://modelscope.cn/models/facebook/sam3
pip install -U modelscope
modelscope download facebook/sam3 \
--local-dir ./checkpoints/sam3
export SAM3_CHECKPOINT_PATH=$PWD/checkpoints/sam3/sam3.pt
export LIBERO_TYPE=pro
# Run one task: libero_object_swap, task 2, seed 0, using Claude Code# with Claude Opus 4.8.
rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
--cuda-device 0 --planner claude_code --model claude-opus-4-8

See the planner docs to configure other planners (api, codex) and model providers. For the exploration workflow and local-memory evaluation, see LIBERO exploration mode.

Interactive CLI mode

Add --interactive (-i) to steer the agent live from your terminal. At the you> prompt, the built-in task is pre-filled — press Enter to use it or replace it with your own — then type any message while it runs to steer the agent at the next turn (/help lists commands; /quit or Ctrl-D ends). Requires an interactive terminal (TTY).

rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
--planner claude_code --model claude-opus-4-8 --interactive

Live Dashboard

Add --dashboard to start a local Dashboard and print its URL in the terminal. Open the URL and confirm the configuration; once the services are ready, start a task with /rpent-task <suite> <task> <seed>. The page streams agent reasoning, camera views, and the action timeline, and you can submit another task after the current one finishes. Use --dashboard-language zh-cn for the Chinese UI.

rpent --robot libero --dashboard --dashboard-language zh-cn \
--planner claude_code --model claude-opus-4-8

For a complete list of CLI options, see the Key CLI options table in the Quick Start docs. RoboCasa and RoboTwin use their own entrypoints and CLI — see the RoboCasa and RoboTwin docs.

For more detailed documentation, see the RPent documentation.

Citation and Acknowledgement

If you find RPent or Harness VLA helpful, please cite the paper:

@article{zhang2026harnessvla,
title={Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents},
author={Zhang, Yixian and Zhang, Huanming and Gao, Feng and Li, Xiao and Liu, Zhihao and Zhu, Chunyang and Qiu, Jiaxing and Yan, Yuchen and Liu, Jiyuan and Tang, Wenhao and Fang, Zhengru and Nie, Yi and Wei, Changxu and Wang, Yu and Ding, Wenbo and Yu, Chao},
journal={arXiv preprint arXiv:2607.08448},
year={2026},
url={https://arxiv.org/abs/2607.08448}
}

RPent builds on the simulators, VLA models, and training infrastructure of RLinf, and on the agent SDKs of the broader open-source community — pydantic-ai, the Claude Agent SDK, and the OpenAI Codex SDK. Thanks to the teams behind LIBERO, RoboCasa, robosuite, MuJoCo, and openpi.

About

RPent: Agentic Infrastructure for the Physical World

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

RPent-logo
Hugging Face

English简体中文

RPent: Agentic Infrastructure for the Physical World

RPent (Recursive Physical Agent) is an open framework for building embodied agents that continuously evolve through recursive interaction with the physical world. Rather than prescribing a single foundation model, RPent provides a recursive agent framework that harnesses heterogeneous intelligence, including perception, reasoning, memory, execution, and self-evolution, into a unified physical agent. Through continuous interaction, reflection, and adaptation, RPent enables physical agents to acquire new capabilities and evolve beyond their initial design. We build RPent upon a foundation of service-oriented, standardized, and composable design principles, ensuring the framework remains highly extensible.

RPent framework

Who Should Consider Using RPent?

RPent is built for four kinds of users:

  • Embodied intelligence researchers targeting high success rates on embodied tasks and benchmarks — especially long-horizon manipulation. RPent's memory-guided, agentic composition consistently lifts task success beyond what a frozen VLA delivers alone.
  • Online-learning and reinforcement-learning researchers studying self-evolving embodied agents. RPent's recursive interaction, reflection, and memory-distillation loops provide a ready substrate for continual and reinforcement learning in the physical world.
  • Robotics application developers deploying embodied solutions on real robot hardware. RPent's service-oriented, standardized architecture and customizable agentic control logic improve real-world success rates and shorten the path from prototype to production.
  • End users of deployed embodied agents — the customers of the developers above. Install RPent together with the relevant real-robot extensions and run the predefined tasks out of the box, with no ML expertise required.

What's NEW!

Feature Matrix

Agentic PlannerAction PrimitiveSimulatorReal World
  • Franka
  • SO-101

Quick Start

1. Install RPent with a single pip install.

git clone https://github.com/RLinf/RPent rpent &&cd rpent
pip install -e ".[full]"

.[full] is the default end-to-end stack (openpi Pi0.5 VLA + LIBERO-PRO and RoboCasa365 simulators + SAM 3.0 on the RLinf runtime). If you don't need the whole stack, see the installation docs for narrower extras.

2. Download the LIBERO-PRO simulator assets.

liberopro-download-assets --skip-existing

💡 Slow connection to Hugging Face? Download through the mirror: HF_ENDPOINT=https://hf-mirror.com liberopro-download-assets --skip-existing.

See the installation docs for other simulators.

3. Configure keys and checkpoints, then run.

# Anthropic key; no need to export the base url if you use the official endpoint.export ANTHROPIC_BASE_URL=https://xxx
export ANTHROPIC_API_KEY=sk-xxx
# VLA checkpoint — download from# https://huggingface.co/RLinf/RLinf-Pi05-LIBERO-130-fullshot-SFT
hf download RLinf/RLinf-Pi05-LIBERO-130-fullshot-SFT \
--exclude optimizer.pt \
--local-dir ./checkpoints/RLinf-Pi05-LIBERO-130-fullshot-SFT
export PI05_CHECKPOINT_PATH=$PWD/checkpoints/RLinf-Pi05-LIBERO-130-fullshot-SFT
# SAM 3.0 checkpoint — download from# https://modelscope.cn/models/facebook/sam3
pip install -U modelscope
modelscope download facebook/sam3 \
--local-dir ./checkpoints/sam3
export SAM3_CHECKPOINT_PATH=$PWD/checkpoints/sam3/sam3.pt
export LIBERO_TYPE=pro
# Run one task: libero_object_swap, task 2, seed 0, using Claude Code# with Claude Opus 4.8.
rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
--cuda-device 0 --planner claude_code --model claude-opus-4-8

See the planner docs to configure other planners (api, codex) and model providers. For the exploration workflow and local-memory evaluation, see LIBERO exploration mode.

Interactive CLI mode

Add --interactive (-i) to steer the agent live from your terminal. At the you> prompt, the built-in task is pre-filled — press Enter to use it or replace it with your own — then type any message while it runs to steer the agent at the next turn (/help lists commands; /quit or Ctrl-D ends). Requires an interactive terminal (TTY).

rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
--planner claude_code --model claude-opus-4-8 --interactive

Live Dashboard

Add --dashboard to start a local Dashboard and print its URL in the terminal. Open the URL and confirm the configuration; once the services are ready, start a task with /rpent-task <suite> <task> <seed>. The page streams agent reasoning, camera views, and the action timeline, and you can submit another task after the current one finishes. Use --dashboard-language zh-cn for the Chinese UI.

rpent --robot libero --dashboard --dashboard-language zh-cn \
--planner claude_code --model claude-opus-4-8

For a complete list of CLI options, see the Key CLI options table in the Quick Start docs. RoboCasa and RoboTwin use their own entrypoints and CLI — see the RoboCasa and RoboTwin docs.

For more detailed documentation, see the RPent documentation.

Citation and Acknowledgement

If you find RPent or Harness VLA helpful, please cite the paper:

@article{zhang2026harnessvla,
title={Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents},
author={Zhang, Yixian and Zhang, Huanming and Gao, Feng and Li, Xiao and Liu, Zhihao and Zhu, Chunyang and Qiu, Jiaxing and Yan, Yuchen and Liu, Jiyuan and Tang, Wenhao and Fang, Zhengru and Nie, Yi and Wei, Changxu and Wang, Yu and Ding, Wenbo and Yu, Chao},
journal={arXiv preprint arXiv:2607.08448},
year={2026},
url={https://arxiv.org/abs/2607.08448}
}

RPent builds on the simulators, VLA models, and training infrastructure of RLinf, and on the agent SDKs of the broader open-source community — pydantic-ai, the Claude Agent SDK, and the OpenAI Codex SDK. Thanks to the teams behind LIBERO, RoboCasa, robosuite, MuJoCo, and openpi.

About

RPent: Agentic Infrastructure for the Physical World

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

RPent-logo
Hugging Face

English简体中文

RPent: Agentic Infrastructure for the Physical World

RPent (Recursive Physical Agent) is an open framework for building embodied agents that continuously evolve through recursive interaction with the physical world. Rather than prescribing a single foundation model, RPent provides a recursive agent framework that harnesses heterogeneous intelligence, including perception, reasoning, memory, execution, and self-evolution, into a unified physical agent. Through continuous interaction, reflection, and adaptation, RPent enables physical agents to acquire new capabilities and evolve beyond their initial design. We build RPent upon a foundation of service-oriented, standardized, and composable design principles, ensuring the framework remains highly extensible.

RPent framework

Who Should Consider Using RPent?

RPent is built for four kinds of users:

  • Embodied intelligence researchers targeting high success rates on embodied tasks and benchmarks — especially long-horizon manipulation. RPent's memory-guided, agentic composition consistently lifts task success beyond what a frozen VLA delivers alone.
  • Online-learning and reinforcement-learning researchers studying self-evolving embodied agents. RPent's recursive interaction, reflection, and memory-distillation loops provide a ready substrate for continual and reinforcement learning in the physical world.
  • Robotics application developers deploying embodied solutions on real robot hardware. RPent's service-oriented, standardized architecture and customizable agentic control logic improve real-world success rates and shorten the path from prototype to production.
  • End users of deployed embodied agents — the customers of the developers above. Install RPent together with the relevant real-robot extensions and run the predefined tasks out of the box, with no ML expertise required.

What's NEW!

Feature Matrix

Agentic PlannerAction PrimitiveSimulatorReal World
  • Franka
  • SO-101

Quick Start

1. Install RPent with a single pip install.

git clone https://github.com/RLinf/RPent rpent &&cd rpent
pip install -e ".[full]"

.[full] is the default end-to-end stack (openpi Pi0.5 VLA + LIBERO-PRO and RoboCasa365 simulators + SAM 3.0 on the RLinf runtime). If you don't need the whole stack, see the installation docs for narrower extras.

2. Download the LIBERO-PRO simulator assets.

liberopro-download-assets --skip-existing

💡 Slow connection to Hugging Face? Download through the mirror: HF_ENDPOINT=https://hf-mirror.com liberopro-download-assets --skip-existing.

See the installation docs for other simulators.

3. Configure keys and checkpoints, then run.

# Anthropic key; no need to export the base url if you use the official endpoint.export ANTHROPIC_BASE_URL=https://xxx
export ANTHROPIC_API_KEY=sk-xxx
# VLA checkpoint — download from# https://huggingface.co/RLinf/RLinf-Pi05-LIBERO-130-fullshot-SFT
hf download RLinf/RLinf-Pi05-LIBERO-130-fullshot-SFT \
--exclude optimizer.pt \
--local-dir ./checkpoints/RLinf-Pi05-LIBERO-130-fullshot-SFT
export PI05_CHECKPOINT_PATH=$PWD/checkpoints/RLinf-Pi05-LIBERO-130-fullshot-SFT
# SAM 3.0 checkpoint — download from# https://modelscope.cn/models/facebook/sam3
pip install -U modelscope
modelscope download facebook/sam3 \
--local-dir ./checkpoints/sam3
export SAM3_CHECKPOINT_PATH=$PWD/checkpoints/sam3/sam3.pt
export LIBERO_TYPE=pro
# Run one task: libero_object_swap, task 2, seed 0, using Claude Code# with Claude Opus 4.8.
rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
--cuda-device 0 --planner claude_code --model claude-opus-4-8

See the planner docs to configure other planners (api, codex) and model providers. For the exploration workflow and local-memory evaluation, see LIBERO exploration mode.

Interactive CLI mode

Add --interactive (-i) to steer the agent live from your terminal. At the you> prompt, the built-in task is pre-filled — press Enter to use it or replace it with your own — then type any message while it runs to steer the agent at the next turn (/help lists commands; /quit or Ctrl-D ends). Requires an interactive terminal (TTY).

rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
--planner claude_code --model claude-opus-4-8 --interactive

Live Dashboard

Add --dashboard to start a local Dashboard and print its URL in the terminal. Open the URL and confirm the configuration; once the services are ready, start a task with /rpent-task <suite> <task> <seed>. The page streams agent reasoning, camera views, and the action timeline, and you can submit another task after the current one finishes. Use --dashboard-language zh-cn for the Chinese UI.

rpent --robot libero --dashboard --dashboard-language zh-cn \
--planner claude_code --model claude-opus-4-8

For a complete list of CLI options, see the Key CLI options table in the Quick Start docs. RoboCasa and RoboTwin use their own entrypoints and CLI — see the RoboCasa and RoboTwin docs.

For more detailed documentation, see the RPent documentation.

Citation and Acknowledgement

If you find RPent or Harness VLA helpful, please cite the paper:

@article{zhang2026harnessvla,
title={Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents},
author={Zhang, Yixian and Zhang, Huanming and Gao, Feng and Li, Xiao and Liu, Zhihao and Zhu, Chunyang and Qiu, Jiaxing and Yan, Yuchen and Liu, Jiyuan and Tang, Wenhao and Fang, Zhengru and Nie, Yi and Wei, Changxu and Wang, Yu and Ding, Wenbo and Yu, Chao},
journal={arXiv preprint arXiv:2607.08448},
year={2026},
url={https://arxiv.org/abs/2607.08448}
}

RPent builds on the simulators, VLA models, and training infrastructure of RLinf, and on the agent SDKs of the broader open-source community — pydantic-ai, the Claude Agent SDK, and the OpenAI Codex SDK. Thanks to the teams behind LIBERO, RoboCasa, robosuite, MuJoCo, and openpi.

About

RPent: Agentic Infrastructure for the Physical World

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

RPent-logo
Hugging Face

English简体中文

RPent: Agentic Infrastructure for the Physical World

RPent (Recursive Physical Agent) is an open framework for building embodied agents that continuously evolve through recursive interaction with the physical world. Rather than prescribing a single foundation model, RPent provides a recursive agent framework that harnesses heterogeneous intelligence, including perception, reasoning, memory, execution, and self-evolution, into a unified physical agent. Through continuous interaction, reflection, and adaptation, RPent enables physical agents to acquire new capabilities and evolve beyond their initial design. We build RPent upon a foundation of service-oriented, standardized, and composable design principles, ensuring the framework remains highly extensible.

RPent framework

Who Should Consider Using RPent?

RPent is built for four kinds of users:

  • Embodied intelligence researchers targeting high success rates on embodied tasks and benchmarks — especially long-horizon manipulation. RPent's memory-guided, agentic composition consistently lifts task success beyond what a frozen VLA delivers alone.
  • Online-learning and reinforcement-learning researchers studying self-evolving embodied agents. RPent's recursive interaction, reflection, and memory-distillation loops provide a ready substrate for continual and reinforcement learning in the physical world.
  • Robotics application developers deploying embodied solutions on real robot hardware. RPent's service-oriented, standardized architecture and customizable agentic control logic improve real-world success rates and shorten the path from prototype to production.
  • End users of deployed embodied agents — the customers of the developers above. Install RPent together with the relevant real-robot extensions and run the predefined tasks out of the box, with no ML expertise required.

What's NEW!

Feature Matrix

Agentic PlannerAction PrimitiveSimulatorReal World
  • Franka
  • SO-101

Quick Start

1. Install RPent with a single pip install.

git clone https://github.com/RLinf/RPent rpent &&cd rpent
pip install -e ".[full]"

.[full] is the default end-to-end stack (openpi Pi0.5 VLA + LIBERO-PRO and RoboCasa365 simulators + SAM 3.0 on the RLinf runtime). If you don't need the whole stack, see the installation docs for narrower extras.

2. Download the LIBERO-PRO simulator assets.

liberopro-download-assets --skip-existing

💡 Slow connection to Hugging Face? Download through the mirror: HF_ENDPOINT=https://hf-mirror.com liberopro-download-assets --skip-existing.

See the installation docs for other simulators.

3. Configure keys and checkpoints, then run.

# Anthropic key; no need to export the base url if you use the official endpoint.export ANTHROPIC_BASE_URL=https://xxx
export ANTHROPIC_API_KEY=sk-xxx
# VLA checkpoint — download from# https://huggingface.co/RLinf/RLinf-Pi05-LIBERO-130-fullshot-SFT
hf download RLinf/RLinf-Pi05-LIBERO-130-fullshot-SFT \
--exclude optimizer.pt \
--local-dir ./checkpoints/RLinf-Pi05-LIBERO-130-fullshot-SFT
export PI05_CHECKPOINT_PATH=$PWD/checkpoints/RLinf-Pi05-LIBERO-130-fullshot-SFT
# SAM 3.0 checkpoint — download from# https://modelscope.cn/models/facebook/sam3
pip install -U modelscope
modelscope download facebook/sam3 \
--local-dir ./checkpoints/sam3
export SAM3_CHECKPOINT_PATH=$PWD/checkpoints/sam3/sam3.pt
export LIBERO_TYPE=pro
# Run one task: libero_object_swap, task 2, seed 0, using Claude Code# with Claude Opus 4.8.
rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
--cuda-device 0 --planner claude_code --model claude-opus-4-8

See the planner docs to configure other planners (api, codex) and model providers. For the exploration workflow and local-memory evaluation, see LIBERO exploration mode.

Interactive CLI mode

Add --interactive (-i) to steer the agent live from your terminal. At the you> prompt, the built-in task is pre-filled — press Enter to use it or replace it with your own — then type any message while it runs to steer the agent at the next turn (/help lists commands; /quit or Ctrl-D ends). Requires an interactive terminal (TTY).

rpent --robot libero --suite libero_object_swap --task 2 --seed 0 \
--planner claude_code --model claude-opus-4-8 --interactive

Live Dashboard

Add --dashboard to start a local Dashboard and print its URL in the terminal. Open the URL and confirm the configuration; once the services are ready, start a task with /rpent-task <suite> <task> <seed>. The page streams agent reasoning, camera views, and the action timeline, and you can submit another task after the current one finishes. Use --dashboard-language zh-cn for the Chinese UI.

rpent --robot libero --dashboard --dashboard-language zh-cn \
--planner claude_code --model claude-opus-4-8

For a complete list of CLI options, see the Key CLI options table in the Quick Start docs. RoboCasa and RoboTwin use their own entrypoints and CLI — see the RoboCasa and RoboTwin docs.

For more detailed documentation, see the RPent documentation.

Citation and Acknowledgement

If you find RPent or Harness VLA helpful, please cite the paper:

@article{zhang2026harnessvla,
title={Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents},
author={Zhang, Yixian and Zhang, Huanming and Gao, Feng and Li, Xiao and Liu, Zhihao and Zhu, Chunyang and Qiu, Jiaxing and Yan, Yuchen and Liu, Jiyuan and Tang, Wenhao and Fang, Zhengru and Nie, Yi and Wei, Changxu and Wang, Yu and Ding, Wenbo and Yu, Chao},
journal={arXiv preprint arXiv:2607.08448},
year={2026},
url={https://arxiv.org/abs/2607.08448}
}

RPent builds on the simulators, VLA models, and training infrastructure of RLinf, and on the agent SDKs of the broader open-source community — pydantic-ai, the Claude Agent SDK, and the OpenAI Codex SDK. Thanks to the teams behind LIBERO, RoboCasa, robosuite, MuJoCo, and openpi.

About

RPent: Agentic Infrastructure for the Physical World

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages