Skip to content

Latest commit

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

NeoRL2

LicenseLicense

The NEORL2 repository is an extension of the offline reinforcement learning benchmark NeoRL. The NEORL2 repository contains datasets for training and corresponding environments for testing the trained policies. The current datasets are collected from seven open-source environments: Pipeline, Simglucose, RocketRecovery, RandomFrictionHopper, DMSD, Fusion and SafetyHalfCheetah tasks. We perform online training using reinforcement learning algorithms or PID policies on these tasks and then select suboptimal policies with returns ranging from 50% to 80% of the expert's return to generate offline datasets for each task. These suboptimal policy-sampled datasets better align with real-world task scenarios compared to random or expert policy datasets.

The following paper provides details about the NeoRL‑2 benchmark:

Songyi Gao, Zuolin Tu, Rong‑Jun Qin, Yi‑Hao Sun, Xiong‑Hui Chen, and Yang Yu.
NeoRL‑2: Near Real‑World Benchmarks for Offline Reinforcement Learning with Extended Realistic Scenarios.
arXiv preprintarXiv:2503.19267, 2025.

The dataset is released in huggingface neorl2

Install NeoRL2 interface

NeoRL2 interface can be installed as follows:

git clone https://agit.ai/Polixir/neorl2.git
cd neorl
pip install -e .

After installation, Pipeline、Simglucose、RocketRecover、DMSD and Fusion environments will be available. However, the "RandomFrictionHopper" and "SafetyHalfCheetah" tasks rely on MuJoCo. If you need to use these two environments, it is necessary to obtain a license and follow the setup instructions, and then run:

pip install -e .[mujoco]

Using NeoRL2

NeoRL2 uses the OpenAI Gym API. Tasks can be created as follows:

import neorl2
import gymnasium as gym
# Create an environment
env = gym.make("Pipeline")
env.reset()
env.step(env.action_space.sample())

After creating the environment, you can use the get_dataset() function to obtain training data and validation data:

train_data, val_data = env.get_dataset()

Each environment supports setting and getting the reward function and done function of the environment, which is very useful for adjusting the environment settings when needed.

# Set reward function
env.set_reward_func(reward_func)
# Get reward function
env.get_reward_func(reward_func)
# Set done function
env.get_done_func(done_func)
# Get done function
env.set_done_func(done_func)

You can use the following environments now:

Env Nameobservation shapeaction shapehave donemax timesteps
Pipeline521False1000
Simglucose311True480
RocketRecovery72True500
RandomFrictionHopper133True1000
DMSD62False100
Fusion156False100
SafetyHalfCheetah186False1000

Data in NeoRL2

In NeoRL2, training data and validation data returned by get_dataset() function are dict with the same format:

  • obs: An N by observation dimensional array of current step's observation.

  • next_obs: An N by observation dimensional array of next step's observation.

  • action: An N by action dimensional array of actions.

  • reward: An N dimensional array of rewards.

  • done: An N dimensional array of episode termination flags.

  • index: An trajectory number-dimensional array. The numbers in index indicate the beginning of trajectories.

Reference

Simglucose: Jinyu Xie. Simglucose v0.2.1 (2018) [Online]. Available: https://github.com/jxx123/simglucose. Accessed on: 5-17-2024. code

DMSD: Char, Ian, et al. "Correlated Trajectory Uncertainty for Adaptive Sequential Decision Making." NeurIPS 2023 Workshop on Adaptive Experimental Design and Active Learning in the Real World. 2023. papercode

MuJoCo: Todorov E, Erez T, Tassa Y. "Mujoco: A Physics Engine for Model-based Control." Proceedings of the 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, pp. 5026-5033, 2012. paperwebsite

Gym: Brockman, Greg, et al. "Openai gym." arXiv preprint arXiv:1606.01540 (2016). papercode

Citation

Please use the following bibtex for citations:

@misc{gao2025neorl2nearrealworldbenchmarks,
title={NeoRL-2: Near Real-World Benchmarks for Offline Reinforcement Learning with Extended Realistic Scenarios}, author={Songyi Gao and Zuolin Tu and Rong-Jun Qin and Yi-Hao Sun and Xiong-Hui Chen and Yang Yu},
year={2025},
eprint={2503.19267},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2503.19267}, }

Licenses

All datasets are licensed under the Creative Commons Attribution 4.0 License (CC BY), and code is licensed under the Apache 2.0 License.

About

No description, website, or topics provided.

Resources

Stars

20 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - polixir/NeoRL2 · GitHub
Skip to content

Latest commit

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

NeoRL2

LicenseLicense

The NEORL2 repository is an extension of the offline reinforcement learning benchmark NeoRL. The NEORL2 repository contains datasets for training and corresponding environments for testing the trained policies. The current datasets are collected from seven open-source environments: Pipeline, Simglucose, RocketRecovery, RandomFrictionHopper, DMSD, Fusion and SafetyHalfCheetah tasks. We perform online training using reinforcement learning algorithms or PID policies on these tasks and then select suboptimal policies with returns ranging from 50% to 80% of the expert's return to generate offline datasets for each task. These suboptimal policy-sampled datasets better align with real-world task scenarios compared to random or expert policy datasets.

The following paper provides details about the NeoRL‑2 benchmark:

Songyi Gao, Zuolin Tu, Rong‑Jun Qin, Yi‑Hao Sun, Xiong‑Hui Chen, and Yang Yu.
NeoRL‑2: Near Real‑World Benchmarks for Offline Reinforcement Learning with Extended Realistic Scenarios.
arXiv preprintarXiv:2503.19267, 2025.

The dataset is released in huggingface neorl2

Install NeoRL2 interface

NeoRL2 interface can be installed as follows:

git clone https://agit.ai/Polixir/neorl2.git
cd neorl
pip install -e .

After installation, Pipeline、Simglucose、RocketRecover、DMSD and Fusion environments will be available. However, the "RandomFrictionHopper" and "SafetyHalfCheetah" tasks rely on MuJoCo. If you need to use these two environments, it is necessary to obtain a license and follow the setup instructions, and then run:

pip install -e .[mujoco]

Using NeoRL2

NeoRL2 uses the OpenAI Gym API. Tasks can be created as follows:

import neorl2
import gymnasium as gym
# Create an environment
env = gym.make("Pipeline")
env.reset()
env.step(env.action_space.sample())

After creating the environment, you can use the get_dataset() function to obtain training data and validation data:

train_data, val_data = env.get_dataset()

Each environment supports setting and getting the reward function and done function of the environment, which is very useful for adjusting the environment settings when needed.

# Set reward function
env.set_reward_func(reward_func)
# Get reward function
env.get_reward_func(reward_func)
# Set done function
env.get_done_func(done_func)
# Get done function
env.set_done_func(done_func)

You can use the following environments now:

Env Nameobservation shapeaction shapehave donemax timesteps
Pipeline521False1000
Simglucose311True480
RocketRecovery72True500
RandomFrictionHopper133True1000
DMSD62False100
Fusion156False100
SafetyHalfCheetah186False1000

Data in NeoRL2

In NeoRL2, training data and validation data returned by get_dataset() function are dict with the same format:

  • obs: An N by observation dimensional array of current step's observation.

  • next_obs: An N by observation dimensional array of next step's observation.

  • action: An N by action dimensional array of actions.

  • reward: An N dimensional array of rewards.

  • done: An N dimensional array of episode termination flags.

  • index: An trajectory number-dimensional array. The numbers in index indicate the beginning of trajectories.

Reference

Simglucose: Jinyu Xie. Simglucose v0.2.1 (2018) [Online]. Available: https://github.com/jxx123/simglucose. Accessed on: 5-17-2024. code

DMSD: Char, Ian, et al. "Correlated Trajectory Uncertainty for Adaptive Sequential Decision Making." NeurIPS 2023 Workshop on Adaptive Experimental Design and Active Learning in the Real World. 2023. papercode

MuJoCo: Todorov E, Erez T, Tassa Y. "Mujoco: A Physics Engine for Model-based Control." Proceedings of the 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, pp. 5026-5033, 2012. paperwebsite

Gym: Brockman, Greg, et al. "Openai gym." arXiv preprint arXiv:1606.01540 (2016). papercode

Citation

Please use the following bibtex for citations:

@misc{gao2025neorl2nearrealworldbenchmarks,
title={NeoRL-2: Near Real-World Benchmarks for Offline Reinforcement Learning with Extended Realistic Scenarios}, author={Songyi Gao and Zuolin Tu and Rong-Jun Qin and Yi-Hao Sun and Xiong-Hui Chen and Yang Yu},
year={2025},
eprint={2503.19267},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2503.19267}, }

Licenses

All datasets are licensed under the Creative Commons Attribution 4.0 License (CC BY), and code is licensed under the Apache 2.0 License.

About

No description, website, or topics provided.

Resources

Stars

20 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - polixir/NeoRL2 · GitHub
Skip to content

Latest commit

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

NeoRL2

LicenseLicense

The NEORL2 repository is an extension of the offline reinforcement learning benchmark NeoRL. The NEORL2 repository contains datasets for training and corresponding environments for testing the trained policies. The current datasets are collected from seven open-source environments: Pipeline, Simglucose, RocketRecovery, RandomFrictionHopper, DMSD, Fusion and SafetyHalfCheetah tasks. We perform online training using reinforcement learning algorithms or PID policies on these tasks and then select suboptimal policies with returns ranging from 50% to 80% of the expert's return to generate offline datasets for each task. These suboptimal policy-sampled datasets better align with real-world task scenarios compared to random or expert policy datasets.

The following paper provides details about the NeoRL‑2 benchmark:

Songyi Gao, Zuolin Tu, Rong‑Jun Qin, Yi‑Hao Sun, Xiong‑Hui Chen, and Yang Yu.
NeoRL‑2: Near Real‑World Benchmarks for Offline Reinforcement Learning with Extended Realistic Scenarios.
arXiv preprintarXiv:2503.19267, 2025.

The dataset is released in huggingface neorl2

Install NeoRL2 interface

NeoRL2 interface can be installed as follows:

git clone https://agit.ai/Polixir/neorl2.git
cd neorl
pip install -e .

After installation, Pipeline、Simglucose、RocketRecover、DMSD and Fusion environments will be available. However, the "RandomFrictionHopper" and "SafetyHalfCheetah" tasks rely on MuJoCo. If you need to use these two environments, it is necessary to obtain a license and follow the setup instructions, and then run:

pip install -e .[mujoco]

Using NeoRL2

NeoRL2 uses the OpenAI Gym API. Tasks can be created as follows:

import neorl2
import gymnasium as gym
# Create an environment
env = gym.make("Pipeline")
env.reset()
env.step(env.action_space.sample())

After creating the environment, you can use the get_dataset() function to obtain training data and validation data:

train_data, val_data = env.get_dataset()

Each environment supports setting and getting the reward function and done function of the environment, which is very useful for adjusting the environment settings when needed.

# Set reward function
env.set_reward_func(reward_func)
# Get reward function
env.get_reward_func(reward_func)
# Set done function
env.get_done_func(done_func)
# Get done function
env.set_done_func(done_func)

You can use the following environments now:

Env Nameobservation shapeaction shapehave donemax timesteps
Pipeline521False1000
Simglucose311True480
RocketRecovery72True500
RandomFrictionHopper133True1000
DMSD62False100
Fusion156False100
SafetyHalfCheetah186False1000

Data in NeoRL2

In NeoRL2, training data and validation data returned by get_dataset() function are dict with the same format:

  • obs: An N by observation dimensional array of current step's observation.

  • next_obs: An N by observation dimensional array of next step's observation.

  • action: An N by action dimensional array of actions.

  • reward: An N dimensional array of rewards.

  • done: An N dimensional array of episode termination flags.

  • index: An trajectory number-dimensional array. The numbers in index indicate the beginning of trajectories.

Reference

Simglucose: Jinyu Xie. Simglucose v0.2.1 (2018) [Online]. Available: https://github.com/jxx123/simglucose. Accessed on: 5-17-2024. code

DMSD: Char, Ian, et al. "Correlated Trajectory Uncertainty for Adaptive Sequential Decision Making." NeurIPS 2023 Workshop on Adaptive Experimental Design and Active Learning in the Real World. 2023. papercode

MuJoCo: Todorov E, Erez T, Tassa Y. "Mujoco: A Physics Engine for Model-based Control." Proceedings of the 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, pp. 5026-5033, 2012. paperwebsite

Gym: Brockman, Greg, et al. "Openai gym." arXiv preprint arXiv:1606.01540 (2016). papercode

Citation

Please use the following bibtex for citations:

@misc{gao2025neorl2nearrealworldbenchmarks,
title={NeoRL-2: Near Real-World Benchmarks for Offline Reinforcement Learning with Extended Realistic Scenarios}, author={Songyi Gao and Zuolin Tu and Rong-Jun Qin and Yi-Hao Sun and Xiong-Hui Chen and Yang Yu},
year={2025},
eprint={2503.19267},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2503.19267}, }

Licenses

All datasets are licensed under the Creative Commons Attribution 4.0 License (CC BY), and code is licensed under the Apache 2.0 License.

About

No description, website, or topics provided.

Resources

Stars

20 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - polixir/NeoRL2 · GitHub
Skip to content

Latest commit

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

NeoRL2

LicenseLicense

The NEORL2 repository is an extension of the offline reinforcement learning benchmark NeoRL. The NEORL2 repository contains datasets for training and corresponding environments for testing the trained policies. The current datasets are collected from seven open-source environments: Pipeline, Simglucose, RocketRecovery, RandomFrictionHopper, DMSD, Fusion and SafetyHalfCheetah tasks. We perform online training using reinforcement learning algorithms or PID policies on these tasks and then select suboptimal policies with returns ranging from 50% to 80% of the expert's return to generate offline datasets for each task. These suboptimal policy-sampled datasets better align with real-world task scenarios compared to random or expert policy datasets.

The following paper provides details about the NeoRL‑2 benchmark:

Songyi Gao, Zuolin Tu, Rong‑Jun Qin, Yi‑Hao Sun, Xiong‑Hui Chen, and Yang Yu.
NeoRL‑2: Near Real‑World Benchmarks for Offline Reinforcement Learning with Extended Realistic Scenarios.
arXiv preprintarXiv:2503.19267, 2025.

The dataset is released in huggingface neorl2

Install NeoRL2 interface

NeoRL2 interface can be installed as follows:

git clone https://agit.ai/Polixir/neorl2.git
cd neorl
pip install -e .

After installation, Pipeline、Simglucose、RocketRecover、DMSD and Fusion environments will be available. However, the "RandomFrictionHopper" and "SafetyHalfCheetah" tasks rely on MuJoCo. If you need to use these two environments, it is necessary to obtain a license and follow the setup instructions, and then run:

pip install -e .[mujoco]

Using NeoRL2

NeoRL2 uses the OpenAI Gym API. Tasks can be created as follows:

import neorl2
import gymnasium as gym
# Create an environment
env = gym.make("Pipeline")
env.reset()
env.step(env.action_space.sample())

After creating the environment, you can use the get_dataset() function to obtain training data and validation data:

train_data, val_data = env.get_dataset()

Each environment supports setting and getting the reward function and done function of the environment, which is very useful for adjusting the environment settings when needed.

# Set reward function
env.set_reward_func(reward_func)
# Get reward function
env.get_reward_func(reward_func)
# Set done function
env.get_done_func(done_func)
# Get done function
env.set_done_func(done_func)

You can use the following environments now:

Env Nameobservation shapeaction shapehave donemax timesteps
Pipeline521False1000
Simglucose311True480
RocketRecovery72True500
RandomFrictionHopper133True1000
DMSD62False100
Fusion156False100
SafetyHalfCheetah186False1000

Data in NeoRL2

In NeoRL2, training data and validation data returned by get_dataset() function are dict with the same format:

  • obs: An N by observation dimensional array of current step's observation.

  • next_obs: An N by observation dimensional array of next step's observation.

  • action: An N by action dimensional array of actions.

  • reward: An N dimensional array of rewards.

  • done: An N dimensional array of episode termination flags.

  • index: An trajectory number-dimensional array. The numbers in index indicate the beginning of trajectories.

Reference

Simglucose: Jinyu Xie. Simglucose v0.2.1 (2018) [Online]. Available: https://github.com/jxx123/simglucose. Accessed on: 5-17-2024. code

DMSD: Char, Ian, et al. "Correlated Trajectory Uncertainty for Adaptive Sequential Decision Making." NeurIPS 2023 Workshop on Adaptive Experimental Design and Active Learning in the Real World. 2023. papercode

MuJoCo: Todorov E, Erez T, Tassa Y. "Mujoco: A Physics Engine for Model-based Control." Proceedings of the 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, pp. 5026-5033, 2012. paperwebsite

Gym: Brockman, Greg, et al. "Openai gym." arXiv preprint arXiv:1606.01540 (2016). papercode

Citation

Please use the following bibtex for citations:

@misc{gao2025neorl2nearrealworldbenchmarks,
title={NeoRL-2: Near Real-World Benchmarks for Offline Reinforcement Learning with Extended Realistic Scenarios}, author={Songyi Gao and Zuolin Tu and Rong-Jun Qin and Yi-Hao Sun and Xiong-Hui Chen and Yang Yu},
year={2025},
eprint={2503.19267},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2503.19267}, }

Licenses

All datasets are licensed under the Creative Commons Attribution 4.0 License (CC BY), and code is licensed under the Apache 2.0 License.

About

No description, website, or topics provided.

Resources

Stars

20 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - polixir/NeoRL2 · GitHub
Skip to content

Latest commit

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

NeoRL2

LicenseLicense

The NEORL2 repository is an extension of the offline reinforcement learning benchmark NeoRL. The NEORL2 repository contains datasets for training and corresponding environments for testing the trained policies. The current datasets are collected from seven open-source environments: Pipeline, Simglucose, RocketRecovery, RandomFrictionHopper, DMSD, Fusion and SafetyHalfCheetah tasks. We perform online training using reinforcement learning algorithms or PID policies on these tasks and then select suboptimal policies with returns ranging from 50% to 80% of the expert's return to generate offline datasets for each task. These suboptimal policy-sampled datasets better align with real-world task scenarios compared to random or expert policy datasets.

The following paper provides details about the NeoRL‑2 benchmark:

Songyi Gao, Zuolin Tu, Rong‑Jun Qin, Yi‑Hao Sun, Xiong‑Hui Chen, and Yang Yu.
NeoRL‑2: Near Real‑World Benchmarks for Offline Reinforcement Learning with Extended Realistic Scenarios.
arXiv preprintarXiv:2503.19267, 2025.

The dataset is released in huggingface neorl2

Install NeoRL2 interface

NeoRL2 interface can be installed as follows:

git clone https://agit.ai/Polixir/neorl2.git
cd neorl
pip install -e .

After installation, Pipeline、Simglucose、RocketRecover、DMSD and Fusion environments will be available. However, the "RandomFrictionHopper" and "SafetyHalfCheetah" tasks rely on MuJoCo. If you need to use these two environments, it is necessary to obtain a license and follow the setup instructions, and then run:

pip install -e .[mujoco]

Using NeoRL2

NeoRL2 uses the OpenAI Gym API. Tasks can be created as follows:

import neorl2
import gymnasium as gym
# Create an environment
env = gym.make("Pipeline")
env.reset()
env.step(env.action_space.sample())

After creating the environment, you can use the get_dataset() function to obtain training data and validation data:

train_data, val_data = env.get_dataset()

Each environment supports setting and getting the reward function and done function of the environment, which is very useful for adjusting the environment settings when needed.

# Set reward function
env.set_reward_func(reward_func)
# Get reward function
env.get_reward_func(reward_func)
# Set done function
env.get_done_func(done_func)
# Get done function
env.set_done_func(done_func)

You can use the following environments now:

Env Nameobservation shapeaction shapehave donemax timesteps
Pipeline521False1000
Simglucose311True480
RocketRecovery72True500
RandomFrictionHopper133True1000
DMSD62False100
Fusion156False100
SafetyHalfCheetah186False1000

Data in NeoRL2

In NeoRL2, training data and validation data returned by get_dataset() function are dict with the same format:

  • obs: An N by observation dimensional array of current step's observation.

  • next_obs: An N by observation dimensional array of next step's observation.

  • action: An N by action dimensional array of actions.

  • reward: An N dimensional array of rewards.

  • done: An N dimensional array of episode termination flags.

  • index: An trajectory number-dimensional array. The numbers in index indicate the beginning of trajectories.

Reference

Simglucose: Jinyu Xie. Simglucose v0.2.1 (2018) [Online]. Available: https://github.com/jxx123/simglucose. Accessed on: 5-17-2024. code

DMSD: Char, Ian, et al. "Correlated Trajectory Uncertainty for Adaptive Sequential Decision Making." NeurIPS 2023 Workshop on Adaptive Experimental Design and Active Learning in the Real World. 2023. papercode

MuJoCo: Todorov E, Erez T, Tassa Y. "Mujoco: A Physics Engine for Model-based Control." Proceedings of the 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, pp. 5026-5033, 2012. paperwebsite

Gym: Brockman, Greg, et al. "Openai gym." arXiv preprint arXiv:1606.01540 (2016). papercode

Citation

Please use the following bibtex for citations:

@misc{gao2025neorl2nearrealworldbenchmarks,
title={NeoRL-2: Near Real-World Benchmarks for Offline Reinforcement Learning with Extended Realistic Scenarios}, author={Songyi Gao and Zuolin Tu and Rong-Jun Qin and Yi-Hao Sun and Xiong-Hui Chen and Yang Yu},
year={2025},
eprint={2503.19267},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2503.19267}, }

Licenses

All datasets are licensed under the Creative Commons Attribution 4.0 License (CC BY), and code is licensed under the Apache 2.0 License.

About

No description, website, or topics provided.

Resources

Stars

20 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - polixir/NeoRL2 · GitHub
Skip to content

Latest commit

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

NeoRL2

LicenseLicense

The NEORL2 repository is an extension of the offline reinforcement learning benchmark NeoRL. The NEORL2 repository contains datasets for training and corresponding environments for testing the trained policies. The current datasets are collected from seven open-source environments: Pipeline, Simglucose, RocketRecovery, RandomFrictionHopper, DMSD, Fusion and SafetyHalfCheetah tasks. We perform online training using reinforcement learning algorithms or PID policies on these tasks and then select suboptimal policies with returns ranging from 50% to 80% of the expert's return to generate offline datasets for each task. These suboptimal policy-sampled datasets better align with real-world task scenarios compared to random or expert policy datasets.

The following paper provides details about the NeoRL‑2 benchmark:

Songyi Gao, Zuolin Tu, Rong‑Jun Qin, Yi‑Hao Sun, Xiong‑Hui Chen, and Yang Yu.
NeoRL‑2: Near Real‑World Benchmarks for Offline Reinforcement Learning with Extended Realistic Scenarios.
arXiv preprintarXiv:2503.19267, 2025.

The dataset is released in huggingface neorl2

Install NeoRL2 interface

NeoRL2 interface can be installed as follows:

git clone https://agit.ai/Polixir/neorl2.git
cd neorl
pip install -e .

After installation, Pipeline、Simglucose、RocketRecover、DMSD and Fusion environments will be available. However, the "RandomFrictionHopper" and "SafetyHalfCheetah" tasks rely on MuJoCo. If you need to use these two environments, it is necessary to obtain a license and follow the setup instructions, and then run:

pip install -e .[mujoco]

Using NeoRL2

NeoRL2 uses the OpenAI Gym API. Tasks can be created as follows:

import neorl2
import gymnasium as gym
# Create an environment
env = gym.make("Pipeline")
env.reset()
env.step(env.action_space.sample())

After creating the environment, you can use the get_dataset() function to obtain training data and validation data:

train_data, val_data = env.get_dataset()

Each environment supports setting and getting the reward function and done function of the environment, which is very useful for adjusting the environment settings when needed.

# Set reward function
env.set_reward_func(reward_func)
# Get reward function
env.get_reward_func(reward_func)
# Set done function
env.get_done_func(done_func)
# Get done function
env.set_done_func(done_func)

You can use the following environments now:

Env Nameobservation shapeaction shapehave donemax timesteps
Pipeline521False1000
Simglucose311True480
RocketRecovery72True500
RandomFrictionHopper133True1000
DMSD62False100
Fusion156False100
SafetyHalfCheetah186False1000

Data in NeoRL2

In NeoRL2, training data and validation data returned by get_dataset() function are dict with the same format:

  • obs: An N by observation dimensional array of current step's observation.

  • next_obs: An N by observation dimensional array of next step's observation.

  • action: An N by action dimensional array of actions.

  • reward: An N dimensional array of rewards.

  • done: An N dimensional array of episode termination flags.

  • index: An trajectory number-dimensional array. The numbers in index indicate the beginning of trajectories.

Reference

Simglucose: Jinyu Xie. Simglucose v0.2.1 (2018) [Online]. Available: https://github.com/jxx123/simglucose. Accessed on: 5-17-2024. code

DMSD: Char, Ian, et al. "Correlated Trajectory Uncertainty for Adaptive Sequential Decision Making." NeurIPS 2023 Workshop on Adaptive Experimental Design and Active Learning in the Real World. 2023. papercode

MuJoCo: Todorov E, Erez T, Tassa Y. "Mujoco: A Physics Engine for Model-based Control." Proceedings of the 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, pp. 5026-5033, 2012. paperwebsite

Gym: Brockman, Greg, et al. "Openai gym." arXiv preprint arXiv:1606.01540 (2016). papercode

Citation

Please use the following bibtex for citations:

@misc{gao2025neorl2nearrealworldbenchmarks,
title={NeoRL-2: Near Real-World Benchmarks for Offline Reinforcement Learning with Extended Realistic Scenarios}, author={Songyi Gao and Zuolin Tu and Rong-Jun Qin and Yi-Hao Sun and Xiong-Hui Chen and Yang Yu},
year={2025},
eprint={2503.19267},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2503.19267}, }

Licenses

All datasets are licensed under the Creative Commons Attribution 4.0 License (CC BY), and code is licensed under the Apache 2.0 License.

About

No description, website, or topics provided.

Resources

Stars

20 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); })(); GitHub - polixir/NeoRL2 · GitHub
Skip to content

Latest commit

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

NeoRL2

LicenseLicense

The NEORL2 repository is an extension of the offline reinforcement learning benchmark NeoRL. The NEORL2 repository contains datasets for training and corresponding environments for testing the trained policies. The current datasets are collected from seven open-source environments: Pipeline, Simglucose, RocketRecovery, RandomFrictionHopper, DMSD, Fusion and SafetyHalfCheetah tasks. We perform online training using reinforcement learning algorithms or PID policies on these tasks and then select suboptimal policies with returns ranging from 50% to 80% of the expert's return to generate offline datasets for each task. These suboptimal policy-sampled datasets better align with real-world task scenarios compared to random or expert policy datasets.

The following paper provides details about the NeoRL‑2 benchmark:

Songyi Gao, Zuolin Tu, Rong‑Jun Qin, Yi‑Hao Sun, Xiong‑Hui Chen, and Yang Yu.
NeoRL‑2: Near Real‑World Benchmarks for Offline Reinforcement Learning with Extended Realistic Scenarios.
arXiv preprintarXiv:2503.19267, 2025.

The dataset is released in huggingface neorl2

Install NeoRL2 interface

NeoRL2 interface can be installed as follows:

git clone https://agit.ai/Polixir/neorl2.git
cd neorl
pip install -e .

After installation, Pipeline、Simglucose、RocketRecover、DMSD and Fusion environments will be available. However, the "RandomFrictionHopper" and "SafetyHalfCheetah" tasks rely on MuJoCo. If you need to use these two environments, it is necessary to obtain a license and follow the setup instructions, and then run:

pip install -e .[mujoco]

Using NeoRL2

NeoRL2 uses the OpenAI Gym API. Tasks can be created as follows:

import neorl2
import gymnasium as gym
# Create an environment
env = gym.make("Pipeline")
env.reset()
env.step(env.action_space.sample())

After creating the environment, you can use the get_dataset() function to obtain training data and validation data:

train_data, val_data = env.get_dataset()

Each environment supports setting and getting the reward function and done function of the environment, which is very useful for adjusting the environment settings when needed.

# Set reward function
env.set_reward_func(reward_func)
# Get reward function
env.get_reward_func(reward_func)
# Set done function
env.get_done_func(done_func)
# Get done function
env.set_done_func(done_func)

You can use the following environments now:

Env Nameobservation shapeaction shapehave donemax timesteps
Pipeline521False1000
Simglucose311True480
RocketRecovery72True500
RandomFrictionHopper133True1000
DMSD62False100
Fusion156False100
SafetyHalfCheetah186False1000

Data in NeoRL2

In NeoRL2, training data and validation data returned by get_dataset() function are dict with the same format:

  • obs: An N by observation dimensional array of current step's observation.

  • next_obs: An N by observation dimensional array of next step's observation.

  • action: An N by action dimensional array of actions.

  • reward: An N dimensional array of rewards.

  • done: An N dimensional array of episode termination flags.

  • index: An trajectory number-dimensional array. The numbers in index indicate the beginning of trajectories.

Reference

Simglucose: Jinyu Xie. Simglucose v0.2.1 (2018) [Online]. Available: https://github.com/jxx123/simglucose. Accessed on: 5-17-2024. code

DMSD: Char, Ian, et al. "Correlated Trajectory Uncertainty for Adaptive Sequential Decision Making." NeurIPS 2023 Workshop on Adaptive Experimental Design and Active Learning in the Real World. 2023. papercode

MuJoCo: Todorov E, Erez T, Tassa Y. "Mujoco: A Physics Engine for Model-based Control." Proceedings of the 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, pp. 5026-5033, 2012. paperwebsite

Gym: Brockman, Greg, et al. "Openai gym." arXiv preprint arXiv:1606.01540 (2016). papercode

Citation

Please use the following bibtex for citations:

@misc{gao2025neorl2nearrealworldbenchmarks,
title={NeoRL-2: Near Real-World Benchmarks for Offline Reinforcement Learning with Extended Realistic Scenarios}, author={Songyi Gao and Zuolin Tu and Rong-Jun Qin and Yi-Hao Sun and Xiong-Hui Chen and Yang Yu},
year={2025},
eprint={2503.19267},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2503.19267}, }

Licenses

All datasets are licensed under the Creative Commons Attribution 4.0 License (CC BY), and code is licensed under the Apache 2.0 License.

About

No description, website, or topics provided.

Resources

Stars

20 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages