Repository files navigation

Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models

arXivPythonLicenseStatic Badge

Nils Blank, Moritz Reuss, Marcel Rühle, Ömer Erdinç Yağmurlu, Fabian Wenzel, Oier Mees, Rudolf Lioutikov


This is the official repository for NILS: Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models.

We present NILS: Natural language Instruction Labeling for Scalability, a framework to label long-horizon robot demonstrations with natural language instructions. Given a long-horizon demonstration, the framework detects keystates and segments the demonstration into indivudal tasks, while generating language instructions.

Quickstart

To label your own demonstration videos, use the following command:

python annotate.py \
dataset.path=<path to folder containing videos> \
dataset.name=<name of the dataset> \
GPU_IDS=[0]<List of gpu_ids to use> \
PROC_PER_GPU=<Number of processes per GPU> \
embodiment_prompt_grounding_dino="the black robotic gripper" <Embodiment prompt for grounding dino> \
embodiment_prompt_clipseg=<Embodiment Prompt for ClipSeg>

This command will use the default settings. The framework can be tuned to specific datasets via configs. We will explain the process in more detail below.

Framework Usage

This section will describe how you can use and adapt the framework to annotate your own datasets.

Creating a Dataset

To create a new dataset, create a new torch.utils.data.Dataset class that inherits from torch.utils.data.Dataset. The class should have the following attributes:

  • paths: A list of paths to the folders containing the data. This will later be used to store the annotations.
  • name: The name of the dataset, used for logging.

The __getitem__ method should return a dictionary with the following keys:

  • rgb_static: np.ndarray containing the RGB images.
  • paths: Path to the currently processed data.
  • frame_names: Original frame indices. Will be used for saving the detected keystates and language annotations.
  • gripper_actions (optional): Binary gripper actions, if available.
  • keystates (optional): Frame indices of keystates, if available.

Incorporating prior knowledge

To improve the frameworks accuracy, you can incorporate prior knowledge, such as objects in the scene, available tasks or keystates of long-horizon demonstrations.

Objects

Create a list of objects in format

pot:silver
spoon:yellow
...

and specify the path to the file in the configuration file (object_list). By default, the framework will still check for objects in the scene and will not output objects with a high overlap with predefined objects. If you only want to use predefined objects, set only_predefined_object to True

Tasks

Create a list of tasks in format

place pot on stove
...

and specify the path to the file in the configuration file (task_list).

Keystates

If you want to load ground truth keystates, adjust the Dataset to return the keystates in the dictionary. The framework will then use the keystates to annotate the scene.

Config

We use Hydra as config manager. The main config file is located in conf/base.yaml.

You can control the number keystate heuristics by changing the entries in keystate_predictor and keystate_predictors_for_voting.

  • GPU_IDS What GPUs to use for labeling
  • PROC_PER_GPU Number of processes per GPU. Increase depending on your VRAM.
  • n_splits The framework currently has memory leaks, which is why we frequently have to reinitalize the framework. This setting divides the dataset into n_splits chunks which are processed in parallel. The number of processes used is GPU_IDS * PROC_PER_GPU

Repository Structure

  • nils/ contains the NILS codebase.
  • nils/annotator contains the main code of the framework and its components.
  • nils/my_datasets contain the datasets used in the paper.
  • nils/specialist_models contains the foundation models used to annotate the scene.
  • conf/ contains the configuration files for the framework.
  • scripts/ contains scripts to start annotation of datasets.
  • scripts/experiments contains the code used to conduct the experiments in the paper.

Installation

git clone --recurse-submodules git@github.com:intuitive-robots/NILS.git
pip install torch==2.3.1
conda create -n "NILS" python=3.9
conda activate NILS pip install -r requirements.txt
pip install -U openmim
mim install mmcv

The framework relies on several specialst models, which you need to install manually

DEVA

cd dependencies/Tracking-Anything-with-DEVA
pip install -e . --no-dependencies
bash scripts/download_models.sh

SAM2

cd nils/specialst_models/sam2/checkpoints
./download_ckpts.sh

DepthAnythingv2

cd dependencies/Depth-Anything-V2
mkdir checkpoints;cd checkpoints
wget https://huggingface.co/depth-anything/Depth-Anything-V2-Metric-Hypersim-Large/resolve/main/depth_anything_v2_metric_hypersim_vitl.pth?download=true -O depth_anything_v2_metric_hypersim_vitl.pth
export PYTHONPATH=$PYTHONPATH:path/to/Depth-Anything-V2/metric_depth

Unimatch (GMFLOW)

cd dependencies/unimatch
export PYTHONPATH=$PYTHONPATH:path/to/unimatch
mkdir pretrained;cd pretrained
wget https://s3.eu-central-1.amazonaws.com/avg-projects/unimatch/pretrained/gmflow-scale2-regrefine6-mixdata-train320x576-4e7b215d.pth

DINOv2

git clone https://github.com/facebookresearch/dinov2.git
export PYTHONPATH=$PYTHONPATH:path/to/dinov2

Citation

If you find our framework useful in your work, please cite our paper:

@inproceedings{
blank2024scaling,
title={Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models},
author={Nils Blank and Moritz Reuss and Marcel R{\"u}hle and {\"O}mer Erdin{\c{c}} Ya{\u{g}}murlu and Fabian Wenzel and Oier Mees and Rudolf Lioutikov},
booktitle={8th Annual Conference on Robot Learning},
year={2024},
url={https://openreview.net/forum?id=EdVNB2kHv1}
}

About

[CoRL 2024] Official code for "Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models"

Resources

Stars

36 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models

arXivPythonLicenseStatic Badge

Nils Blank, Moritz Reuss, Marcel Rühle, Ömer Erdinç Yağmurlu, Fabian Wenzel, Oier Mees, Rudolf Lioutikov


This is the official repository for NILS: Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models.

We present NILS: Natural language Instruction Labeling for Scalability, a framework to label long-horizon robot demonstrations with natural language instructions. Given a long-horizon demonstration, the framework detects keystates and segments the demonstration into indivudal tasks, while generating language instructions.

Quickstart

To label your own demonstration videos, use the following command:

python annotate.py \
dataset.path=<path to folder containing videos> \
dataset.name=<name of the dataset> \
GPU_IDS=[0]<List of gpu_ids to use> \
PROC_PER_GPU=<Number of processes per GPU> \
embodiment_prompt_grounding_dino="the black robotic gripper" <Embodiment prompt for grounding dino> \
embodiment_prompt_clipseg=<Embodiment Prompt for ClipSeg>

This command will use the default settings. The framework can be tuned to specific datasets via configs. We will explain the process in more detail below.

Framework Usage

This section will describe how you can use and adapt the framework to annotate your own datasets.

Creating a Dataset

To create a new dataset, create a new torch.utils.data.Dataset class that inherits from torch.utils.data.Dataset. The class should have the following attributes:

  • paths: A list of paths to the folders containing the data. This will later be used to store the annotations.
  • name: The name of the dataset, used for logging.

The __getitem__ method should return a dictionary with the following keys:

  • rgb_static: np.ndarray containing the RGB images.
  • paths: Path to the currently processed data.
  • frame_names: Original frame indices. Will be used for saving the detected keystates and language annotations.
  • gripper_actions (optional): Binary gripper actions, if available.
  • keystates (optional): Frame indices of keystates, if available.

Incorporating prior knowledge

To improve the frameworks accuracy, you can incorporate prior knowledge, such as objects in the scene, available tasks or keystates of long-horizon demonstrations.

Objects

Create a list of objects in format

pot:silver
spoon:yellow
...

and specify the path to the file in the configuration file (object_list). By default, the framework will still check for objects in the scene and will not output objects with a high overlap with predefined objects. If you only want to use predefined objects, set only_predefined_object to True

Tasks

Create a list of tasks in format

place pot on stove
...

and specify the path to the file in the configuration file (task_list).

Keystates

If you want to load ground truth keystates, adjust the Dataset to return the keystates in the dictionary. The framework will then use the keystates to annotate the scene.

Config

We use Hydra as config manager. The main config file is located in conf/base.yaml.

You can control the number keystate heuristics by changing the entries in keystate_predictor and keystate_predictors_for_voting.

  • GPU_IDS What GPUs to use for labeling
  • PROC_PER_GPU Number of processes per GPU. Increase depending on your VRAM.
  • n_splits The framework currently has memory leaks, which is why we frequently have to reinitalize the framework. This setting divides the dataset into n_splits chunks which are processed in parallel. The number of processes used is GPU_IDS * PROC_PER_GPU

Repository Structure

  • nils/ contains the NILS codebase.
  • nils/annotator contains the main code of the framework and its components.
  • nils/my_datasets contain the datasets used in the paper.
  • nils/specialist_models contains the foundation models used to annotate the scene.
  • conf/ contains the configuration files for the framework.
  • scripts/ contains scripts to start annotation of datasets.
  • scripts/experiments contains the code used to conduct the experiments in the paper.

Installation

git clone --recurse-submodules git@github.com:intuitive-robots/NILS.git
pip install torch==2.3.1
conda create -n "NILS" python=3.9
conda activate NILS pip install -r requirements.txt
pip install -U openmim
mim install mmcv

The framework relies on several specialst models, which you need to install manually

DEVA

cd dependencies/Tracking-Anything-with-DEVA
pip install -e . --no-dependencies
bash scripts/download_models.sh

SAM2

cd nils/specialst_models/sam2/checkpoints
./download_ckpts.sh

DepthAnythingv2

cd dependencies/Depth-Anything-V2
mkdir checkpoints;cd checkpoints
wget https://huggingface.co/depth-anything/Depth-Anything-V2-Metric-Hypersim-Large/resolve/main/depth_anything_v2_metric_hypersim_vitl.pth?download=true -O depth_anything_v2_metric_hypersim_vitl.pth
export PYTHONPATH=$PYTHONPATH:path/to/Depth-Anything-V2/metric_depth

Unimatch (GMFLOW)

cd dependencies/unimatch
export PYTHONPATH=$PYTHONPATH:path/to/unimatch
mkdir pretrained;cd pretrained
wget https://s3.eu-central-1.amazonaws.com/avg-projects/unimatch/pretrained/gmflow-scale2-regrefine6-mixdata-train320x576-4e7b215d.pth

DINOv2

git clone https://github.com/facebookresearch/dinov2.git
export PYTHONPATH=$PYTHONPATH:path/to/dinov2

Citation

If you find our framework useful in your work, please cite our paper:

@inproceedings{
blank2024scaling,
title={Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models},
author={Nils Blank and Moritz Reuss and Marcel R{\"u}hle and {\"O}mer Erdin{\c{c}} Ya{\u{g}}murlu and Fabian Wenzel and Oier Mees and Rudolf Lioutikov},
booktitle={8th Annual Conference on Robot Learning},
year={2024},
url={https://openreview.net/forum?id=EdVNB2kHv1}
}

About

[CoRL 2024] Official code for "Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models"

Resources

Stars

36 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models

arXivPythonLicenseStatic Badge

Nils Blank, Moritz Reuss, Marcel Rühle, Ömer Erdinç Yağmurlu, Fabian Wenzel, Oier Mees, Rudolf Lioutikov


This is the official repository for NILS: Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models.

We present NILS: Natural language Instruction Labeling for Scalability, a framework to label long-horizon robot demonstrations with natural language instructions. Given a long-horizon demonstration, the framework detects keystates and segments the demonstration into indivudal tasks, while generating language instructions.

Quickstart

To label your own demonstration videos, use the following command:

python annotate.py \
dataset.path=<path to folder containing videos> \
dataset.name=<name of the dataset> \
GPU_IDS=[0]<List of gpu_ids to use> \
PROC_PER_GPU=<Number of processes per GPU> \
embodiment_prompt_grounding_dino="the black robotic gripper" <Embodiment prompt for grounding dino> \
embodiment_prompt_clipseg=<Embodiment Prompt for ClipSeg>

This command will use the default settings. The framework can be tuned to specific datasets via configs. We will explain the process in more detail below.

Framework Usage

This section will describe how you can use and adapt the framework to annotate your own datasets.

Creating a Dataset

To create a new dataset, create a new torch.utils.data.Dataset class that inherits from torch.utils.data.Dataset. The class should have the following attributes:

  • paths: A list of paths to the folders containing the data. This will later be used to store the annotations.
  • name: The name of the dataset, used for logging.

The __getitem__ method should return a dictionary with the following keys:

  • rgb_static: np.ndarray containing the RGB images.
  • paths: Path to the currently processed data.
  • frame_names: Original frame indices. Will be used for saving the detected keystates and language annotations.
  • gripper_actions (optional): Binary gripper actions, if available.
  • keystates (optional): Frame indices of keystates, if available.

Incorporating prior knowledge

To improve the frameworks accuracy, you can incorporate prior knowledge, such as objects in the scene, available tasks or keystates of long-horizon demonstrations.

Objects

Create a list of objects in format

pot:silver
spoon:yellow
...

and specify the path to the file in the configuration file (object_list). By default, the framework will still check for objects in the scene and will not output objects with a high overlap with predefined objects. If you only want to use predefined objects, set only_predefined_object to True

Tasks

Create a list of tasks in format

place pot on stove
...

and specify the path to the file in the configuration file (task_list).

Keystates

If you want to load ground truth keystates, adjust the Dataset to return the keystates in the dictionary. The framework will then use the keystates to annotate the scene.

Config

We use Hydra as config manager. The main config file is located in conf/base.yaml.

You can control the number keystate heuristics by changing the entries in keystate_predictor and keystate_predictors_for_voting.

  • GPU_IDS What GPUs to use for labeling
  • PROC_PER_GPU Number of processes per GPU. Increase depending on your VRAM.
  • n_splits The framework currently has memory leaks, which is why we frequently have to reinitalize the framework. This setting divides the dataset into n_splits chunks which are processed in parallel. The number of processes used is GPU_IDS * PROC_PER_GPU

Repository Structure

  • nils/ contains the NILS codebase.
  • nils/annotator contains the main code of the framework and its components.
  • nils/my_datasets contain the datasets used in the paper.
  • nils/specialist_models contains the foundation models used to annotate the scene.
  • conf/ contains the configuration files for the framework.
  • scripts/ contains scripts to start annotation of datasets.
  • scripts/experiments contains the code used to conduct the experiments in the paper.

Installation

git clone --recurse-submodules git@github.com:intuitive-robots/NILS.git
pip install torch==2.3.1
conda create -n "NILS" python=3.9
conda activate NILS pip install -r requirements.txt
pip install -U openmim
mim install mmcv

The framework relies on several specialst models, which you need to install manually

DEVA

cd dependencies/Tracking-Anything-with-DEVA
pip install -e . --no-dependencies
bash scripts/download_models.sh

SAM2

cd nils/specialst_models/sam2/checkpoints
./download_ckpts.sh

DepthAnythingv2

cd dependencies/Depth-Anything-V2
mkdir checkpoints;cd checkpoints
wget https://huggingface.co/depth-anything/Depth-Anything-V2-Metric-Hypersim-Large/resolve/main/depth_anything_v2_metric_hypersim_vitl.pth?download=true -O depth_anything_v2_metric_hypersim_vitl.pth
export PYTHONPATH=$PYTHONPATH:path/to/Depth-Anything-V2/metric_depth

Unimatch (GMFLOW)

cd dependencies/unimatch
export PYTHONPATH=$PYTHONPATH:path/to/unimatch
mkdir pretrained;cd pretrained
wget https://s3.eu-central-1.amazonaws.com/avg-projects/unimatch/pretrained/gmflow-scale2-regrefine6-mixdata-train320x576-4e7b215d.pth

DINOv2

git clone https://github.com/facebookresearch/dinov2.git
export PYTHONPATH=$PYTHONPATH:path/to/dinov2

Citation

If you find our framework useful in your work, please cite our paper:

@inproceedings{
blank2024scaling,
title={Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models},
author={Nils Blank and Moritz Reuss and Marcel R{\"u}hle and {\"O}mer Erdin{\c{c}} Ya{\u{g}}murlu and Fabian Wenzel and Oier Mees and Rudolf Lioutikov},
booktitle={8th Annual Conference on Robot Learning},
year={2024},
url={https://openreview.net/forum?id=EdVNB2kHv1}
}

About

[CoRL 2024] Official code for "Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models"

Resources

Stars

36 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models

arXivPythonLicenseStatic Badge

Nils Blank, Moritz Reuss, Marcel Rühle, Ömer Erdinç Yağmurlu, Fabian Wenzel, Oier Mees, Rudolf Lioutikov


This is the official repository for NILS: Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models.

We present NILS: Natural language Instruction Labeling for Scalability, a framework to label long-horizon robot demonstrations with natural language instructions. Given a long-horizon demonstration, the framework detects keystates and segments the demonstration into indivudal tasks, while generating language instructions.

Quickstart

To label your own demonstration videos, use the following command:

python annotate.py \
dataset.path=<path to folder containing videos> \
dataset.name=<name of the dataset> \
GPU_IDS=[0]<List of gpu_ids to use> \
PROC_PER_GPU=<Number of processes per GPU> \
embodiment_prompt_grounding_dino="the black robotic gripper" <Embodiment prompt for grounding dino> \
embodiment_prompt_clipseg=<Embodiment Prompt for ClipSeg>

This command will use the default settings. The framework can be tuned to specific datasets via configs. We will explain the process in more detail below.

Framework Usage

This section will describe how you can use and adapt the framework to annotate your own datasets.

Creating a Dataset

To create a new dataset, create a new torch.utils.data.Dataset class that inherits from torch.utils.data.Dataset. The class should have the following attributes:

  • paths: A list of paths to the folders containing the data. This will later be used to store the annotations.
  • name: The name of the dataset, used for logging.

The __getitem__ method should return a dictionary with the following keys:

  • rgb_static: np.ndarray containing the RGB images.
  • paths: Path to the currently processed data.
  • frame_names: Original frame indices. Will be used for saving the detected keystates and language annotations.
  • gripper_actions (optional): Binary gripper actions, if available.
  • keystates (optional): Frame indices of keystates, if available.

Incorporating prior knowledge

To improve the frameworks accuracy, you can incorporate prior knowledge, such as objects in the scene, available tasks or keystates of long-horizon demonstrations.

Objects

Create a list of objects in format

pot:silver
spoon:yellow
...

and specify the path to the file in the configuration file (object_list). By default, the framework will still check for objects in the scene and will not output objects with a high overlap with predefined objects. If you only want to use predefined objects, set only_predefined_object to True

Tasks

Create a list of tasks in format

place pot on stove
...

and specify the path to the file in the configuration file (task_list).

Keystates

If you want to load ground truth keystates, adjust the Dataset to return the keystates in the dictionary. The framework will then use the keystates to annotate the scene.

Config

We use Hydra as config manager. The main config file is located in conf/base.yaml.

You can control the number keystate heuristics by changing the entries in keystate_predictor and keystate_predictors_for_voting.

  • GPU_IDS What GPUs to use for labeling
  • PROC_PER_GPU Number of processes per GPU. Increase depending on your VRAM.
  • n_splits The framework currently has memory leaks, which is why we frequently have to reinitalize the framework. This setting divides the dataset into n_splits chunks which are processed in parallel. The number of processes used is GPU_IDS * PROC_PER_GPU

Repository Structure

  • nils/ contains the NILS codebase.
  • nils/annotator contains the main code of the framework and its components.
  • nils/my_datasets contain the datasets used in the paper.
  • nils/specialist_models contains the foundation models used to annotate the scene.
  • conf/ contains the configuration files for the framework.
  • scripts/ contains scripts to start annotation of datasets.
  • scripts/experiments contains the code used to conduct the experiments in the paper.

Installation

git clone --recurse-submodules git@github.com:intuitive-robots/NILS.git
pip install torch==2.3.1
conda create -n "NILS" python=3.9
conda activate NILS pip install -r requirements.txt
pip install -U openmim
mim install mmcv

The framework relies on several specialst models, which you need to install manually

DEVA

cd dependencies/Tracking-Anything-with-DEVA
pip install -e . --no-dependencies
bash scripts/download_models.sh

SAM2

cd nils/specialst_models/sam2/checkpoints
./download_ckpts.sh

DepthAnythingv2

cd dependencies/Depth-Anything-V2
mkdir checkpoints;cd checkpoints
wget https://huggingface.co/depth-anything/Depth-Anything-V2-Metric-Hypersim-Large/resolve/main/depth_anything_v2_metric_hypersim_vitl.pth?download=true -O depth_anything_v2_metric_hypersim_vitl.pth
export PYTHONPATH=$PYTHONPATH:path/to/Depth-Anything-V2/metric_depth

Unimatch (GMFLOW)

cd dependencies/unimatch
export PYTHONPATH=$PYTHONPATH:path/to/unimatch
mkdir pretrained;cd pretrained
wget https://s3.eu-central-1.amazonaws.com/avg-projects/unimatch/pretrained/gmflow-scale2-regrefine6-mixdata-train320x576-4e7b215d.pth

DINOv2

git clone https://github.com/facebookresearch/dinov2.git
export PYTHONPATH=$PYTHONPATH:path/to/dinov2

Citation

If you find our framework useful in your work, please cite our paper:

@inproceedings{
blank2024scaling,
title={Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models},
author={Nils Blank and Moritz Reuss and Marcel R{\"u}hle and {\"O}mer Erdin{\c{c}} Ya{\u{g}}murlu and Fabian Wenzel and Oier Mees and Rudolf Lioutikov},
booktitle={8th Annual Conference on Robot Learning},
year={2024},
url={https://openreview.net/forum?id=EdVNB2kHv1}
}

About

[CoRL 2024] Official code for "Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models"

Resources

Stars

36 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models

arXivPythonLicenseStatic Badge

Nils Blank, Moritz Reuss, Marcel Rühle, Ömer Erdinç Yağmurlu, Fabian Wenzel, Oier Mees, Rudolf Lioutikov


This is the official repository for NILS: Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models.

We present NILS: Natural language Instruction Labeling for Scalability, a framework to label long-horizon robot demonstrations with natural language instructions. Given a long-horizon demonstration, the framework detects keystates and segments the demonstration into indivudal tasks, while generating language instructions.

Quickstart

To label your own demonstration videos, use the following command:

python annotate.py \
dataset.path=<path to folder containing videos> \
dataset.name=<name of the dataset> \
GPU_IDS=[0]<List of gpu_ids to use> \
PROC_PER_GPU=<Number of processes per GPU> \
embodiment_prompt_grounding_dino="the black robotic gripper" <Embodiment prompt for grounding dino> \
embodiment_prompt_clipseg=<Embodiment Prompt for ClipSeg>

This command will use the default settings. The framework can be tuned to specific datasets via configs. We will explain the process in more detail below.

Framework Usage

This section will describe how you can use and adapt the framework to annotate your own datasets.

Creating a Dataset

To create a new dataset, create a new torch.utils.data.Dataset class that inherits from torch.utils.data.Dataset. The class should have the following attributes:

  • paths: A list of paths to the folders containing the data. This will later be used to store the annotations.
  • name: The name of the dataset, used for logging.

The __getitem__ method should return a dictionary with the following keys:

  • rgb_static: np.ndarray containing the RGB images.
  • paths: Path to the currently processed data.
  • frame_names: Original frame indices. Will be used for saving the detected keystates and language annotations.
  • gripper_actions (optional): Binary gripper actions, if available.
  • keystates (optional): Frame indices of keystates, if available.

Incorporating prior knowledge

To improve the frameworks accuracy, you can incorporate prior knowledge, such as objects in the scene, available tasks or keystates of long-horizon demonstrations.

Objects

Create a list of objects in format

pot:silver
spoon:yellow
...

and specify the path to the file in the configuration file (object_list). By default, the framework will still check for objects in the scene and will not output objects with a high overlap with predefined objects. If you only want to use predefined objects, set only_predefined_object to True

Tasks

Create a list of tasks in format

place pot on stove
...

and specify the path to the file in the configuration file (task_list).

Keystates

If you want to load ground truth keystates, adjust the Dataset to return the keystates in the dictionary. The framework will then use the keystates to annotate the scene.

Config

We use Hydra as config manager. The main config file is located in conf/base.yaml.

You can control the number keystate heuristics by changing the entries in keystate_predictor and keystate_predictors_for_voting.

  • GPU_IDS What GPUs to use for labeling
  • PROC_PER_GPU Number of processes per GPU. Increase depending on your VRAM.
  • n_splits The framework currently has memory leaks, which is why we frequently have to reinitalize the framework. This setting divides the dataset into n_splits chunks which are processed in parallel. The number of processes used is GPU_IDS * PROC_PER_GPU

Repository Structure

  • nils/ contains the NILS codebase.
  • nils/annotator contains the main code of the framework and its components.
  • nils/my_datasets contain the datasets used in the paper.
  • nils/specialist_models contains the foundation models used to annotate the scene.
  • conf/ contains the configuration files for the framework.
  • scripts/ contains scripts to start annotation of datasets.
  • scripts/experiments contains the code used to conduct the experiments in the paper.

Installation

git clone --recurse-submodules git@github.com:intuitive-robots/NILS.git
pip install torch==2.3.1
conda create -n "NILS" python=3.9
conda activate NILS pip install -r requirements.txt
pip install -U openmim
mim install mmcv

The framework relies on several specialst models, which you need to install manually

DEVA

cd dependencies/Tracking-Anything-with-DEVA
pip install -e . --no-dependencies
bash scripts/download_models.sh

SAM2

cd nils/specialst_models/sam2/checkpoints
./download_ckpts.sh

DepthAnythingv2

cd dependencies/Depth-Anything-V2
mkdir checkpoints;cd checkpoints
wget https://huggingface.co/depth-anything/Depth-Anything-V2-Metric-Hypersim-Large/resolve/main/depth_anything_v2_metric_hypersim_vitl.pth?download=true -O depth_anything_v2_metric_hypersim_vitl.pth
export PYTHONPATH=$PYTHONPATH:path/to/Depth-Anything-V2/metric_depth

Unimatch (GMFLOW)

cd dependencies/unimatch
export PYTHONPATH=$PYTHONPATH:path/to/unimatch
mkdir pretrained;cd pretrained
wget https://s3.eu-central-1.amazonaws.com/avg-projects/unimatch/pretrained/gmflow-scale2-regrefine6-mixdata-train320x576-4e7b215d.pth

DINOv2

git clone https://github.com/facebookresearch/dinov2.git
export PYTHONPATH=$PYTHONPATH:path/to/dinov2

Citation

If you find our framework useful in your work, please cite our paper:

@inproceedings{
blank2024scaling,
title={Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models},
author={Nils Blank and Moritz Reuss and Marcel R{\"u}hle and {\"O}mer Erdin{\c{c}} Ya{\u{g}}murlu and Fabian Wenzel and Oier Mees and Rudolf Lioutikov},
booktitle={8th Annual Conference on Robot Learning},
year={2024},
url={https://openreview.net/forum?id=EdVNB2kHv1}
}

About

[CoRL 2024] Official code for "Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models"

Resources

Stars

36 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models

arXivPythonLicenseStatic Badge

Nils Blank, Moritz Reuss, Marcel Rühle, Ömer Erdinç Yağmurlu, Fabian Wenzel, Oier Mees, Rudolf Lioutikov


This is the official repository for NILS: Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models.

We present NILS: Natural language Instruction Labeling for Scalability, a framework to label long-horizon robot demonstrations with natural language instructions. Given a long-horizon demonstration, the framework detects keystates and segments the demonstration into indivudal tasks, while generating language instructions.

Quickstart

To label your own demonstration videos, use the following command:

python annotate.py \
dataset.path=<path to folder containing videos> \
dataset.name=<name of the dataset> \
GPU_IDS=[0]<List of gpu_ids to use> \
PROC_PER_GPU=<Number of processes per GPU> \
embodiment_prompt_grounding_dino="the black robotic gripper" <Embodiment prompt for grounding dino> \
embodiment_prompt_clipseg=<Embodiment Prompt for ClipSeg>

This command will use the default settings. The framework can be tuned to specific datasets via configs. We will explain the process in more detail below.

Framework Usage

This section will describe how you can use and adapt the framework to annotate your own datasets.

Creating a Dataset

To create a new dataset, create a new torch.utils.data.Dataset class that inherits from torch.utils.data.Dataset. The class should have the following attributes:

  • paths: A list of paths to the folders containing the data. This will later be used to store the annotations.
  • name: The name of the dataset, used for logging.

The __getitem__ method should return a dictionary with the following keys:

  • rgb_static: np.ndarray containing the RGB images.
  • paths: Path to the currently processed data.
  • frame_names: Original frame indices. Will be used for saving the detected keystates and language annotations.
  • gripper_actions (optional): Binary gripper actions, if available.
  • keystates (optional): Frame indices of keystates, if available.

Incorporating prior knowledge

To improve the frameworks accuracy, you can incorporate prior knowledge, such as objects in the scene, available tasks or keystates of long-horizon demonstrations.

Objects

Create a list of objects in format

pot:silver
spoon:yellow
...

and specify the path to the file in the configuration file (object_list). By default, the framework will still check for objects in the scene and will not output objects with a high overlap with predefined objects. If you only want to use predefined objects, set only_predefined_object to True

Tasks

Create a list of tasks in format

place pot on stove
...

and specify the path to the file in the configuration file (task_list).

Keystates

If you want to load ground truth keystates, adjust the Dataset to return the keystates in the dictionary. The framework will then use the keystates to annotate the scene.

Config

We use Hydra as config manager. The main config file is located in conf/base.yaml.

You can control the number keystate heuristics by changing the entries in keystate_predictor and keystate_predictors_for_voting.

  • GPU_IDS What GPUs to use for labeling
  • PROC_PER_GPU Number of processes per GPU. Increase depending on your VRAM.
  • n_splits The framework currently has memory leaks, which is why we frequently have to reinitalize the framework. This setting divides the dataset into n_splits chunks which are processed in parallel. The number of processes used is GPU_IDS * PROC_PER_GPU

Repository Structure

  • nils/ contains the NILS codebase.
  • nils/annotator contains the main code of the framework and its components.
  • nils/my_datasets contain the datasets used in the paper.
  • nils/specialist_models contains the foundation models used to annotate the scene.
  • conf/ contains the configuration files for the framework.
  • scripts/ contains scripts to start annotation of datasets.
  • scripts/experiments contains the code used to conduct the experiments in the paper.

Installation

git clone --recurse-submodules git@github.com:intuitive-robots/NILS.git
pip install torch==2.3.1
conda create -n "NILS" python=3.9
conda activate NILS pip install -r requirements.txt
pip install -U openmim
mim install mmcv

The framework relies on several specialst models, which you need to install manually

DEVA

cd dependencies/Tracking-Anything-with-DEVA
pip install -e . --no-dependencies
bash scripts/download_models.sh

SAM2

cd nils/specialst_models/sam2/checkpoints
./download_ckpts.sh

DepthAnythingv2

cd dependencies/Depth-Anything-V2
mkdir checkpoints;cd checkpoints
wget https://huggingface.co/depth-anything/Depth-Anything-V2-Metric-Hypersim-Large/resolve/main/depth_anything_v2_metric_hypersim_vitl.pth?download=true -O depth_anything_v2_metric_hypersim_vitl.pth
export PYTHONPATH=$PYTHONPATH:path/to/Depth-Anything-V2/metric_depth

Unimatch (GMFLOW)

cd dependencies/unimatch
export PYTHONPATH=$PYTHONPATH:path/to/unimatch
mkdir pretrained;cd pretrained
wget https://s3.eu-central-1.amazonaws.com/avg-projects/unimatch/pretrained/gmflow-scale2-regrefine6-mixdata-train320x576-4e7b215d.pth

DINOv2

git clone https://github.com/facebookresearch/dinov2.git
export PYTHONPATH=$PYTHONPATH:path/to/dinov2

Citation

If you find our framework useful in your work, please cite our paper:

@inproceedings{
blank2024scaling,
title={Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models},
author={Nils Blank and Moritz Reuss and Marcel R{\"u}hle and {\"O}mer Erdin{\c{c}} Ya{\u{g}}murlu and Fabian Wenzel and Oier Mees and Rudolf Lioutikov},
booktitle={8th Annual Conference on Robot Learning},
year={2024},
url={https://openreview.net/forum?id=EdVNB2kHv1}
}

About

[CoRL 2024] Official code for "Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models"

Resources

Stars

36 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models

arXivPythonLicenseStatic Badge

Nils Blank, Moritz Reuss, Marcel Rühle, Ömer Erdinç Yağmurlu, Fabian Wenzel, Oier Mees, Rudolf Lioutikov


This is the official repository for NILS: Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models.

We present NILS: Natural language Instruction Labeling for Scalability, a framework to label long-horizon robot demonstrations with natural language instructions. Given a long-horizon demonstration, the framework detects keystates and segments the demonstration into indivudal tasks, while generating language instructions.

Quickstart

To label your own demonstration videos, use the following command:

python annotate.py \
dataset.path=<path to folder containing videos> \
dataset.name=<name of the dataset> \
GPU_IDS=[0]<List of gpu_ids to use> \
PROC_PER_GPU=<Number of processes per GPU> \
embodiment_prompt_grounding_dino="the black robotic gripper" <Embodiment prompt for grounding dino> \
embodiment_prompt_clipseg=<Embodiment Prompt for ClipSeg>

This command will use the default settings. The framework can be tuned to specific datasets via configs. We will explain the process in more detail below.

Framework Usage

This section will describe how you can use and adapt the framework to annotate your own datasets.

Creating a Dataset

To create a new dataset, create a new torch.utils.data.Dataset class that inherits from torch.utils.data.Dataset. The class should have the following attributes:

  • paths: A list of paths to the folders containing the data. This will later be used to store the annotations.
  • name: The name of the dataset, used for logging.

The __getitem__ method should return a dictionary with the following keys:

  • rgb_static: np.ndarray containing the RGB images.
  • paths: Path to the currently processed data.
  • frame_names: Original frame indices. Will be used for saving the detected keystates and language annotations.
  • gripper_actions (optional): Binary gripper actions, if available.
  • keystates (optional): Frame indices of keystates, if available.

Incorporating prior knowledge

To improve the frameworks accuracy, you can incorporate prior knowledge, such as objects in the scene, available tasks or keystates of long-horizon demonstrations.

Objects

Create a list of objects in format

pot:silver
spoon:yellow
...

and specify the path to the file in the configuration file (object_list). By default, the framework will still check for objects in the scene and will not output objects with a high overlap with predefined objects. If you only want to use predefined objects, set only_predefined_object to True

Tasks

Create a list of tasks in format

place pot on stove
...

and specify the path to the file in the configuration file (task_list).

Keystates

If you want to load ground truth keystates, adjust the Dataset to return the keystates in the dictionary. The framework will then use the keystates to annotate the scene.

Config

We use Hydra as config manager. The main config file is located in conf/base.yaml.

You can control the number keystate heuristics by changing the entries in keystate_predictor and keystate_predictors_for_voting.

  • GPU_IDS What GPUs to use for labeling
  • PROC_PER_GPU Number of processes per GPU. Increase depending on your VRAM.
  • n_splits The framework currently has memory leaks, which is why we frequently have to reinitalize the framework. This setting divides the dataset into n_splits chunks which are processed in parallel. The number of processes used is GPU_IDS * PROC_PER_GPU

Repository Structure

  • nils/ contains the NILS codebase.
  • nils/annotator contains the main code of the framework and its components.
  • nils/my_datasets contain the datasets used in the paper.
  • nils/specialist_models contains the foundation models used to annotate the scene.
  • conf/ contains the configuration files for the framework.
  • scripts/ contains scripts to start annotation of datasets.
  • scripts/experiments contains the code used to conduct the experiments in the paper.

Installation

git clone --recurse-submodules git@github.com:intuitive-robots/NILS.git
pip install torch==2.3.1
conda create -n "NILS" python=3.9
conda activate NILS pip install -r requirements.txt
pip install -U openmim
mim install mmcv

The framework relies on several specialst models, which you need to install manually

DEVA

cd dependencies/Tracking-Anything-with-DEVA
pip install -e . --no-dependencies
bash scripts/download_models.sh

SAM2

cd nils/specialst_models/sam2/checkpoints
./download_ckpts.sh

DepthAnythingv2

cd dependencies/Depth-Anything-V2
mkdir checkpoints;cd checkpoints
wget https://huggingface.co/depth-anything/Depth-Anything-V2-Metric-Hypersim-Large/resolve/main/depth_anything_v2_metric_hypersim_vitl.pth?download=true -O depth_anything_v2_metric_hypersim_vitl.pth
export PYTHONPATH=$PYTHONPATH:path/to/Depth-Anything-V2/metric_depth

Unimatch (GMFLOW)

cd dependencies/unimatch
export PYTHONPATH=$PYTHONPATH:path/to/unimatch
mkdir pretrained;cd pretrained
wget https://s3.eu-central-1.amazonaws.com/avg-projects/unimatch/pretrained/gmflow-scale2-regrefine6-mixdata-train320x576-4e7b215d.pth

DINOv2

git clone https://github.com/facebookresearch/dinov2.git
export PYTHONPATH=$PYTHONPATH:path/to/dinov2

Citation

If you find our framework useful in your work, please cite our paper:

@inproceedings{
blank2024scaling,
title={Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models},
author={Nils Blank and Moritz Reuss and Marcel R{\"u}hle and {\"O}mer Erdin{\c{c}} Ya{\u{g}}murlu and Fabian Wenzel and Oier Mees and Rudolf Lioutikov},
booktitle={8th Annual Conference on Robot Learning},
year={2024},
url={https://openreview.net/forum?id=EdVNB2kHv1}
}

About

[CoRL 2024] Official code for "Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models"

Resources

Stars

36 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models

arXivPythonLicenseStatic Badge

Nils Blank, Moritz Reuss, Marcel Rühle, Ömer Erdinç Yağmurlu, Fabian Wenzel, Oier Mees, Rudolf Lioutikov


This is the official repository for NILS: Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models.

We present NILS: Natural language Instruction Labeling for Scalability, a framework to label long-horizon robot demonstrations with natural language instructions. Given a long-horizon demonstration, the framework detects keystates and segments the demonstration into indivudal tasks, while generating language instructions.

Quickstart

To label your own demonstration videos, use the following command:

python annotate.py \
dataset.path=<path to folder containing videos> \
dataset.name=<name of the dataset> \
GPU_IDS=[0]<List of gpu_ids to use> \
PROC_PER_GPU=<Number of processes per GPU> \
embodiment_prompt_grounding_dino="the black robotic gripper" <Embodiment prompt for grounding dino> \
embodiment_prompt_clipseg=<Embodiment Prompt for ClipSeg>

This command will use the default settings. The framework can be tuned to specific datasets via configs. We will explain the process in more detail below.

Framework Usage

This section will describe how you can use and adapt the framework to annotate your own datasets.

Creating a Dataset

To create a new dataset, create a new torch.utils.data.Dataset class that inherits from torch.utils.data.Dataset. The class should have the following attributes:

  • paths: A list of paths to the folders containing the data. This will later be used to store the annotations.
  • name: The name of the dataset, used for logging.

The __getitem__ method should return a dictionary with the following keys:

  • rgb_static: np.ndarray containing the RGB images.
  • paths: Path to the currently processed data.
  • frame_names: Original frame indices. Will be used for saving the detected keystates and language annotations.
  • gripper_actions (optional): Binary gripper actions, if available.
  • keystates (optional): Frame indices of keystates, if available.

Incorporating prior knowledge

To improve the frameworks accuracy, you can incorporate prior knowledge, such as objects in the scene, available tasks or keystates of long-horizon demonstrations.

Objects

Create a list of objects in format

pot:silver
spoon:yellow
...

and specify the path to the file in the configuration file (object_list). By default, the framework will still check for objects in the scene and will not output objects with a high overlap with predefined objects. If you only want to use predefined objects, set only_predefined_object to True

Tasks

Create a list of tasks in format

place pot on stove
...

and specify the path to the file in the configuration file (task_list).

Keystates

If you want to load ground truth keystates, adjust the Dataset to return the keystates in the dictionary. The framework will then use the keystates to annotate the scene.

Config

We use Hydra as config manager. The main config file is located in conf/base.yaml.

You can control the number keystate heuristics by changing the entries in keystate_predictor and keystate_predictors_for_voting.

  • GPU_IDS What GPUs to use for labeling
  • PROC_PER_GPU Number of processes per GPU. Increase depending on your VRAM.
  • n_splits The framework currently has memory leaks, which is why we frequently have to reinitalize the framework. This setting divides the dataset into n_splits chunks which are processed in parallel. The number of processes used is GPU_IDS * PROC_PER_GPU

Repository Structure

  • nils/ contains the NILS codebase.
  • nils/annotator contains the main code of the framework and its components.
  • nils/my_datasets contain the datasets used in the paper.
  • nils/specialist_models contains the foundation models used to annotate the scene.
  • conf/ contains the configuration files for the framework.
  • scripts/ contains scripts to start annotation of datasets.
  • scripts/experiments contains the code used to conduct the experiments in the paper.

Installation

git clone --recurse-submodules git@github.com:intuitive-robots/NILS.git
pip install torch==2.3.1
conda create -n "NILS" python=3.9
conda activate NILS pip install -r requirements.txt
pip install -U openmim
mim install mmcv

The framework relies on several specialst models, which you need to install manually

DEVA

cd dependencies/Tracking-Anything-with-DEVA
pip install -e . --no-dependencies
bash scripts/download_models.sh

SAM2

cd nils/specialst_models/sam2/checkpoints
./download_ckpts.sh

DepthAnythingv2

cd dependencies/Depth-Anything-V2
mkdir checkpoints;cd checkpoints
wget https://huggingface.co/depth-anything/Depth-Anything-V2-Metric-Hypersim-Large/resolve/main/depth_anything_v2_metric_hypersim_vitl.pth?download=true -O depth_anything_v2_metric_hypersim_vitl.pth
export PYTHONPATH=$PYTHONPATH:path/to/Depth-Anything-V2/metric_depth

Unimatch (GMFLOW)

cd dependencies/unimatch
export PYTHONPATH=$PYTHONPATH:path/to/unimatch
mkdir pretrained;cd pretrained
wget https://s3.eu-central-1.amazonaws.com/avg-projects/unimatch/pretrained/gmflow-scale2-regrefine6-mixdata-train320x576-4e7b215d.pth

DINOv2

git clone https://github.com/facebookresearch/dinov2.git
export PYTHONPATH=$PYTHONPATH:path/to/dinov2

Citation

If you find our framework useful in your work, please cite our paper:

@inproceedings{
blank2024scaling,
title={Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models},
author={Nils Blank and Moritz Reuss and Marcel R{\"u}hle and {\"O}mer Erdin{\c{c}} Ya{\u{g}}murlu and Fabian Wenzel and Oier Mees and Rudolf Lioutikov},
booktitle={8th Annual Conference on Robot Learning},
year={2024},
url={https://openreview.net/forum?id=EdVNB2kHv1}
}

About

[CoRL 2024] Official code for "Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models"

Resources

Stars

36 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages