Repository files navigation

Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

This is the official repository for the paper: Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter.

Paper | Video

We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches learn end-to-end policies with features from vision foundation models, requiring large datasets. Others combine foundation models in a zero-shot setting, suffering from cascading errors. In addition, they primarily leverage vision and language foundation models, focusing less on action priors. In this paper, we aim to develop an effective policy by integrating foundation priors from vision, language, and action. We propose A2, an action prior alignment method that aligns unconditioned action priors with 3D vision-language priors by learning one attention layer. The alignment formulation enables our policy to train with less data and preserve zero-shot generalization capabilities. We show that a shared policy for both pick and place actions enhances the performance for each task, and introduce a policy adaptation scheme to accommodate the multi-modal nature of actions. Extensive experiments in simulation and the real-world show that our policy achieves higher task success rates with fewer steps for both pick and place tasks in clutter, effectively generalizing to unseen objects and language instructions.

system overview

Contact

Any question, please let me know: kcxu@zju.edu.cn

Setup

Installation

  • Ubuntu 20.04
  • Torch 1.10.1
  • Cuda 11.3
  • GTX 4090 is tested
git clone git@github.com:xukechun/Action-Prior-Alignment.git
cd Action-Prior-Alignment
conda create -n a2 python=3.8
conda activate a2
pip install -r requirements.txt
python setup.py develop
cd models/graspnet/pointnet2
python setup.py install
cd ../knn
python setup.py install

Potential Issues of Installation

  • When installing graspnetAPI, the following problem might occur:
× python setup.py egg_info did not run successfully.
│ exit code: 1
╰─> [18 lines of output]
The 'sklearn' PyPI package is deprecated, use 'scikit-learn'
rather than 'sklearn' for pip commands.

solution:

export SKLEARN_ALLOW_DEPRECATED_SKLEARN_PACKAGE_INSTALL=True
  • Check the compatible version of torch and torchvision of your machine (especially the cuda vision) if the following problem occurs:
RuntimeError: CUDA error: no kernel image is available for execution on the device

solution: to install torch with the right cuda version

Easy Installation

If you use conda, we provide our conda environment produced by conda-pack in this link. NOTE: This environment is compatiable with CUDA 11.3.

Then you can easily build and activate the conda environment by

cd 'PATH OF YOUR CONDA ENVS'
mkdir a2
tar -xzvf a2.tar.gz -C a2
conda activate a2
python setup.py develop
cd models/graspnet/pointnet2
python setup.py install
cd ../knn
python setup.py install

Assets

We provide the processed object models in this link. Please download the file and unzip it in the assets folder.

Data and Pre-trained Models

We provide our training data in this link. Please download the file and unzip it in the data folder.

We provide our testing cases in this link. Please download the file and unzip it in the testing_cases folder.

We provide our pre-trained models in this link. Please download the file and unzip it in the logs folder.

Data Collection

  • For pick data
bash scripts/data_collection/collect_data_grasp.sh
  • For place data
bash scripts/data_collection/collect_data_place.sh

Training

  • Unified training for pick and place
bash scripts/train/train_clutter_gp_unified.sh
  • Adaptation for place
bash scripts/train/train_clutter_gp_adaptive.sh

Evaluation

To test the pre-trained model, simply change the location of --model_path:

  • Pick
bash scripts/test/test_grasp.sh
  • Place
bash scripts/test/test_place.sh
  • Pick and place
bash scripts/test/test_pickplace.sh

Citation

If you find this work useful, please consider citing:

@article{xu2025efficient,
title={Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter},
author={Xu, Kechun and Xia, Xunlong and Wang, Kaixuan and Yang, Yifei and Mao, Yunxuan and Deng, Bing and Xiong, Rong and Wang, Yue},
journal={arXiv preprint arXiv:2503.09423},
year={2025}
}

About

[TASE 2025] Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

Resources

Stars

36 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

This is the official repository for the paper: Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter.

Paper | Video

We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches learn end-to-end policies with features from vision foundation models, requiring large datasets. Others combine foundation models in a zero-shot setting, suffering from cascading errors. In addition, they primarily leverage vision and language foundation models, focusing less on action priors. In this paper, we aim to develop an effective policy by integrating foundation priors from vision, language, and action. We propose A2, an action prior alignment method that aligns unconditioned action priors with 3D vision-language priors by learning one attention layer. The alignment formulation enables our policy to train with less data and preserve zero-shot generalization capabilities. We show that a shared policy for both pick and place actions enhances the performance for each task, and introduce a policy adaptation scheme to accommodate the multi-modal nature of actions. Extensive experiments in simulation and the real-world show that our policy achieves higher task success rates with fewer steps for both pick and place tasks in clutter, effectively generalizing to unseen objects and language instructions.

system overview

Contact

Any question, please let me know: kcxu@zju.edu.cn

Setup

Installation

  • Ubuntu 20.04
  • Torch 1.10.1
  • Cuda 11.3
  • GTX 4090 is tested
git clone git@github.com:xukechun/Action-Prior-Alignment.git
cd Action-Prior-Alignment
conda create -n a2 python=3.8
conda activate a2
pip install -r requirements.txt
python setup.py develop
cd models/graspnet/pointnet2
python setup.py install
cd ../knn
python setup.py install

Potential Issues of Installation

  • When installing graspnetAPI, the following problem might occur:
× python setup.py egg_info did not run successfully.
│ exit code: 1
╰─> [18 lines of output]
The 'sklearn' PyPI package is deprecated, use 'scikit-learn'
rather than 'sklearn' for pip commands.

solution:

export SKLEARN_ALLOW_DEPRECATED_SKLEARN_PACKAGE_INSTALL=True
  • Check the compatible version of torch and torchvision of your machine (especially the cuda vision) if the following problem occurs:
RuntimeError: CUDA error: no kernel image is available for execution on the device

solution: to install torch with the right cuda version

Easy Installation

If you use conda, we provide our conda environment produced by conda-pack in this link. NOTE: This environment is compatiable with CUDA 11.3.

Then you can easily build and activate the conda environment by

cd 'PATH OF YOUR CONDA ENVS'
mkdir a2
tar -xzvf a2.tar.gz -C a2
conda activate a2
python setup.py develop
cd models/graspnet/pointnet2
python setup.py install
cd ../knn
python setup.py install

Assets

We provide the processed object models in this link. Please download the file and unzip it in the assets folder.

Data and Pre-trained Models

We provide our training data in this link. Please download the file and unzip it in the data folder.

We provide our testing cases in this link. Please download the file and unzip it in the testing_cases folder.

We provide our pre-trained models in this link. Please download the file and unzip it in the logs folder.

Data Collection

  • For pick data
bash scripts/data_collection/collect_data_grasp.sh
  • For place data
bash scripts/data_collection/collect_data_place.sh

Training

  • Unified training for pick and place
bash scripts/train/train_clutter_gp_unified.sh
  • Adaptation for place
bash scripts/train/train_clutter_gp_adaptive.sh

Evaluation

To test the pre-trained model, simply change the location of --model_path:

  • Pick
bash scripts/test/test_grasp.sh
  • Place
bash scripts/test/test_place.sh
  • Pick and place
bash scripts/test/test_pickplace.sh

Citation

If you find this work useful, please consider citing:

@article{xu2025efficient,
title={Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter},
author={Xu, Kechun and Xia, Xunlong and Wang, Kaixuan and Yang, Yifei and Mao, Yunxuan and Deng, Bing and Xiong, Rong and Wang, Yue},
journal={arXiv preprint arXiv:2503.09423},
year={2025}
}

About

[TASE 2025] Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

Resources

Stars

36 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

This is the official repository for the paper: Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter.

Paper | Video

We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches learn end-to-end policies with features from vision foundation models, requiring large datasets. Others combine foundation models in a zero-shot setting, suffering from cascading errors. In addition, they primarily leverage vision and language foundation models, focusing less on action priors. In this paper, we aim to develop an effective policy by integrating foundation priors from vision, language, and action. We propose A2, an action prior alignment method that aligns unconditioned action priors with 3D vision-language priors by learning one attention layer. The alignment formulation enables our policy to train with less data and preserve zero-shot generalization capabilities. We show that a shared policy for both pick and place actions enhances the performance for each task, and introduce a policy adaptation scheme to accommodate the multi-modal nature of actions. Extensive experiments in simulation and the real-world show that our policy achieves higher task success rates with fewer steps for both pick and place tasks in clutter, effectively generalizing to unseen objects and language instructions.

system overview

Contact

Any question, please let me know: kcxu@zju.edu.cn

Setup

Installation

  • Ubuntu 20.04
  • Torch 1.10.1
  • Cuda 11.3
  • GTX 4090 is tested
git clone git@github.com:xukechun/Action-Prior-Alignment.git
cd Action-Prior-Alignment
conda create -n a2 python=3.8
conda activate a2
pip install -r requirements.txt
python setup.py develop
cd models/graspnet/pointnet2
python setup.py install
cd ../knn
python setup.py install

Potential Issues of Installation

  • When installing graspnetAPI, the following problem might occur:
× python setup.py egg_info did not run successfully.
│ exit code: 1
╰─> [18 lines of output]
The 'sklearn' PyPI package is deprecated, use 'scikit-learn'
rather than 'sklearn' for pip commands.

solution:

export SKLEARN_ALLOW_DEPRECATED_SKLEARN_PACKAGE_INSTALL=True
  • Check the compatible version of torch and torchvision of your machine (especially the cuda vision) if the following problem occurs:
RuntimeError: CUDA error: no kernel image is available for execution on the device

solution: to install torch with the right cuda version

Easy Installation

If you use conda, we provide our conda environment produced by conda-pack in this link. NOTE: This environment is compatiable with CUDA 11.3.

Then you can easily build and activate the conda environment by

cd 'PATH OF YOUR CONDA ENVS'
mkdir a2
tar -xzvf a2.tar.gz -C a2
conda activate a2
python setup.py develop
cd models/graspnet/pointnet2
python setup.py install
cd ../knn
python setup.py install

Assets

We provide the processed object models in this link. Please download the file and unzip it in the assets folder.

Data and Pre-trained Models

We provide our training data in this link. Please download the file and unzip it in the data folder.

We provide our testing cases in this link. Please download the file and unzip it in the testing_cases folder.

We provide our pre-trained models in this link. Please download the file and unzip it in the logs folder.

Data Collection

  • For pick data
bash scripts/data_collection/collect_data_grasp.sh
  • For place data
bash scripts/data_collection/collect_data_place.sh

Training

  • Unified training for pick and place
bash scripts/train/train_clutter_gp_unified.sh
  • Adaptation for place
bash scripts/train/train_clutter_gp_adaptive.sh

Evaluation

To test the pre-trained model, simply change the location of --model_path:

  • Pick
bash scripts/test/test_grasp.sh
  • Place
bash scripts/test/test_place.sh
  • Pick and place
bash scripts/test/test_pickplace.sh

Citation

If you find this work useful, please consider citing:

@article{xu2025efficient,
title={Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter},
author={Xu, Kechun and Xia, Xunlong and Wang, Kaixuan and Yang, Yifei and Mao, Yunxuan and Deng, Bing and Xiong, Rong and Wang, Yue},
journal={arXiv preprint arXiv:2503.09423},
year={2025}
}

About

[TASE 2025] Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

Resources

Stars

36 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

This is the official repository for the paper: Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter.

Paper | Video

We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches learn end-to-end policies with features from vision foundation models, requiring large datasets. Others combine foundation models in a zero-shot setting, suffering from cascading errors. In addition, they primarily leverage vision and language foundation models, focusing less on action priors. In this paper, we aim to develop an effective policy by integrating foundation priors from vision, language, and action. We propose A2, an action prior alignment method that aligns unconditioned action priors with 3D vision-language priors by learning one attention layer. The alignment formulation enables our policy to train with less data and preserve zero-shot generalization capabilities. We show that a shared policy for both pick and place actions enhances the performance for each task, and introduce a policy adaptation scheme to accommodate the multi-modal nature of actions. Extensive experiments in simulation and the real-world show that our policy achieves higher task success rates with fewer steps for both pick and place tasks in clutter, effectively generalizing to unseen objects and language instructions.

system overview

Contact

Any question, please let me know: kcxu@zju.edu.cn

Setup

Installation

  • Ubuntu 20.04
  • Torch 1.10.1
  • Cuda 11.3
  • GTX 4090 is tested
git clone git@github.com:xukechun/Action-Prior-Alignment.git
cd Action-Prior-Alignment
conda create -n a2 python=3.8
conda activate a2
pip install -r requirements.txt
python setup.py develop
cd models/graspnet/pointnet2
python setup.py install
cd ../knn
python setup.py install

Potential Issues of Installation

  • When installing graspnetAPI, the following problem might occur:
× python setup.py egg_info did not run successfully.
│ exit code: 1
╰─> [18 lines of output]
The 'sklearn' PyPI package is deprecated, use 'scikit-learn'
rather than 'sklearn' for pip commands.

solution:

export SKLEARN_ALLOW_DEPRECATED_SKLEARN_PACKAGE_INSTALL=True
  • Check the compatible version of torch and torchvision of your machine (especially the cuda vision) if the following problem occurs:
RuntimeError: CUDA error: no kernel image is available for execution on the device

solution: to install torch with the right cuda version

Easy Installation

If you use conda, we provide our conda environment produced by conda-pack in this link. NOTE: This environment is compatiable with CUDA 11.3.

Then you can easily build and activate the conda environment by

cd 'PATH OF YOUR CONDA ENVS'
mkdir a2
tar -xzvf a2.tar.gz -C a2
conda activate a2
python setup.py develop
cd models/graspnet/pointnet2
python setup.py install
cd ../knn
python setup.py install

Assets

We provide the processed object models in this link. Please download the file and unzip it in the assets folder.

Data and Pre-trained Models

We provide our training data in this link. Please download the file and unzip it in the data folder.

We provide our testing cases in this link. Please download the file and unzip it in the testing_cases folder.

We provide our pre-trained models in this link. Please download the file and unzip it in the logs folder.

Data Collection

  • For pick data
bash scripts/data_collection/collect_data_grasp.sh
  • For place data
bash scripts/data_collection/collect_data_place.sh

Training

  • Unified training for pick and place
bash scripts/train/train_clutter_gp_unified.sh
  • Adaptation for place
bash scripts/train/train_clutter_gp_adaptive.sh

Evaluation

To test the pre-trained model, simply change the location of --model_path:

  • Pick
bash scripts/test/test_grasp.sh
  • Place
bash scripts/test/test_place.sh
  • Pick and place
bash scripts/test/test_pickplace.sh

Citation

If you find this work useful, please consider citing:

@article{xu2025efficient,
title={Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter},
author={Xu, Kechun and Xia, Xunlong and Wang, Kaixuan and Yang, Yifei and Mao, Yunxuan and Deng, Bing and Xiong, Rong and Wang, Yue},
journal={arXiv preprint arXiv:2503.09423},
year={2025}
}

About

[TASE 2025] Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

Resources

Stars

36 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

This is the official repository for the paper: Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter.

Paper | Video

We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches learn end-to-end policies with features from vision foundation models, requiring large datasets. Others combine foundation models in a zero-shot setting, suffering from cascading errors. In addition, they primarily leverage vision and language foundation models, focusing less on action priors. In this paper, we aim to develop an effective policy by integrating foundation priors from vision, language, and action. We propose A2, an action prior alignment method that aligns unconditioned action priors with 3D vision-language priors by learning one attention layer. The alignment formulation enables our policy to train with less data and preserve zero-shot generalization capabilities. We show that a shared policy for both pick and place actions enhances the performance for each task, and introduce a policy adaptation scheme to accommodate the multi-modal nature of actions. Extensive experiments in simulation and the real-world show that our policy achieves higher task success rates with fewer steps for both pick and place tasks in clutter, effectively generalizing to unseen objects and language instructions.

system overview

Contact

Any question, please let me know: kcxu@zju.edu.cn

Setup

Installation

  • Ubuntu 20.04
  • Torch 1.10.1
  • Cuda 11.3
  • GTX 4090 is tested
git clone git@github.com:xukechun/Action-Prior-Alignment.git
cd Action-Prior-Alignment
conda create -n a2 python=3.8
conda activate a2
pip install -r requirements.txt
python setup.py develop
cd models/graspnet/pointnet2
python setup.py install
cd ../knn
python setup.py install

Potential Issues of Installation

  • When installing graspnetAPI, the following problem might occur:
× python setup.py egg_info did not run successfully.
│ exit code: 1
╰─> [18 lines of output]
The 'sklearn' PyPI package is deprecated, use 'scikit-learn'
rather than 'sklearn' for pip commands.

solution:

export SKLEARN_ALLOW_DEPRECATED_SKLEARN_PACKAGE_INSTALL=True
  • Check the compatible version of torch and torchvision of your machine (especially the cuda vision) if the following problem occurs:
RuntimeError: CUDA error: no kernel image is available for execution on the device

solution: to install torch with the right cuda version

Easy Installation

If you use conda, we provide our conda environment produced by conda-pack in this link. NOTE: This environment is compatiable with CUDA 11.3.

Then you can easily build and activate the conda environment by

cd 'PATH OF YOUR CONDA ENVS'
mkdir a2
tar -xzvf a2.tar.gz -C a2
conda activate a2
python setup.py develop
cd models/graspnet/pointnet2
python setup.py install
cd ../knn
python setup.py install

Assets

We provide the processed object models in this link. Please download the file and unzip it in the assets folder.

Data and Pre-trained Models

We provide our training data in this link. Please download the file and unzip it in the data folder.

We provide our testing cases in this link. Please download the file and unzip it in the testing_cases folder.

We provide our pre-trained models in this link. Please download the file and unzip it in the logs folder.

Data Collection

  • For pick data
bash scripts/data_collection/collect_data_grasp.sh
  • For place data
bash scripts/data_collection/collect_data_place.sh

Training

  • Unified training for pick and place
bash scripts/train/train_clutter_gp_unified.sh
  • Adaptation for place
bash scripts/train/train_clutter_gp_adaptive.sh

Evaluation

To test the pre-trained model, simply change the location of --model_path:

  • Pick
bash scripts/test/test_grasp.sh
  • Place
bash scripts/test/test_place.sh
  • Pick and place
bash scripts/test/test_pickplace.sh

Citation

If you find this work useful, please consider citing:

@article{xu2025efficient,
title={Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter},
author={Xu, Kechun and Xia, Xunlong and Wang, Kaixuan and Yang, Yifei and Mao, Yunxuan and Deng, Bing and Xiong, Rong and Wang, Yue},
journal={arXiv preprint arXiv:2503.09423},
year={2025}
}

About

[TASE 2025] Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

Resources

Stars

36 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

This is the official repository for the paper: Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter.

Paper | Video

We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches learn end-to-end policies with features from vision foundation models, requiring large datasets. Others combine foundation models in a zero-shot setting, suffering from cascading errors. In addition, they primarily leverage vision and language foundation models, focusing less on action priors. In this paper, we aim to develop an effective policy by integrating foundation priors from vision, language, and action. We propose A2, an action prior alignment method that aligns unconditioned action priors with 3D vision-language priors by learning one attention layer. The alignment formulation enables our policy to train with less data and preserve zero-shot generalization capabilities. We show that a shared policy for both pick and place actions enhances the performance for each task, and introduce a policy adaptation scheme to accommodate the multi-modal nature of actions. Extensive experiments in simulation and the real-world show that our policy achieves higher task success rates with fewer steps for both pick and place tasks in clutter, effectively generalizing to unseen objects and language instructions.

system overview

Contact

Any question, please let me know: kcxu@zju.edu.cn

Setup

Installation

  • Ubuntu 20.04
  • Torch 1.10.1
  • Cuda 11.3
  • GTX 4090 is tested
git clone git@github.com:xukechun/Action-Prior-Alignment.git
cd Action-Prior-Alignment
conda create -n a2 python=3.8
conda activate a2
pip install -r requirements.txt
python setup.py develop
cd models/graspnet/pointnet2
python setup.py install
cd ../knn
python setup.py install

Potential Issues of Installation

  • When installing graspnetAPI, the following problem might occur:
× python setup.py egg_info did not run successfully.
│ exit code: 1
╰─> [18 lines of output]
The 'sklearn' PyPI package is deprecated, use 'scikit-learn'
rather than 'sklearn' for pip commands.

solution:

export SKLEARN_ALLOW_DEPRECATED_SKLEARN_PACKAGE_INSTALL=True
  • Check the compatible version of torch and torchvision of your machine (especially the cuda vision) if the following problem occurs:
RuntimeError: CUDA error: no kernel image is available for execution on the device

solution: to install torch with the right cuda version

Easy Installation

If you use conda, we provide our conda environment produced by conda-pack in this link. NOTE: This environment is compatiable with CUDA 11.3.

Then you can easily build and activate the conda environment by

cd 'PATH OF YOUR CONDA ENVS'
mkdir a2
tar -xzvf a2.tar.gz -C a2
conda activate a2
python setup.py develop
cd models/graspnet/pointnet2
python setup.py install
cd ../knn
python setup.py install

Assets

We provide the processed object models in this link. Please download the file and unzip it in the assets folder.

Data and Pre-trained Models

We provide our training data in this link. Please download the file and unzip it in the data folder.

We provide our testing cases in this link. Please download the file and unzip it in the testing_cases folder.

We provide our pre-trained models in this link. Please download the file and unzip it in the logs folder.

Data Collection

  • For pick data
bash scripts/data_collection/collect_data_grasp.sh
  • For place data
bash scripts/data_collection/collect_data_place.sh

Training

  • Unified training for pick and place
bash scripts/train/train_clutter_gp_unified.sh
  • Adaptation for place
bash scripts/train/train_clutter_gp_adaptive.sh

Evaluation

To test the pre-trained model, simply change the location of --model_path:

  • Pick
bash scripts/test/test_grasp.sh
  • Place
bash scripts/test/test_place.sh
  • Pick and place
bash scripts/test/test_pickplace.sh

Citation

If you find this work useful, please consider citing:

@article{xu2025efficient,
title={Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter},
author={Xu, Kechun and Xia, Xunlong and Wang, Kaixuan and Yang, Yifei and Mao, Yunxuan and Deng, Bing and Xiong, Rong and Wang, Yue},
journal={arXiv preprint arXiv:2503.09423},
year={2025}
}

About

[TASE 2025] Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

Resources

Stars

36 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

This is the official repository for the paper: Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter.

Paper | Video

We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches learn end-to-end policies with features from vision foundation models, requiring large datasets. Others combine foundation models in a zero-shot setting, suffering from cascading errors. In addition, they primarily leverage vision and language foundation models, focusing less on action priors. In this paper, we aim to develop an effective policy by integrating foundation priors from vision, language, and action. We propose A2, an action prior alignment method that aligns unconditioned action priors with 3D vision-language priors by learning one attention layer. The alignment formulation enables our policy to train with less data and preserve zero-shot generalization capabilities. We show that a shared policy for both pick and place actions enhances the performance for each task, and introduce a policy adaptation scheme to accommodate the multi-modal nature of actions. Extensive experiments in simulation and the real-world show that our policy achieves higher task success rates with fewer steps for both pick and place tasks in clutter, effectively generalizing to unseen objects and language instructions.

system overview

Contact

Any question, please let me know: kcxu@zju.edu.cn

Setup

Installation

  • Ubuntu 20.04
  • Torch 1.10.1
  • Cuda 11.3
  • GTX 4090 is tested
git clone git@github.com:xukechun/Action-Prior-Alignment.git
cd Action-Prior-Alignment
conda create -n a2 python=3.8
conda activate a2
pip install -r requirements.txt
python setup.py develop
cd models/graspnet/pointnet2
python setup.py install
cd ../knn
python setup.py install

Potential Issues of Installation

  • When installing graspnetAPI, the following problem might occur:
× python setup.py egg_info did not run successfully.
│ exit code: 1
╰─> [18 lines of output]
The 'sklearn' PyPI package is deprecated, use 'scikit-learn'
rather than 'sklearn' for pip commands.

solution:

export SKLEARN_ALLOW_DEPRECATED_SKLEARN_PACKAGE_INSTALL=True
  • Check the compatible version of torch and torchvision of your machine (especially the cuda vision) if the following problem occurs:
RuntimeError: CUDA error: no kernel image is available for execution on the device

solution: to install torch with the right cuda version

Easy Installation

If you use conda, we provide our conda environment produced by conda-pack in this link. NOTE: This environment is compatiable with CUDA 11.3.

Then you can easily build and activate the conda environment by

cd 'PATH OF YOUR CONDA ENVS'
mkdir a2
tar -xzvf a2.tar.gz -C a2
conda activate a2
python setup.py develop
cd models/graspnet/pointnet2
python setup.py install
cd ../knn
python setup.py install

Assets

We provide the processed object models in this link. Please download the file and unzip it in the assets folder.

Data and Pre-trained Models

We provide our training data in this link. Please download the file and unzip it in the data folder.

We provide our testing cases in this link. Please download the file and unzip it in the testing_cases folder.

We provide our pre-trained models in this link. Please download the file and unzip it in the logs folder.

Data Collection

  • For pick data
bash scripts/data_collection/collect_data_grasp.sh
  • For place data
bash scripts/data_collection/collect_data_place.sh

Training

  • Unified training for pick and place
bash scripts/train/train_clutter_gp_unified.sh
  • Adaptation for place
bash scripts/train/train_clutter_gp_adaptive.sh

Evaluation

To test the pre-trained model, simply change the location of --model_path:

  • Pick
bash scripts/test/test_grasp.sh
  • Place
bash scripts/test/test_place.sh
  • Pick and place
bash scripts/test/test_pickplace.sh

Citation

If you find this work useful, please consider citing:

@article{xu2025efficient,
title={Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter},
author={Xu, Kechun and Xia, Xunlong and Wang, Kaixuan and Yang, Yifei and Mao, Yunxuan and Deng, Bing and Xiong, Rong and Wang, Yue},
journal={arXiv preprint arXiv:2503.09423},
year={2025}
}

About

[TASE 2025] Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

Resources

Stars

36 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

This is the official repository for the paper: Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter.

Paper | Video

We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches learn end-to-end policies with features from vision foundation models, requiring large datasets. Others combine foundation models in a zero-shot setting, suffering from cascading errors. In addition, they primarily leverage vision and language foundation models, focusing less on action priors. In this paper, we aim to develop an effective policy by integrating foundation priors from vision, language, and action. We propose A2, an action prior alignment method that aligns unconditioned action priors with 3D vision-language priors by learning one attention layer. The alignment formulation enables our policy to train with less data and preserve zero-shot generalization capabilities. We show that a shared policy for both pick and place actions enhances the performance for each task, and introduce a policy adaptation scheme to accommodate the multi-modal nature of actions. Extensive experiments in simulation and the real-world show that our policy achieves higher task success rates with fewer steps for both pick and place tasks in clutter, effectively generalizing to unseen objects and language instructions.

system overview

Contact

Any question, please let me know: kcxu@zju.edu.cn

Setup

Installation

  • Ubuntu 20.04
  • Torch 1.10.1
  • Cuda 11.3
  • GTX 4090 is tested
git clone git@github.com:xukechun/Action-Prior-Alignment.git
cd Action-Prior-Alignment
conda create -n a2 python=3.8
conda activate a2
pip install -r requirements.txt
python setup.py develop
cd models/graspnet/pointnet2
python setup.py install
cd ../knn
python setup.py install

Potential Issues of Installation

  • When installing graspnetAPI, the following problem might occur:
× python setup.py egg_info did not run successfully.
│ exit code: 1
╰─> [18 lines of output]
The 'sklearn' PyPI package is deprecated, use 'scikit-learn'
rather than 'sklearn' for pip commands.

solution:

export SKLEARN_ALLOW_DEPRECATED_SKLEARN_PACKAGE_INSTALL=True
  • Check the compatible version of torch and torchvision of your machine (especially the cuda vision) if the following problem occurs:
RuntimeError: CUDA error: no kernel image is available for execution on the device

solution: to install torch with the right cuda version

Easy Installation

If you use conda, we provide our conda environment produced by conda-pack in this link. NOTE: This environment is compatiable with CUDA 11.3.

Then you can easily build and activate the conda environment by

cd 'PATH OF YOUR CONDA ENVS'
mkdir a2
tar -xzvf a2.tar.gz -C a2
conda activate a2
python setup.py develop
cd models/graspnet/pointnet2
python setup.py install
cd ../knn
python setup.py install

Assets

We provide the processed object models in this link. Please download the file and unzip it in the assets folder.

Data and Pre-trained Models

We provide our training data in this link. Please download the file and unzip it in the data folder.

We provide our testing cases in this link. Please download the file and unzip it in the testing_cases folder.

We provide our pre-trained models in this link. Please download the file and unzip it in the logs folder.

Data Collection

  • For pick data
bash scripts/data_collection/collect_data_grasp.sh
  • For place data
bash scripts/data_collection/collect_data_place.sh

Training

  • Unified training for pick and place
bash scripts/train/train_clutter_gp_unified.sh
  • Adaptation for place
bash scripts/train/train_clutter_gp_adaptive.sh

Evaluation

To test the pre-trained model, simply change the location of --model_path:

  • Pick
bash scripts/test/test_grasp.sh
  • Place
bash scripts/test/test_place.sh
  • Pick and place
bash scripts/test/test_pickplace.sh

Citation

If you find this work useful, please consider citing:

@article{xu2025efficient,
title={Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter},
author={Xu, Kechun and Xia, Xunlong and Wang, Kaixuan and Yang, Yifei and Mao, Yunxuan and Deng, Bing and Xiong, Rong and Wang, Yue},
journal={arXiv preprint arXiv:2503.09423},
year={2025}
}

About

[TASE 2025] Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

Resources

Stars

36 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages