Repository files navigation

HumanSD


This repository contains the implementation of the ICCV2023 paper:

HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation[Project Page][Paper][Code][Video][Data]
Xuan Ju∗12, Ailing Zeng∗1, Chenchen Zhao∗2, Jianan Wang1, Lei Zhang1, Qiang Xu2
Equal contribution 1International Digital Economy Academy 2The Chinese University of Hong Kong

In this work, we propose a native skeleton-guided diffusion model for controllable HIG called HumanSD. Instead of performing image editing with dual-branch diffusion, we fine-tune the original SD model using a novel heatmap-guided denoising loss. This strategy effectively and efficiently strengthens the given skeleton condition during model training while mitigating the catastrophic forgetting effects. HumanSD is fine-tuned on the assembly of three large-scale human-centric datasets with text-imagepose information, two of which are established in this work.


  • (a) a generation by the pre-trained pose-less text-guided stable diffusion (SD)
  • (b) pose skeleton images as the condition to ControlNet and our proposed HumanSD
  • (c) a generation by ControlNet
  • (d) a generation by HumanSD (ours). ControlNet and HumanSD receive both text and pose conditions.

HumanSD shows its superiorities in terms of (I) challenging poses, (II) accurate painting styles, (III) pose control capability, (IV) multi-person scenarios, and (V) delicate details.

Table of Contents

TODO

News!! Our paper have been accepted by ICCV2023! Training code is released.

  • Release inference code and pretrained models
  • Release Gradio UI demo
  • Public training data (LAION-Human)
  • Release training code

Model Overview

Getting Started

Environment Requirement

HumanSD has been implemented and tested on Pytorch 1.12.1 with python 3.9.

Clone the repo:

git clone git@github.com:IDEA-Research/HumanSD.git

We recommend you first install pytorch following official instructions. For example:

# conda
conda install pytorch==1.12.1 torchvision==0.13.1 torchaudio==0.12.1 cudatoolkit=11.3 -c pytorch

Then, you can install required packages thourgh:

pip install -r requirements.txt

You also need to install MMPose following here. Noted that you only need to install MMPose as a python package. PS: Because of the update of MMPose, we recommend you to install 0.29.0 version of MMPose.

Model and Checkpoints

Download necessary checkpoints of HumanSD, which can be found here. The data structure should be like:

|-- humansd_data
|-- checkpoints
|-- higherhrnet_w48_humanart_512x512_udp.pth
|-- v2-1_512-ema-pruned.ckpt
|-- humansd-v1.ckpt

Noted that v2-1_512-ema-pruned.ckpt should be download from Stable Diffusion.

Quick Demo

You can run demo either through command line or gradio.

You can run demo through command line with:

python scripts/pose2img.py --prompt "oil painting of girls dancing on the stage" --pose_file assets/pose/demo.npz

You can also run demo compared with ControlNet and T2I-Adapter:

python scripts/pose2img.py --prompt "oil painting of girls dancing on the stage" --pose_file assets/pose/demo.npz --controlnet --t2i

You can run gradio demo through:

python scripts/gradio/pose2img.py

We have also provided the comparison of ControlNet and T2I-Adapter, you can run all these methods in one demo. But you need to download corresponding model and checkpoints following:

To compare ControlNet, and T2I-Adpater's results. (1) You need to initialize ControlNet and T2I-Adapter as submodule using
git submodule init
git submodule update

(2) Then download checkpoints from: a. T2I-Adapter b. ControlNet. And put them into humansd_data/checkpoints

Then, run:

python scripts/gradio/pose2img.py --controlnet --t2i

Noted that you may have to modify some code in T2I-Adapter due to the path conflict.

e.g., use

from comparison_models.T2IAdapter.ldm.models.diffusion.ddim import DDIMSampler

instead of

from T2IAdapter.ldm.models.diffusion.ddim import DDIMSampler

Dataset

You may refer to the code here for loading the data.

Laion-Human

You may apply for access of Laion-Human here. Noted that we have provide the pose annotations, images' .parquet file and mapping file, please download the images according to .parquet. The key in .parquet is the corresponding image index. For example, image with key=338717 in 00033.parquet is corresponding to images/00000/000338717.jpg.

After downloading the images and pose, you need to extract zip files and make it looks like:

|-- humansd_data
|-- datasets
|-- Laion |-- Aesthetics_Human
|-- images
|-- 00000.parquet
|-- 00001.parquet
|-- ...
|-- pose
|-- 00000
|-- 000000000.npz
|-- 000000001.npz
|-- ...
|-- 00001
|-- ... |-- mapping_file_training.json 

Then, you can use python utils/download_data.py to download all images.

Then, the file data structure should be like:

|-- humansd_data
|-- datasets
|-- Laion |-- Aesthetics_Human
|-- images
|-- 00000.parquet
|-- 00001.parquet
|-- ...
|-- 00000
|-- 000000000.jpg
|-- 000000001.jpg
|-- ...
|-- 00001
|-- ...
|-- pose
|-- 00000
|-- 000000000.npz
|-- 000000001.npz
|-- ...
|-- 00001
|-- ... |-- mapping_file_training.json 

If you download the LAION-Aesthetics in tar files, which is different from our data structure, we recommend you extract the tar file through code:

importtarfiletar_file="00000.tar"# 00000.tar - 00286.tarpresent_tar_path=f"xxxxxx/{tar_file}"save_dir="humansd_data/datasets/Laion/Aesthetics_Human/images"withtarfile.open(present_tar_path, "r") astar_file:
forpresent_fileintar_file.getmembers():
ifpresent_file.name.endswith(".jpg"):
print(f" image:- {present_file.name} -")
image_save_path=os.path.join(save_dir,tar_file.replace(".tar",""),present_file.name)
present_image_fp=TarIO.TarIO(present_tar_path, present_file.name)
present_image=Image.open(present_image_fp)
present_image_numpy=cv2.cvtColor(np.array(present_image),cv2.COLOR_RGB2BGR)
ifnotos.path.exists(os.path.dirname(image_save_path)):
os.makedirs(os.path.dirname(image_save_path))
cv2.imwrite(image_save_path,present_image_numpy)

Human-Art

You may download Human-Art dataset here.

The file data structure should be like:

|-- humansd_data
|-- datasets
|-- HumanArt |-- images
|-- 2D_virtual_human
|-- cartoon
|-- 000000000007.jpg
|-- 000000000019.jpg
|-- ...
|-- digital_art
|-- ...
|-- 3D_virtual_human
|-- real_human
|-- pose
|-- 2D_virtual_human
|-- cartoon
|-- 000000000007.npz
|-- 000000000019.npz
|-- ...
|-- digital_art
|-- ...
|-- 3D_virtual_human
|-- real_human
|-- mapping_file_training.json |-- mapping_file_validation.json 

Training

Note that the datasets and checkpoints should be downloaded and prepared before training.

Run the commands below to start training:

python main.py --base configs/humansd/humansd-finetune.yaml -t --gpus 0,1 --name finetune_humansd

If you want to finetune without heat-map-guided diffusion loss for ablation, you can run the following commands:

python main.py --base configs/humansd/humansd-finetune-originalloss.yaml -t --gpus 0,1 --name finetune_humansd_original_loss

Quantitative Results

Metrics can be calculated through:

python scripts/pose2img_metrics.py --outdir outputs/metrics --config utils/metrics/metrics.yaml --ckpt path_to_ckpt

Qualitative Results

  • (a) a generation by the pre-trained text-guided stable diffusion (SD)
  • (b) pose skeleton images as the condition to ControlNet, T2I-Adapter and our proposed HumanSD
  • (c) a generation by ControlNet
  • (d) a generation by T2I-Adapter
  • (e) a generation by HumanSD (ours).

ControlNet, T2I-Adapter, and HumanSD receive both text and pose conditions.

Natural Scene

Sketch Scene

Shadow Play Scene

Children Drawing Scene

Oil Painting Scene

Watercolor Scene

Digital Art Scene

Relief Scene

Sculpture Scene

Cite Us

@article{ju2023humansd,
title={Human{SD}: A Native Skeleton-Guided Diffusion Model for Human Image Generation},
author={Ju, Xuan and Zeng, Ailing and Zhao, Chenchen and Wang, Jianan and Zhang, Lei and Xu, Qiang},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
year={2023}
}
@inproceedings{ju2023human,
title={Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial Scenes},
author={Ju, Xuan and Zeng, Ailing and Wang, Jianan and Xu, Qiang and Zhang, Lei},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year={2023},
}

Acknowledgement

About

[ICCV 2023] The official implementation of paper "HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation"

Topics

Resources

Stars

306 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

HumanSD


This repository contains the implementation of the ICCV2023 paper:

HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation[Project Page][Paper][Code][Video][Data]
Xuan Ju∗12, Ailing Zeng∗1, Chenchen Zhao∗2, Jianan Wang1, Lei Zhang1, Qiang Xu2
Equal contribution 1International Digital Economy Academy 2The Chinese University of Hong Kong

In this work, we propose a native skeleton-guided diffusion model for controllable HIG called HumanSD. Instead of performing image editing with dual-branch diffusion, we fine-tune the original SD model using a novel heatmap-guided denoising loss. This strategy effectively and efficiently strengthens the given skeleton condition during model training while mitigating the catastrophic forgetting effects. HumanSD is fine-tuned on the assembly of three large-scale human-centric datasets with text-imagepose information, two of which are established in this work.


  • (a) a generation by the pre-trained pose-less text-guided stable diffusion (SD)
  • (b) pose skeleton images as the condition to ControlNet and our proposed HumanSD
  • (c) a generation by ControlNet
  • (d) a generation by HumanSD (ours). ControlNet and HumanSD receive both text and pose conditions.

HumanSD shows its superiorities in terms of (I) challenging poses, (II) accurate painting styles, (III) pose control capability, (IV) multi-person scenarios, and (V) delicate details.

Table of Contents

TODO

News!! Our paper have been accepted by ICCV2023! Training code is released.

  • Release inference code and pretrained models
  • Release Gradio UI demo
  • Public training data (LAION-Human)
  • Release training code

Model Overview

Getting Started

Environment Requirement

HumanSD has been implemented and tested on Pytorch 1.12.1 with python 3.9.

Clone the repo:

git clone git@github.com:IDEA-Research/HumanSD.git

We recommend you first install pytorch following official instructions. For example:

# conda
conda install pytorch==1.12.1 torchvision==0.13.1 torchaudio==0.12.1 cudatoolkit=11.3 -c pytorch

Then, you can install required packages thourgh:

pip install -r requirements.txt

You also need to install MMPose following here. Noted that you only need to install MMPose as a python package. PS: Because of the update of MMPose, we recommend you to install 0.29.0 version of MMPose.

Model and Checkpoints

Download necessary checkpoints of HumanSD, which can be found here. The data structure should be like:

|-- humansd_data
|-- checkpoints
|-- higherhrnet_w48_humanart_512x512_udp.pth
|-- v2-1_512-ema-pruned.ckpt
|-- humansd-v1.ckpt

Noted that v2-1_512-ema-pruned.ckpt should be download from Stable Diffusion.

Quick Demo

You can run demo either through command line or gradio.

You can run demo through command line with:

python scripts/pose2img.py --prompt "oil painting of girls dancing on the stage" --pose_file assets/pose/demo.npz

You can also run demo compared with ControlNet and T2I-Adapter:

python scripts/pose2img.py --prompt "oil painting of girls dancing on the stage" --pose_file assets/pose/demo.npz --controlnet --t2i

You can run gradio demo through:

python scripts/gradio/pose2img.py

We have also provided the comparison of ControlNet and T2I-Adapter, you can run all these methods in one demo. But you need to download corresponding model and checkpoints following:

To compare ControlNet, and T2I-Adpater's results. (1) You need to initialize ControlNet and T2I-Adapter as submodule using
git submodule init
git submodule update

(2) Then download checkpoints from: a. T2I-Adapter b. ControlNet. And put them into humansd_data/checkpoints

Then, run:

python scripts/gradio/pose2img.py --controlnet --t2i

Noted that you may have to modify some code in T2I-Adapter due to the path conflict.

e.g., use

from comparison_models.T2IAdapter.ldm.models.diffusion.ddim import DDIMSampler

instead of

from T2IAdapter.ldm.models.diffusion.ddim import DDIMSampler

Dataset

You may refer to the code here for loading the data.

Laion-Human

You may apply for access of Laion-Human here. Noted that we have provide the pose annotations, images' .parquet file and mapping file, please download the images according to .parquet. The key in .parquet is the corresponding image index. For example, image with key=338717 in 00033.parquet is corresponding to images/00000/000338717.jpg.

After downloading the images and pose, you need to extract zip files and make it looks like:

|-- humansd_data
|-- datasets
|-- Laion |-- Aesthetics_Human
|-- images
|-- 00000.parquet
|-- 00001.parquet
|-- ...
|-- pose
|-- 00000
|-- 000000000.npz
|-- 000000001.npz
|-- ...
|-- 00001
|-- ... |-- mapping_file_training.json 

Then, you can use python utils/download_data.py to download all images.

Then, the file data structure should be like:

|-- humansd_data
|-- datasets
|-- Laion |-- Aesthetics_Human
|-- images
|-- 00000.parquet
|-- 00001.parquet
|-- ...
|-- 00000
|-- 000000000.jpg
|-- 000000001.jpg
|-- ...
|-- 00001
|-- ...
|-- pose
|-- 00000
|-- 000000000.npz
|-- 000000001.npz
|-- ...
|-- 00001
|-- ... |-- mapping_file_training.json 

If you download the LAION-Aesthetics in tar files, which is different from our data structure, we recommend you extract the tar file through code:

importtarfiletar_file="00000.tar"# 00000.tar - 00286.tarpresent_tar_path=f"xxxxxx/{tar_file}"save_dir="humansd_data/datasets/Laion/Aesthetics_Human/images"withtarfile.open(present_tar_path, "r") astar_file:
forpresent_fileintar_file.getmembers():
ifpresent_file.name.endswith(".jpg"):
print(f" image:- {present_file.name} -")
image_save_path=os.path.join(save_dir,tar_file.replace(".tar",""),present_file.name)
present_image_fp=TarIO.TarIO(present_tar_path, present_file.name)
present_image=Image.open(present_image_fp)
present_image_numpy=cv2.cvtColor(np.array(present_image),cv2.COLOR_RGB2BGR)
ifnotos.path.exists(os.path.dirname(image_save_path)):
os.makedirs(os.path.dirname(image_save_path))
cv2.imwrite(image_save_path,present_image_numpy)

Human-Art

You may download Human-Art dataset here.

The file data structure should be like:

|-- humansd_data
|-- datasets
|-- HumanArt |-- images
|-- 2D_virtual_human
|-- cartoon
|-- 000000000007.jpg
|-- 000000000019.jpg
|-- ...
|-- digital_art
|-- ...
|-- 3D_virtual_human
|-- real_human
|-- pose
|-- 2D_virtual_human
|-- cartoon
|-- 000000000007.npz
|-- 000000000019.npz
|-- ...
|-- digital_art
|-- ...
|-- 3D_virtual_human
|-- real_human
|-- mapping_file_training.json |-- mapping_file_validation.json 

Training

Note that the datasets and checkpoints should be downloaded and prepared before training.

Run the commands below to start training:

python main.py --base configs/humansd/humansd-finetune.yaml -t --gpus 0,1 --name finetune_humansd

If you want to finetune without heat-map-guided diffusion loss for ablation, you can run the following commands:

python main.py --base configs/humansd/humansd-finetune-originalloss.yaml -t --gpus 0,1 --name finetune_humansd_original_loss

Quantitative Results

Metrics can be calculated through:

python scripts/pose2img_metrics.py --outdir outputs/metrics --config utils/metrics/metrics.yaml --ckpt path_to_ckpt

Qualitative Results

  • (a) a generation by the pre-trained text-guided stable diffusion (SD)
  • (b) pose skeleton images as the condition to ControlNet, T2I-Adapter and our proposed HumanSD
  • (c) a generation by ControlNet
  • (d) a generation by T2I-Adapter
  • (e) a generation by HumanSD (ours).

ControlNet, T2I-Adapter, and HumanSD receive both text and pose conditions.

Natural Scene

Sketch Scene

Shadow Play Scene

Children Drawing Scene

Oil Painting Scene

Watercolor Scene

Digital Art Scene

Relief Scene

Sculpture Scene

Cite Us

@article{ju2023humansd,
title={Human{SD}: A Native Skeleton-Guided Diffusion Model for Human Image Generation},
author={Ju, Xuan and Zeng, Ailing and Zhao, Chenchen and Wang, Jianan and Zhang, Lei and Xu, Qiang},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
year={2023}
}
@inproceedings{ju2023human,
title={Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial Scenes},
author={Ju, Xuan and Zeng, Ailing and Wang, Jianan and Xu, Qiang and Zhang, Lei},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year={2023},
}

Acknowledgement

About

[ICCV 2023] The official implementation of paper "HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation"

Topics

Resources

Stars

306 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

HumanSD


This repository contains the implementation of the ICCV2023 paper:

HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation[Project Page][Paper][Code][Video][Data]
Xuan Ju∗12, Ailing Zeng∗1, Chenchen Zhao∗2, Jianan Wang1, Lei Zhang1, Qiang Xu2
Equal contribution 1International Digital Economy Academy 2The Chinese University of Hong Kong

In this work, we propose a native skeleton-guided diffusion model for controllable HIG called HumanSD. Instead of performing image editing with dual-branch diffusion, we fine-tune the original SD model using a novel heatmap-guided denoising loss. This strategy effectively and efficiently strengthens the given skeleton condition during model training while mitigating the catastrophic forgetting effects. HumanSD is fine-tuned on the assembly of three large-scale human-centric datasets with text-imagepose information, two of which are established in this work.


  • (a) a generation by the pre-trained pose-less text-guided stable diffusion (SD)
  • (b) pose skeleton images as the condition to ControlNet and our proposed HumanSD
  • (c) a generation by ControlNet
  • (d) a generation by HumanSD (ours). ControlNet and HumanSD receive both text and pose conditions.

HumanSD shows its superiorities in terms of (I) challenging poses, (II) accurate painting styles, (III) pose control capability, (IV) multi-person scenarios, and (V) delicate details.

Table of Contents

TODO

News!! Our paper have been accepted by ICCV2023! Training code is released.

  • Release inference code and pretrained models
  • Release Gradio UI demo
  • Public training data (LAION-Human)
  • Release training code

Model Overview

Getting Started

Environment Requirement

HumanSD has been implemented and tested on Pytorch 1.12.1 with python 3.9.

Clone the repo:

git clone git@github.com:IDEA-Research/HumanSD.git

We recommend you first install pytorch following official instructions. For example:

# conda
conda install pytorch==1.12.1 torchvision==0.13.1 torchaudio==0.12.1 cudatoolkit=11.3 -c pytorch

Then, you can install required packages thourgh:

pip install -r requirements.txt

You also need to install MMPose following here. Noted that you only need to install MMPose as a python package. PS: Because of the update of MMPose, we recommend you to install 0.29.0 version of MMPose.

Model and Checkpoints

Download necessary checkpoints of HumanSD, which can be found here. The data structure should be like:

|-- humansd_data
|-- checkpoints
|-- higherhrnet_w48_humanart_512x512_udp.pth
|-- v2-1_512-ema-pruned.ckpt
|-- humansd-v1.ckpt

Noted that v2-1_512-ema-pruned.ckpt should be download from Stable Diffusion.

Quick Demo

You can run demo either through command line or gradio.

You can run demo through command line with:

python scripts/pose2img.py --prompt "oil painting of girls dancing on the stage" --pose_file assets/pose/demo.npz

You can also run demo compared with ControlNet and T2I-Adapter:

python scripts/pose2img.py --prompt "oil painting of girls dancing on the stage" --pose_file assets/pose/demo.npz --controlnet --t2i

You can run gradio demo through:

python scripts/gradio/pose2img.py

We have also provided the comparison of ControlNet and T2I-Adapter, you can run all these methods in one demo. But you need to download corresponding model and checkpoints following:

To compare ControlNet, and T2I-Adpater's results. (1) You need to initialize ControlNet and T2I-Adapter as submodule using
git submodule init
git submodule update

(2) Then download checkpoints from: a. T2I-Adapter b. ControlNet. And put them into humansd_data/checkpoints

Then, run:

python scripts/gradio/pose2img.py --controlnet --t2i

Noted that you may have to modify some code in T2I-Adapter due to the path conflict.

e.g., use

from comparison_models.T2IAdapter.ldm.models.diffusion.ddim import DDIMSampler

instead of

from T2IAdapter.ldm.models.diffusion.ddim import DDIMSampler

Dataset

You may refer to the code here for loading the data.

Laion-Human

You may apply for access of Laion-Human here. Noted that we have provide the pose annotations, images' .parquet file and mapping file, please download the images according to .parquet. The key in .parquet is the corresponding image index. For example, image with key=338717 in 00033.parquet is corresponding to images/00000/000338717.jpg.

After downloading the images and pose, you need to extract zip files and make it looks like:

|-- humansd_data
|-- datasets
|-- Laion |-- Aesthetics_Human
|-- images
|-- 00000.parquet
|-- 00001.parquet
|-- ...
|-- pose
|-- 00000
|-- 000000000.npz
|-- 000000001.npz
|-- ...
|-- 00001
|-- ... |-- mapping_file_training.json 

Then, you can use python utils/download_data.py to download all images.

Then, the file data structure should be like:

|-- humansd_data
|-- datasets
|-- Laion |-- Aesthetics_Human
|-- images
|-- 00000.parquet
|-- 00001.parquet
|-- ...
|-- 00000
|-- 000000000.jpg
|-- 000000001.jpg
|-- ...
|-- 00001
|-- ...
|-- pose
|-- 00000
|-- 000000000.npz
|-- 000000001.npz
|-- ...
|-- 00001
|-- ... |-- mapping_file_training.json 

If you download the LAION-Aesthetics in tar files, which is different from our data structure, we recommend you extract the tar file through code:

importtarfiletar_file="00000.tar"# 00000.tar - 00286.tarpresent_tar_path=f"xxxxxx/{tar_file}"save_dir="humansd_data/datasets/Laion/Aesthetics_Human/images"withtarfile.open(present_tar_path, "r") astar_file:
forpresent_fileintar_file.getmembers():
ifpresent_file.name.endswith(".jpg"):
print(f" image:- {present_file.name} -")
image_save_path=os.path.join(save_dir,tar_file.replace(".tar",""),present_file.name)
present_image_fp=TarIO.TarIO(present_tar_path, present_file.name)
present_image=Image.open(present_image_fp)
present_image_numpy=cv2.cvtColor(np.array(present_image),cv2.COLOR_RGB2BGR)
ifnotos.path.exists(os.path.dirname(image_save_path)):
os.makedirs(os.path.dirname(image_save_path))
cv2.imwrite(image_save_path,present_image_numpy)

Human-Art

You may download Human-Art dataset here.

The file data structure should be like:

|-- humansd_data
|-- datasets
|-- HumanArt |-- images
|-- 2D_virtual_human
|-- cartoon
|-- 000000000007.jpg
|-- 000000000019.jpg
|-- ...
|-- digital_art
|-- ...
|-- 3D_virtual_human
|-- real_human
|-- pose
|-- 2D_virtual_human
|-- cartoon
|-- 000000000007.npz
|-- 000000000019.npz
|-- ...
|-- digital_art
|-- ...
|-- 3D_virtual_human
|-- real_human
|-- mapping_file_training.json |-- mapping_file_validation.json 

Training

Note that the datasets and checkpoints should be downloaded and prepared before training.

Run the commands below to start training:

python main.py --base configs/humansd/humansd-finetune.yaml -t --gpus 0,1 --name finetune_humansd

If you want to finetune without heat-map-guided diffusion loss for ablation, you can run the following commands:

python main.py --base configs/humansd/humansd-finetune-originalloss.yaml -t --gpus 0,1 --name finetune_humansd_original_loss

Quantitative Results

Metrics can be calculated through:

python scripts/pose2img_metrics.py --outdir outputs/metrics --config utils/metrics/metrics.yaml --ckpt path_to_ckpt

Qualitative Results

  • (a) a generation by the pre-trained text-guided stable diffusion (SD)
  • (b) pose skeleton images as the condition to ControlNet, T2I-Adapter and our proposed HumanSD
  • (c) a generation by ControlNet
  • (d) a generation by T2I-Adapter
  • (e) a generation by HumanSD (ours).

ControlNet, T2I-Adapter, and HumanSD receive both text and pose conditions.

Natural Scene

Sketch Scene

Shadow Play Scene

Children Drawing Scene

Oil Painting Scene

Watercolor Scene

Digital Art Scene

Relief Scene

Sculpture Scene

Cite Us

@article{ju2023humansd,
title={Human{SD}: A Native Skeleton-Guided Diffusion Model for Human Image Generation},
author={Ju, Xuan and Zeng, Ailing and Zhao, Chenchen and Wang, Jianan and Zhang, Lei and Xu, Qiang},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
year={2023}
}
@inproceedings{ju2023human,
title={Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial Scenes},
author={Ju, Xuan and Zeng, Ailing and Wang, Jianan and Xu, Qiang and Zhang, Lei},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year={2023},
}

Acknowledgement

About

[ICCV 2023] The official implementation of paper "HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation"

Topics

Resources

Stars

306 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

HumanSD


This repository contains the implementation of the ICCV2023 paper:

HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation[Project Page][Paper][Code][Video][Data]
Xuan Ju∗12, Ailing Zeng∗1, Chenchen Zhao∗2, Jianan Wang1, Lei Zhang1, Qiang Xu2
Equal contribution 1International Digital Economy Academy 2The Chinese University of Hong Kong

In this work, we propose a native skeleton-guided diffusion model for controllable HIG called HumanSD. Instead of performing image editing with dual-branch diffusion, we fine-tune the original SD model using a novel heatmap-guided denoising loss. This strategy effectively and efficiently strengthens the given skeleton condition during model training while mitigating the catastrophic forgetting effects. HumanSD is fine-tuned on the assembly of three large-scale human-centric datasets with text-imagepose information, two of which are established in this work.


  • (a) a generation by the pre-trained pose-less text-guided stable diffusion (SD)
  • (b) pose skeleton images as the condition to ControlNet and our proposed HumanSD
  • (c) a generation by ControlNet
  • (d) a generation by HumanSD (ours). ControlNet and HumanSD receive both text and pose conditions.

HumanSD shows its superiorities in terms of (I) challenging poses, (II) accurate painting styles, (III) pose control capability, (IV) multi-person scenarios, and (V) delicate details.

Table of Contents

TODO

News!! Our paper have been accepted by ICCV2023! Training code is released.

  • Release inference code and pretrained models
  • Release Gradio UI demo
  • Public training data (LAION-Human)
  • Release training code

Model Overview

Getting Started

Environment Requirement

HumanSD has been implemented and tested on Pytorch 1.12.1 with python 3.9.

Clone the repo:

git clone git@github.com:IDEA-Research/HumanSD.git

We recommend you first install pytorch following official instructions. For example:

# conda
conda install pytorch==1.12.1 torchvision==0.13.1 torchaudio==0.12.1 cudatoolkit=11.3 -c pytorch

Then, you can install required packages thourgh:

pip install -r requirements.txt

You also need to install MMPose following here. Noted that you only need to install MMPose as a python package. PS: Because of the update of MMPose, we recommend you to install 0.29.0 version of MMPose.

Model and Checkpoints

Download necessary checkpoints of HumanSD, which can be found here. The data structure should be like:

|-- humansd_data
|-- checkpoints
|-- higherhrnet_w48_humanart_512x512_udp.pth
|-- v2-1_512-ema-pruned.ckpt
|-- humansd-v1.ckpt

Noted that v2-1_512-ema-pruned.ckpt should be download from Stable Diffusion.

Quick Demo

You can run demo either through command line or gradio.

You can run demo through command line with:

python scripts/pose2img.py --prompt "oil painting of girls dancing on the stage" --pose_file assets/pose/demo.npz

You can also run demo compared with ControlNet and T2I-Adapter:

python scripts/pose2img.py --prompt "oil painting of girls dancing on the stage" --pose_file assets/pose/demo.npz --controlnet --t2i

You can run gradio demo through:

python scripts/gradio/pose2img.py

We have also provided the comparison of ControlNet and T2I-Adapter, you can run all these methods in one demo. But you need to download corresponding model and checkpoints following:

To compare ControlNet, and T2I-Adpater's results. (1) You need to initialize ControlNet and T2I-Adapter as submodule using
git submodule init
git submodule update

(2) Then download checkpoints from: a. T2I-Adapter b. ControlNet. And put them into humansd_data/checkpoints

Then, run:

python scripts/gradio/pose2img.py --controlnet --t2i

Noted that you may have to modify some code in T2I-Adapter due to the path conflict.

e.g., use

from comparison_models.T2IAdapter.ldm.models.diffusion.ddim import DDIMSampler

instead of

from T2IAdapter.ldm.models.diffusion.ddim import DDIMSampler

Dataset

You may refer to the code here for loading the data.

Laion-Human

You may apply for access of Laion-Human here. Noted that we have provide the pose annotations, images' .parquet file and mapping file, please download the images according to .parquet. The key in .parquet is the corresponding image index. For example, image with key=338717 in 00033.parquet is corresponding to images/00000/000338717.jpg.

After downloading the images and pose, you need to extract zip files and make it looks like:

|-- humansd_data
|-- datasets
|-- Laion |-- Aesthetics_Human
|-- images
|-- 00000.parquet
|-- 00001.parquet
|-- ...
|-- pose
|-- 00000
|-- 000000000.npz
|-- 000000001.npz
|-- ...
|-- 00001
|-- ... |-- mapping_file_training.json 

Then, you can use python utils/download_data.py to download all images.

Then, the file data structure should be like:

|-- humansd_data
|-- datasets
|-- Laion |-- Aesthetics_Human
|-- images
|-- 00000.parquet
|-- 00001.parquet
|-- ...
|-- 00000
|-- 000000000.jpg
|-- 000000001.jpg
|-- ...
|-- 00001
|-- ...
|-- pose
|-- 00000
|-- 000000000.npz
|-- 000000001.npz
|-- ...
|-- 00001
|-- ... |-- mapping_file_training.json 

If you download the LAION-Aesthetics in tar files, which is different from our data structure, we recommend you extract the tar file through code:

importtarfiletar_file="00000.tar"# 00000.tar - 00286.tarpresent_tar_path=f"xxxxxx/{tar_file}"save_dir="humansd_data/datasets/Laion/Aesthetics_Human/images"withtarfile.open(present_tar_path, "r") astar_file:
forpresent_fileintar_file.getmembers():
ifpresent_file.name.endswith(".jpg"):
print(f" image:- {present_file.name} -")
image_save_path=os.path.join(save_dir,tar_file.replace(".tar",""),present_file.name)
present_image_fp=TarIO.TarIO(present_tar_path, present_file.name)
present_image=Image.open(present_image_fp)
present_image_numpy=cv2.cvtColor(np.array(present_image),cv2.COLOR_RGB2BGR)
ifnotos.path.exists(os.path.dirname(image_save_path)):
os.makedirs(os.path.dirname(image_save_path))
cv2.imwrite(image_save_path,present_image_numpy)

Human-Art

You may download Human-Art dataset here.

The file data structure should be like:

|-- humansd_data
|-- datasets
|-- HumanArt |-- images
|-- 2D_virtual_human
|-- cartoon
|-- 000000000007.jpg
|-- 000000000019.jpg
|-- ...
|-- digital_art
|-- ...
|-- 3D_virtual_human
|-- real_human
|-- pose
|-- 2D_virtual_human
|-- cartoon
|-- 000000000007.npz
|-- 000000000019.npz
|-- ...
|-- digital_art
|-- ...
|-- 3D_virtual_human
|-- real_human
|-- mapping_file_training.json |-- mapping_file_validation.json 

Training

Note that the datasets and checkpoints should be downloaded and prepared before training.

Run the commands below to start training:

python main.py --base configs/humansd/humansd-finetune.yaml -t --gpus 0,1 --name finetune_humansd

If you want to finetune without heat-map-guided diffusion loss for ablation, you can run the following commands:

python main.py --base configs/humansd/humansd-finetune-originalloss.yaml -t --gpus 0,1 --name finetune_humansd_original_loss

Quantitative Results

Metrics can be calculated through:

python scripts/pose2img_metrics.py --outdir outputs/metrics --config utils/metrics/metrics.yaml --ckpt path_to_ckpt

Qualitative Results

  • (a) a generation by the pre-trained text-guided stable diffusion (SD)
  • (b) pose skeleton images as the condition to ControlNet, T2I-Adapter and our proposed HumanSD
  • (c) a generation by ControlNet
  • (d) a generation by T2I-Adapter
  • (e) a generation by HumanSD (ours).

ControlNet, T2I-Adapter, and HumanSD receive both text and pose conditions.

Natural Scene

Sketch Scene

Shadow Play Scene

Children Drawing Scene

Oil Painting Scene

Watercolor Scene

Digital Art Scene

Relief Scene

Sculpture Scene

Cite Us

@article{ju2023humansd,
title={Human{SD}: A Native Skeleton-Guided Diffusion Model for Human Image Generation},
author={Ju, Xuan and Zeng, Ailing and Zhao, Chenchen and Wang, Jianan and Zhang, Lei and Xu, Qiang},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
year={2023}
}
@inproceedings{ju2023human,
title={Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial Scenes},
author={Ju, Xuan and Zeng, Ailing and Wang, Jianan and Xu, Qiang and Zhang, Lei},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year={2023},
}

Acknowledgement

About

[ICCV 2023] The official implementation of paper "HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation"

Topics

Resources

Stars

306 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

HumanSD


This repository contains the implementation of the ICCV2023 paper:

HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation[Project Page][Paper][Code][Video][Data]
Xuan Ju∗12, Ailing Zeng∗1, Chenchen Zhao∗2, Jianan Wang1, Lei Zhang1, Qiang Xu2
Equal contribution 1International Digital Economy Academy 2The Chinese University of Hong Kong

In this work, we propose a native skeleton-guided diffusion model for controllable HIG called HumanSD. Instead of performing image editing with dual-branch diffusion, we fine-tune the original SD model using a novel heatmap-guided denoising loss. This strategy effectively and efficiently strengthens the given skeleton condition during model training while mitigating the catastrophic forgetting effects. HumanSD is fine-tuned on the assembly of three large-scale human-centric datasets with text-imagepose information, two of which are established in this work.


  • (a) a generation by the pre-trained pose-less text-guided stable diffusion (SD)
  • (b) pose skeleton images as the condition to ControlNet and our proposed HumanSD
  • (c) a generation by ControlNet
  • (d) a generation by HumanSD (ours). ControlNet and HumanSD receive both text and pose conditions.

HumanSD shows its superiorities in terms of (I) challenging poses, (II) accurate painting styles, (III) pose control capability, (IV) multi-person scenarios, and (V) delicate details.

Table of Contents

TODO

News!! Our paper have been accepted by ICCV2023! Training code is released.

  • Release inference code and pretrained models
  • Release Gradio UI demo
  • Public training data (LAION-Human)
  • Release training code

Model Overview

Getting Started

Environment Requirement

HumanSD has been implemented and tested on Pytorch 1.12.1 with python 3.9.

Clone the repo:

git clone git@github.com:IDEA-Research/HumanSD.git

We recommend you first install pytorch following official instructions. For example:

# conda
conda install pytorch==1.12.1 torchvision==0.13.1 torchaudio==0.12.1 cudatoolkit=11.3 -c pytorch

Then, you can install required packages thourgh:

pip install -r requirements.txt

You also need to install MMPose following here. Noted that you only need to install MMPose as a python package. PS: Because of the update of MMPose, we recommend you to install 0.29.0 version of MMPose.

Model and Checkpoints

Download necessary checkpoints of HumanSD, which can be found here. The data structure should be like:

|-- humansd_data
|-- checkpoints
|-- higherhrnet_w48_humanart_512x512_udp.pth
|-- v2-1_512-ema-pruned.ckpt
|-- humansd-v1.ckpt

Noted that v2-1_512-ema-pruned.ckpt should be download from Stable Diffusion.

Quick Demo

You can run demo either through command line or gradio.

You can run demo through command line with:

python scripts/pose2img.py --prompt "oil painting of girls dancing on the stage" --pose_file assets/pose/demo.npz

You can also run demo compared with ControlNet and T2I-Adapter:

python scripts/pose2img.py --prompt "oil painting of girls dancing on the stage" --pose_file assets/pose/demo.npz --controlnet --t2i

You can run gradio demo through:

python scripts/gradio/pose2img.py

We have also provided the comparison of ControlNet and T2I-Adapter, you can run all these methods in one demo. But you need to download corresponding model and checkpoints following:

To compare ControlNet, and T2I-Adpater's results. (1) You need to initialize ControlNet and T2I-Adapter as submodule using
git submodule init
git submodule update

(2) Then download checkpoints from: a. T2I-Adapter b. ControlNet. And put them into humansd_data/checkpoints

Then, run:

python scripts/gradio/pose2img.py --controlnet --t2i

Noted that you may have to modify some code in T2I-Adapter due to the path conflict.

e.g., use

from comparison_models.T2IAdapter.ldm.models.diffusion.ddim import DDIMSampler

instead of

from T2IAdapter.ldm.models.diffusion.ddim import DDIMSampler

Dataset

You may refer to the code here for loading the data.

Laion-Human

You may apply for access of Laion-Human here. Noted that we have provide the pose annotations, images' .parquet file and mapping file, please download the images according to .parquet. The key in .parquet is the corresponding image index. For example, image with key=338717 in 00033.parquet is corresponding to images/00000/000338717.jpg.

After downloading the images and pose, you need to extract zip files and make it looks like:

|-- humansd_data
|-- datasets
|-- Laion |-- Aesthetics_Human
|-- images
|-- 00000.parquet
|-- 00001.parquet
|-- ...
|-- pose
|-- 00000
|-- 000000000.npz
|-- 000000001.npz
|-- ...
|-- 00001
|-- ... |-- mapping_file_training.json 

Then, you can use python utils/download_data.py to download all images.

Then, the file data structure should be like:

|-- humansd_data
|-- datasets
|-- Laion |-- Aesthetics_Human
|-- images
|-- 00000.parquet
|-- 00001.parquet
|-- ...
|-- 00000
|-- 000000000.jpg
|-- 000000001.jpg
|-- ...
|-- 00001
|-- ...
|-- pose
|-- 00000
|-- 000000000.npz
|-- 000000001.npz
|-- ...
|-- 00001
|-- ... |-- mapping_file_training.json 

If you download the LAION-Aesthetics in tar files, which is different from our data structure, we recommend you extract the tar file through code:

importtarfiletar_file="00000.tar"# 00000.tar - 00286.tarpresent_tar_path=f"xxxxxx/{tar_file}"save_dir="humansd_data/datasets/Laion/Aesthetics_Human/images"withtarfile.open(present_tar_path, "r") astar_file:
forpresent_fileintar_file.getmembers():
ifpresent_file.name.endswith(".jpg"):
print(f" image:- {present_file.name} -")
image_save_path=os.path.join(save_dir,tar_file.replace(".tar",""),present_file.name)
present_image_fp=TarIO.TarIO(present_tar_path, present_file.name)
present_image=Image.open(present_image_fp)
present_image_numpy=cv2.cvtColor(np.array(present_image),cv2.COLOR_RGB2BGR)
ifnotos.path.exists(os.path.dirname(image_save_path)):
os.makedirs(os.path.dirname(image_save_path))
cv2.imwrite(image_save_path,present_image_numpy)

Human-Art

You may download Human-Art dataset here.

The file data structure should be like:

|-- humansd_data
|-- datasets
|-- HumanArt |-- images
|-- 2D_virtual_human
|-- cartoon
|-- 000000000007.jpg
|-- 000000000019.jpg
|-- ...
|-- digital_art
|-- ...
|-- 3D_virtual_human
|-- real_human
|-- pose
|-- 2D_virtual_human
|-- cartoon
|-- 000000000007.npz
|-- 000000000019.npz
|-- ...
|-- digital_art
|-- ...
|-- 3D_virtual_human
|-- real_human
|-- mapping_file_training.json |-- mapping_file_validation.json 

Training

Note that the datasets and checkpoints should be downloaded and prepared before training.

Run the commands below to start training:

python main.py --base configs/humansd/humansd-finetune.yaml -t --gpus 0,1 --name finetune_humansd

If you want to finetune without heat-map-guided diffusion loss for ablation, you can run the following commands:

python main.py --base configs/humansd/humansd-finetune-originalloss.yaml -t --gpus 0,1 --name finetune_humansd_original_loss

Quantitative Results

Metrics can be calculated through:

python scripts/pose2img_metrics.py --outdir outputs/metrics --config utils/metrics/metrics.yaml --ckpt path_to_ckpt

Qualitative Results

  • (a) a generation by the pre-trained text-guided stable diffusion (SD)
  • (b) pose skeleton images as the condition to ControlNet, T2I-Adapter and our proposed HumanSD
  • (c) a generation by ControlNet
  • (d) a generation by T2I-Adapter
  • (e) a generation by HumanSD (ours).

ControlNet, T2I-Adapter, and HumanSD receive both text and pose conditions.

Natural Scene

Sketch Scene

Shadow Play Scene

Children Drawing Scene

Oil Painting Scene

Watercolor Scene

Digital Art Scene

Relief Scene

Sculpture Scene

Cite Us

@article{ju2023humansd,
title={Human{SD}: A Native Skeleton-Guided Diffusion Model for Human Image Generation},
author={Ju, Xuan and Zeng, Ailing and Zhao, Chenchen and Wang, Jianan and Zhang, Lei and Xu, Qiang},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
year={2023}
}
@inproceedings{ju2023human,
title={Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial Scenes},
author={Ju, Xuan and Zeng, Ailing and Wang, Jianan and Xu, Qiang and Zhang, Lei},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year={2023},
}

Acknowledgement

About

[ICCV 2023] The official implementation of paper "HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation"

Topics

Resources

Stars

306 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

HumanSD


This repository contains the implementation of the ICCV2023 paper:

HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation[Project Page][Paper][Code][Video][Data]
Xuan Ju∗12, Ailing Zeng∗1, Chenchen Zhao∗2, Jianan Wang1, Lei Zhang1, Qiang Xu2
Equal contribution 1International Digital Economy Academy 2The Chinese University of Hong Kong

In this work, we propose a native skeleton-guided diffusion model for controllable HIG called HumanSD. Instead of performing image editing with dual-branch diffusion, we fine-tune the original SD model using a novel heatmap-guided denoising loss. This strategy effectively and efficiently strengthens the given skeleton condition during model training while mitigating the catastrophic forgetting effects. HumanSD is fine-tuned on the assembly of three large-scale human-centric datasets with text-imagepose information, two of which are established in this work.


  • (a) a generation by the pre-trained pose-less text-guided stable diffusion (SD)
  • (b) pose skeleton images as the condition to ControlNet and our proposed HumanSD
  • (c) a generation by ControlNet
  • (d) a generation by HumanSD (ours). ControlNet and HumanSD receive both text and pose conditions.

HumanSD shows its superiorities in terms of (I) challenging poses, (II) accurate painting styles, (III) pose control capability, (IV) multi-person scenarios, and (V) delicate details.

Table of Contents

TODO

News!! Our paper have been accepted by ICCV2023! Training code is released.

  • Release inference code and pretrained models
  • Release Gradio UI demo
  • Public training data (LAION-Human)
  • Release training code

Model Overview

Getting Started

Environment Requirement

HumanSD has been implemented and tested on Pytorch 1.12.1 with python 3.9.

Clone the repo:

git clone git@github.com:IDEA-Research/HumanSD.git

We recommend you first install pytorch following official instructions. For example:

# conda
conda install pytorch==1.12.1 torchvision==0.13.1 torchaudio==0.12.1 cudatoolkit=11.3 -c pytorch

Then, you can install required packages thourgh:

pip install -r requirements.txt

You also need to install MMPose following here. Noted that you only need to install MMPose as a python package. PS: Because of the update of MMPose, we recommend you to install 0.29.0 version of MMPose.

Model and Checkpoints

Download necessary checkpoints of HumanSD, which can be found here. The data structure should be like:

|-- humansd_data
|-- checkpoints
|-- higherhrnet_w48_humanart_512x512_udp.pth
|-- v2-1_512-ema-pruned.ckpt
|-- humansd-v1.ckpt

Noted that v2-1_512-ema-pruned.ckpt should be download from Stable Diffusion.

Quick Demo

You can run demo either through command line or gradio.

You can run demo through command line with:

python scripts/pose2img.py --prompt "oil painting of girls dancing on the stage" --pose_file assets/pose/demo.npz

You can also run demo compared with ControlNet and T2I-Adapter:

python scripts/pose2img.py --prompt "oil painting of girls dancing on the stage" --pose_file assets/pose/demo.npz --controlnet --t2i

You can run gradio demo through:

python scripts/gradio/pose2img.py

We have also provided the comparison of ControlNet and T2I-Adapter, you can run all these methods in one demo. But you need to download corresponding model and checkpoints following:

To compare ControlNet, and T2I-Adpater's results. (1) You need to initialize ControlNet and T2I-Adapter as submodule using
git submodule init
git submodule update

(2) Then download checkpoints from: a. T2I-Adapter b. ControlNet. And put them into humansd_data/checkpoints

Then, run:

python scripts/gradio/pose2img.py --controlnet --t2i

Noted that you may have to modify some code in T2I-Adapter due to the path conflict.

e.g., use

from comparison_models.T2IAdapter.ldm.models.diffusion.ddim import DDIMSampler

instead of

from T2IAdapter.ldm.models.diffusion.ddim import DDIMSampler

Dataset

You may refer to the code here for loading the data.

Laion-Human

You may apply for access of Laion-Human here. Noted that we have provide the pose annotations, images' .parquet file and mapping file, please download the images according to .parquet. The key in .parquet is the corresponding image index. For example, image with key=338717 in 00033.parquet is corresponding to images/00000/000338717.jpg.

After downloading the images and pose, you need to extract zip files and make it looks like:

|-- humansd_data
|-- datasets
|-- Laion |-- Aesthetics_Human
|-- images
|-- 00000.parquet
|-- 00001.parquet
|-- ...
|-- pose
|-- 00000
|-- 000000000.npz
|-- 000000001.npz
|-- ...
|-- 00001
|-- ... |-- mapping_file_training.json 

Then, you can use python utils/download_data.py to download all images.

Then, the file data structure should be like:

|-- humansd_data
|-- datasets
|-- Laion |-- Aesthetics_Human
|-- images
|-- 00000.parquet
|-- 00001.parquet
|-- ...
|-- 00000
|-- 000000000.jpg
|-- 000000001.jpg
|-- ...
|-- 00001
|-- ...
|-- pose
|-- 00000
|-- 000000000.npz
|-- 000000001.npz
|-- ...
|-- 00001
|-- ... |-- mapping_file_training.json 

If you download the LAION-Aesthetics in tar files, which is different from our data structure, we recommend you extract the tar file through code:

importtarfiletar_file="00000.tar"# 00000.tar - 00286.tarpresent_tar_path=f"xxxxxx/{tar_file}"save_dir="humansd_data/datasets/Laion/Aesthetics_Human/images"withtarfile.open(present_tar_path, "r") astar_file:
forpresent_fileintar_file.getmembers():
ifpresent_file.name.endswith(".jpg"):
print(f" image:- {present_file.name} -")
image_save_path=os.path.join(save_dir,tar_file.replace(".tar",""),present_file.name)
present_image_fp=TarIO.TarIO(present_tar_path, present_file.name)
present_image=Image.open(present_image_fp)
present_image_numpy=cv2.cvtColor(np.array(present_image),cv2.COLOR_RGB2BGR)
ifnotos.path.exists(os.path.dirname(image_save_path)):
os.makedirs(os.path.dirname(image_save_path))
cv2.imwrite(image_save_path,present_image_numpy)

Human-Art

You may download Human-Art dataset here.

The file data structure should be like:

|-- humansd_data
|-- datasets
|-- HumanArt |-- images
|-- 2D_virtual_human
|-- cartoon
|-- 000000000007.jpg
|-- 000000000019.jpg
|-- ...
|-- digital_art
|-- ...
|-- 3D_virtual_human
|-- real_human
|-- pose
|-- 2D_virtual_human
|-- cartoon
|-- 000000000007.npz
|-- 000000000019.npz
|-- ...
|-- digital_art
|-- ...
|-- 3D_virtual_human
|-- real_human
|-- mapping_file_training.json |-- mapping_file_validation.json 

Training

Note that the datasets and checkpoints should be downloaded and prepared before training.

Run the commands below to start training:

python main.py --base configs/humansd/humansd-finetune.yaml -t --gpus 0,1 --name finetune_humansd

If you want to finetune without heat-map-guided diffusion loss for ablation, you can run the following commands:

python main.py --base configs/humansd/humansd-finetune-originalloss.yaml -t --gpus 0,1 --name finetune_humansd_original_loss

Quantitative Results

Metrics can be calculated through:

python scripts/pose2img_metrics.py --outdir outputs/metrics --config utils/metrics/metrics.yaml --ckpt path_to_ckpt

Qualitative Results

  • (a) a generation by the pre-trained text-guided stable diffusion (SD)
  • (b) pose skeleton images as the condition to ControlNet, T2I-Adapter and our proposed HumanSD
  • (c) a generation by ControlNet
  • (d) a generation by T2I-Adapter
  • (e) a generation by HumanSD (ours).

ControlNet, T2I-Adapter, and HumanSD receive both text and pose conditions.

Natural Scene

Sketch Scene

Shadow Play Scene

Children Drawing Scene

Oil Painting Scene

Watercolor Scene

Digital Art Scene

Relief Scene

Sculpture Scene

Cite Us

@article{ju2023humansd,
title={Human{SD}: A Native Skeleton-Guided Diffusion Model for Human Image Generation},
author={Ju, Xuan and Zeng, Ailing and Zhao, Chenchen and Wang, Jianan and Zhang, Lei and Xu, Qiang},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
year={2023}
}
@inproceedings{ju2023human,
title={Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial Scenes},
author={Ju, Xuan and Zeng, Ailing and Wang, Jianan and Xu, Qiang and Zhang, Lei},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year={2023},
}

Acknowledgement

About

[ICCV 2023] The official implementation of paper "HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation"

Topics

Resources

Stars

306 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

HumanSD


This repository contains the implementation of the ICCV2023 paper:

HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation[Project Page][Paper][Code][Video][Data]
Xuan Ju∗12, Ailing Zeng∗1, Chenchen Zhao∗2, Jianan Wang1, Lei Zhang1, Qiang Xu2
Equal contribution 1International Digital Economy Academy 2The Chinese University of Hong Kong

In this work, we propose a native skeleton-guided diffusion model for controllable HIG called HumanSD. Instead of performing image editing with dual-branch diffusion, we fine-tune the original SD model using a novel heatmap-guided denoising loss. This strategy effectively and efficiently strengthens the given skeleton condition during model training while mitigating the catastrophic forgetting effects. HumanSD is fine-tuned on the assembly of three large-scale human-centric datasets with text-imagepose information, two of which are established in this work.


  • (a) a generation by the pre-trained pose-less text-guided stable diffusion (SD)
  • (b) pose skeleton images as the condition to ControlNet and our proposed HumanSD
  • (c) a generation by ControlNet
  • (d) a generation by HumanSD (ours). ControlNet and HumanSD receive both text and pose conditions.

HumanSD shows its superiorities in terms of (I) challenging poses, (II) accurate painting styles, (III) pose control capability, (IV) multi-person scenarios, and (V) delicate details.

Table of Contents

TODO

News!! Our paper have been accepted by ICCV2023! Training code is released.

  • Release inference code and pretrained models
  • Release Gradio UI demo
  • Public training data (LAION-Human)
  • Release training code

Model Overview

Getting Started

Environment Requirement

HumanSD has been implemented and tested on Pytorch 1.12.1 with python 3.9.

Clone the repo:

git clone git@github.com:IDEA-Research/HumanSD.git

We recommend you first install pytorch following official instructions. For example:

# conda
conda install pytorch==1.12.1 torchvision==0.13.1 torchaudio==0.12.1 cudatoolkit=11.3 -c pytorch

Then, you can install required packages thourgh:

pip install -r requirements.txt

You also need to install MMPose following here. Noted that you only need to install MMPose as a python package. PS: Because of the update of MMPose, we recommend you to install 0.29.0 version of MMPose.

Model and Checkpoints

Download necessary checkpoints of HumanSD, which can be found here. The data structure should be like:

|-- humansd_data
|-- checkpoints
|-- higherhrnet_w48_humanart_512x512_udp.pth
|-- v2-1_512-ema-pruned.ckpt
|-- humansd-v1.ckpt

Noted that v2-1_512-ema-pruned.ckpt should be download from Stable Diffusion.

Quick Demo

You can run demo either through command line or gradio.

You can run demo through command line with:

python scripts/pose2img.py --prompt "oil painting of girls dancing on the stage" --pose_file assets/pose/demo.npz

You can also run demo compared with ControlNet and T2I-Adapter:

python scripts/pose2img.py --prompt "oil painting of girls dancing on the stage" --pose_file assets/pose/demo.npz --controlnet --t2i

You can run gradio demo through:

python scripts/gradio/pose2img.py

We have also provided the comparison of ControlNet and T2I-Adapter, you can run all these methods in one demo. But you need to download corresponding model and checkpoints following:

To compare ControlNet, and T2I-Adpater's results. (1) You need to initialize ControlNet and T2I-Adapter as submodule using
git submodule init
git submodule update

(2) Then download checkpoints from: a. T2I-Adapter b. ControlNet. And put them into humansd_data/checkpoints

Then, run:

python scripts/gradio/pose2img.py --controlnet --t2i

Noted that you may have to modify some code in T2I-Adapter due to the path conflict.

e.g., use

from comparison_models.T2IAdapter.ldm.models.diffusion.ddim import DDIMSampler

instead of

from T2IAdapter.ldm.models.diffusion.ddim import DDIMSampler

Dataset

You may refer to the code here for loading the data.

Laion-Human

You may apply for access of Laion-Human here. Noted that we have provide the pose annotations, images' .parquet file and mapping file, please download the images according to .parquet. The key in .parquet is the corresponding image index. For example, image with key=338717 in 00033.parquet is corresponding to images/00000/000338717.jpg.

After downloading the images and pose, you need to extract zip files and make it looks like:

|-- humansd_data
|-- datasets
|-- Laion |-- Aesthetics_Human
|-- images
|-- 00000.parquet
|-- 00001.parquet
|-- ...
|-- pose
|-- 00000
|-- 000000000.npz
|-- 000000001.npz
|-- ...
|-- 00001
|-- ... |-- mapping_file_training.json 

Then, you can use python utils/download_data.py to download all images.

Then, the file data structure should be like:

|-- humansd_data
|-- datasets
|-- Laion |-- Aesthetics_Human
|-- images
|-- 00000.parquet
|-- 00001.parquet
|-- ...
|-- 00000
|-- 000000000.jpg
|-- 000000001.jpg
|-- ...
|-- 00001
|-- ...
|-- pose
|-- 00000
|-- 000000000.npz
|-- 000000001.npz
|-- ...
|-- 00001
|-- ... |-- mapping_file_training.json 

If you download the LAION-Aesthetics in tar files, which is different from our data structure, we recommend you extract the tar file through code:

importtarfiletar_file="00000.tar"# 00000.tar - 00286.tarpresent_tar_path=f"xxxxxx/{tar_file}"save_dir="humansd_data/datasets/Laion/Aesthetics_Human/images"withtarfile.open(present_tar_path, "r") astar_file:
forpresent_fileintar_file.getmembers():
ifpresent_file.name.endswith(".jpg"):
print(f" image:- {present_file.name} -")
image_save_path=os.path.join(save_dir,tar_file.replace(".tar",""),present_file.name)
present_image_fp=TarIO.TarIO(present_tar_path, present_file.name)
present_image=Image.open(present_image_fp)
present_image_numpy=cv2.cvtColor(np.array(present_image),cv2.COLOR_RGB2BGR)
ifnotos.path.exists(os.path.dirname(image_save_path)):
os.makedirs(os.path.dirname(image_save_path))
cv2.imwrite(image_save_path,present_image_numpy)

Human-Art

You may download Human-Art dataset here.

The file data structure should be like:

|-- humansd_data
|-- datasets
|-- HumanArt |-- images
|-- 2D_virtual_human
|-- cartoon
|-- 000000000007.jpg
|-- 000000000019.jpg
|-- ...
|-- digital_art
|-- ...
|-- 3D_virtual_human
|-- real_human
|-- pose
|-- 2D_virtual_human
|-- cartoon
|-- 000000000007.npz
|-- 000000000019.npz
|-- ...
|-- digital_art
|-- ...
|-- 3D_virtual_human
|-- real_human
|-- mapping_file_training.json |-- mapping_file_validation.json 

Training

Note that the datasets and checkpoints should be downloaded and prepared before training.

Run the commands below to start training:

python main.py --base configs/humansd/humansd-finetune.yaml -t --gpus 0,1 --name finetune_humansd

If you want to finetune without heat-map-guided diffusion loss for ablation, you can run the following commands:

python main.py --base configs/humansd/humansd-finetune-originalloss.yaml -t --gpus 0,1 --name finetune_humansd_original_loss

Quantitative Results

Metrics can be calculated through:

python scripts/pose2img_metrics.py --outdir outputs/metrics --config utils/metrics/metrics.yaml --ckpt path_to_ckpt

Qualitative Results

  • (a) a generation by the pre-trained text-guided stable diffusion (SD)
  • (b) pose skeleton images as the condition to ControlNet, T2I-Adapter and our proposed HumanSD
  • (c) a generation by ControlNet
  • (d) a generation by T2I-Adapter
  • (e) a generation by HumanSD (ours).

ControlNet, T2I-Adapter, and HumanSD receive both text and pose conditions.

Natural Scene

Sketch Scene

Shadow Play Scene

Children Drawing Scene

Oil Painting Scene

Watercolor Scene

Digital Art Scene

Relief Scene

Sculpture Scene

Cite Us

@article{ju2023humansd,
title={Human{SD}: A Native Skeleton-Guided Diffusion Model for Human Image Generation},
author={Ju, Xuan and Zeng, Ailing and Zhao, Chenchen and Wang, Jianan and Zhang, Lei and Xu, Qiang},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
year={2023}
}
@inproceedings{ju2023human,
title={Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial Scenes},
author={Ju, Xuan and Zeng, Ailing and Wang, Jianan and Xu, Qiang and Zhang, Lei},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year={2023},
}

Acknowledgement

About

[ICCV 2023] The official implementation of paper "HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation"

Topics

Resources

Stars

306 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

HumanSD


This repository contains the implementation of the ICCV2023 paper:

HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation[Project Page][Paper][Code][Video][Data]
Xuan Ju∗12, Ailing Zeng∗1, Chenchen Zhao∗2, Jianan Wang1, Lei Zhang1, Qiang Xu2
Equal contribution 1International Digital Economy Academy 2The Chinese University of Hong Kong

In this work, we propose a native skeleton-guided diffusion model for controllable HIG called HumanSD. Instead of performing image editing with dual-branch diffusion, we fine-tune the original SD model using a novel heatmap-guided denoising loss. This strategy effectively and efficiently strengthens the given skeleton condition during model training while mitigating the catastrophic forgetting effects. HumanSD is fine-tuned on the assembly of three large-scale human-centric datasets with text-imagepose information, two of which are established in this work.


  • (a) a generation by the pre-trained pose-less text-guided stable diffusion (SD)
  • (b) pose skeleton images as the condition to ControlNet and our proposed HumanSD
  • (c) a generation by ControlNet
  • (d) a generation by HumanSD (ours). ControlNet and HumanSD receive both text and pose conditions.

HumanSD shows its superiorities in terms of (I) challenging poses, (II) accurate painting styles, (III) pose control capability, (IV) multi-person scenarios, and (V) delicate details.

Table of Contents

TODO

News!! Our paper have been accepted by ICCV2023! Training code is released.

  • Release inference code and pretrained models
  • Release Gradio UI demo
  • Public training data (LAION-Human)
  • Release training code

Model Overview

Getting Started

Environment Requirement

HumanSD has been implemented and tested on Pytorch 1.12.1 with python 3.9.

Clone the repo:

git clone git@github.com:IDEA-Research/HumanSD.git

We recommend you first install pytorch following official instructions. For example:

# conda
conda install pytorch==1.12.1 torchvision==0.13.1 torchaudio==0.12.1 cudatoolkit=11.3 -c pytorch

Then, you can install required packages thourgh:

pip install -r requirements.txt

You also need to install MMPose following here. Noted that you only need to install MMPose as a python package. PS: Because of the update of MMPose, we recommend you to install 0.29.0 version of MMPose.

Model and Checkpoints

Download necessary checkpoints of HumanSD, which can be found here. The data structure should be like:

|-- humansd_data
|-- checkpoints
|-- higherhrnet_w48_humanart_512x512_udp.pth
|-- v2-1_512-ema-pruned.ckpt
|-- humansd-v1.ckpt

Noted that v2-1_512-ema-pruned.ckpt should be download from Stable Diffusion.

Quick Demo

You can run demo either through command line or gradio.

You can run demo through command line with:

python scripts/pose2img.py --prompt "oil painting of girls dancing on the stage" --pose_file assets/pose/demo.npz

You can also run demo compared with ControlNet and T2I-Adapter:

python scripts/pose2img.py --prompt "oil painting of girls dancing on the stage" --pose_file assets/pose/demo.npz --controlnet --t2i

You can run gradio demo through:

python scripts/gradio/pose2img.py

We have also provided the comparison of ControlNet and T2I-Adapter, you can run all these methods in one demo. But you need to download corresponding model and checkpoints following:

To compare ControlNet, and T2I-Adpater's results. (1) You need to initialize ControlNet and T2I-Adapter as submodule using
git submodule init
git submodule update

(2) Then download checkpoints from: a. T2I-Adapter b. ControlNet. And put them into humansd_data/checkpoints

Then, run:

python scripts/gradio/pose2img.py --controlnet --t2i

Noted that you may have to modify some code in T2I-Adapter due to the path conflict.

e.g., use

from comparison_models.T2IAdapter.ldm.models.diffusion.ddim import DDIMSampler

instead of

from T2IAdapter.ldm.models.diffusion.ddim import DDIMSampler

Dataset

You may refer to the code here for loading the data.

Laion-Human

You may apply for access of Laion-Human here. Noted that we have provide the pose annotations, images' .parquet file and mapping file, please download the images according to .parquet. The key in .parquet is the corresponding image index. For example, image with key=338717 in 00033.parquet is corresponding to images/00000/000338717.jpg.

After downloading the images and pose, you need to extract zip files and make it looks like:

|-- humansd_data
|-- datasets
|-- Laion |-- Aesthetics_Human
|-- images
|-- 00000.parquet
|-- 00001.parquet
|-- ...
|-- pose
|-- 00000
|-- 000000000.npz
|-- 000000001.npz
|-- ...
|-- 00001
|-- ... |-- mapping_file_training.json 

Then, you can use python utils/download_data.py to download all images.

Then, the file data structure should be like:

|-- humansd_data
|-- datasets
|-- Laion |-- Aesthetics_Human
|-- images
|-- 00000.parquet
|-- 00001.parquet
|-- ...
|-- 00000
|-- 000000000.jpg
|-- 000000001.jpg
|-- ...
|-- 00001
|-- ...
|-- pose
|-- 00000
|-- 000000000.npz
|-- 000000001.npz
|-- ...
|-- 00001
|-- ... |-- mapping_file_training.json 

If you download the LAION-Aesthetics in tar files, which is different from our data structure, we recommend you extract the tar file through code:

importtarfiletar_file="00000.tar"# 00000.tar - 00286.tarpresent_tar_path=f"xxxxxx/{tar_file}"save_dir="humansd_data/datasets/Laion/Aesthetics_Human/images"withtarfile.open(present_tar_path, "r") astar_file:
forpresent_fileintar_file.getmembers():
ifpresent_file.name.endswith(".jpg"):
print(f" image:- {present_file.name} -")
image_save_path=os.path.join(save_dir,tar_file.replace(".tar",""),present_file.name)
present_image_fp=TarIO.TarIO(present_tar_path, present_file.name)
present_image=Image.open(present_image_fp)
present_image_numpy=cv2.cvtColor(np.array(present_image),cv2.COLOR_RGB2BGR)
ifnotos.path.exists(os.path.dirname(image_save_path)):
os.makedirs(os.path.dirname(image_save_path))
cv2.imwrite(image_save_path,present_image_numpy)

Human-Art

You may download Human-Art dataset here.

The file data structure should be like:

|-- humansd_data
|-- datasets
|-- HumanArt |-- images
|-- 2D_virtual_human
|-- cartoon
|-- 000000000007.jpg
|-- 000000000019.jpg
|-- ...
|-- digital_art
|-- ...
|-- 3D_virtual_human
|-- real_human
|-- pose
|-- 2D_virtual_human
|-- cartoon
|-- 000000000007.npz
|-- 000000000019.npz
|-- ...
|-- digital_art
|-- ...
|-- 3D_virtual_human
|-- real_human
|-- mapping_file_training.json |-- mapping_file_validation.json 

Training

Note that the datasets and checkpoints should be downloaded and prepared before training.

Run the commands below to start training:

python main.py --base configs/humansd/humansd-finetune.yaml -t --gpus 0,1 --name finetune_humansd

If you want to finetune without heat-map-guided diffusion loss for ablation, you can run the following commands:

python main.py --base configs/humansd/humansd-finetune-originalloss.yaml -t --gpus 0,1 --name finetune_humansd_original_loss

Quantitative Results

Metrics can be calculated through:

python scripts/pose2img_metrics.py --outdir outputs/metrics --config utils/metrics/metrics.yaml --ckpt path_to_ckpt

Qualitative Results

  • (a) a generation by the pre-trained text-guided stable diffusion (SD)
  • (b) pose skeleton images as the condition to ControlNet, T2I-Adapter and our proposed HumanSD
  • (c) a generation by ControlNet
  • (d) a generation by T2I-Adapter
  • (e) a generation by HumanSD (ours).

ControlNet, T2I-Adapter, and HumanSD receive both text and pose conditions.

Natural Scene

Sketch Scene

Shadow Play Scene

Children Drawing Scene

Oil Painting Scene

Watercolor Scene

Digital Art Scene

Relief Scene

Sculpture Scene

Cite Us

@article{ju2023humansd,
title={Human{SD}: A Native Skeleton-Guided Diffusion Model for Human Image Generation},
author={Ju, Xuan and Zeng, Ailing and Zhao, Chenchen and Wang, Jianan and Zhang, Lei and Xu, Qiang},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
year={2023}
}
@inproceedings{ju2023human,
title={Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial Scenes},
author={Ju, Xuan and Zeng, Ailing and Wang, Jianan and Xu, Qiang and Zhang, Lei},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year={2023},
}

Acknowledgement

About

[ICCV 2023] The official implementation of paper "HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation"

Topics

Resources

Stars

306 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages