Repository files navigation

Collaborative Diffusion (CVPR 2023)

This repository contains the implementation of the following paper:

Collaborative Diffusion for Multi-Modal Face Generation and Editing
Ziqi Huang, Kelvin C.K. Chan, Yuming Jiang, Ziwei Liu
IEEE/CVF International Conference on Computer Vision (CVPR), 2023

From MMLab@NTU affiliated with S-Lab, Nanyang Technological University

[Paper] | [Project Page] | [Video]

Overview

We propose Collaborative Diffusion, where users can use multiple modalities to control face generation and editing. (a) Face Generation. Given multi-modal controls, our framework synthesizes high-quality images consistent with the input conditions. (b) Face Editing. Collaborative Diffusion also supports multi-modal editing of real images with promising identity preservation capability.


We use pre-trained uni-modal diffusion models to perform multi-modal guided face generation and editing. At each step of the reverse process (i.e., from timestep t to t − 1), the dynamic diffuser predicts the spatial-varying and temporal-varying influence function to selectively enhance or suppress the contributions of the given modality.

Updates

  • [06/2023] We provide the preprocessed multi-modal annotations here.
  • [05/2023] Training code for Collaborative Diffusion (512x512) released.
  • [04/2023] Project page and video available.
  • [04/2023] Arxiv paper available.
  • [04/2023] Checkpoints for multi-modal face generation (512x512) released.
  • [04/2023] Inference code for multi-modal face generation (512x512) released.

Installation

  1. Clone repo

    git clone https://github.com/ziqihuangg/Collaborative-Diffusion
    cd Collaborative-Diffusion
  2. Create conda environment.
    If you already have an ldm environment installed according to LDM, you do not need to go throught this step (i.e., step 2). You can simply conda activate ldm and jump to step 3.

     conda env create -f environment.yaml
    conda activate codiff
  3. Install dependencies

     pip install transformers==4.19.2 scann kornia==0.6.4 torchmetrics==0.6.0
    conda install -c anaconda git
    pip install git+https://github.com/arogozhnikov/einops.git

Download

Download Checkpoints

  1. Download the pre-trained models from here.

  2. Put the models under pretrained as follows:

    Collaborative-Diffusion
    └── pretrained
    ├── 512_codiff_mask_text.ckpt
    ├── 512_mask.ckpt
    ├── 512_text.ckpt
    └── 512_vae.ckpt
    

Download Datasets

We provide preprocessed data used in this project (see Acknowledgement for data source). You need download them only if you want to reproduce the training of Collaborative Diffusion. You can skip this step if you simply want to use our pre-trained models for inference.

  1. Download the preprocessed training data from here.

  2. Put the datasets under dataset as follows:

    Collaborative-Diffusion
    └── dataset
    ├── image
    | └──image_512_downsampled_from_hq_1024
    ├── text
    | └──captions_hq_beard_and_age_2022-08-19.json
    ├── mask
    | └──CelebAMask-HQ-mask-color-palette_32_nearest_downsampled_from_hq_512_one_hot_2d_tensor
    └── sketch
    └──sketch_1x1024_tensor
    

For more details about the annotations, please refer to CelebA-Dialog.

Generation

You can control face generation using text and segmentation mask.

  1. mask_path is the path to the segmentation mask, and input_text is the text condition.

    python generate_512.py \
    --mask_path test_data/512_masks/27007.png \
    --input_text "This man has beard of medium length. He is in his thirties."
    python generate_512.py \
    --mask_path test_data/512_masks/29980.png \
    --input_text "This woman is in her forties."
  2. You can view different types of intermediate outputs by setting the flags as 1. For example, to view the influence functions, you can set return_influence_function to 1.

    python generate_512.py \
    --mask_path test_data/512_masks/27007.png \
    --input_text "This man has beard of medium length. He is in his thirties." \
    --ddim_steps 10 \
    --batch_size 1 \
    --save_z 1 \
    --return_influence_function 1 \
    --display_x_inter 1 \
    --save_mixed 1

    Note that producing intermediate results might consume a lot of GPU memory, so we suggest setting batch_size to 1, and setting ddim_steps to a smaller value (e.g., 10) to save memory and computation time.

Training

We provide the entire training pipeline, including training the VAE, uni-modal diffusion models, and our proposed dynamic diffusers.

If you are only interested in training dynamic diffusers, you can use our provided checkpoints for VAE and uni-modal diffusion models. Simply skip step 1 and 2 and directly look at step 3.

  1. Train VAE.

    LDM compresses images to the VAE latents to save computational cost, and later train UNet diffusion models on the VAE latents. This step is to reproduce the pretrained/512_vae.ckpt.

    python main.py \
    --logdir 'outputs/512_vae' \
    --base 'configs/512_vae.yaml' \
    -t --gpus 0,1,2,3,
  2. Train the uni-modal diffusion models.

    (1) train text-to-image model. This step is to reproduce the pretrained/512_text.ckpt.

    python main.py \
    --logdir 'outputs/512_text' \
    --base 'configs/512_text.yaml' \
    -t --gpus 0,1,2,3,

    (2) train mask-to-image model. This step is to reproduce the pretrained/512_mask.ckpt.

    python main.py \
    --logdir 'outputs/512_mask' \
    --base 'configs/512_mask.yaml' \
    -t --gpus 0,1,2,3,
  3. Train the dynamic diffusers.

    The dynamic diffusers are the meta-networks that determine how the uni-modal diffusion models collaborate together. This step is to reproduce the pretrained/512_codiff_mask_text.ckpt.

    python main.py \
    --logdir 'outputs/512_codiff_mask_text' \
    --base 'configs/512_codiff_mask_text.yaml' \
    -t --gpus 0,1,2,3,

Citation

If you find our repo useful for your research, please consider citing our paper:

@InProceedings{huang2023collaborative,
author = {Huang, Ziqi and Chan, Kelvin C.K. and Jiang, Yuming and Liu, Ziwei},
title = {Collaborative Diffusion for Multi-Modal Face Generation and Editing},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year = {2023},
}

Acknowledgement

The codebase is maintained by Ziqi Huang.

This project is built on top of LDM. We trained on data provided by CelebA-HQ, CelebA-Dialog, CelebAMask-HQ, and MM-CelebA-HQ-Dataset.

About

Collaborative Diffusion (CVPR 2023)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Collaborative Diffusion (CVPR 2023)

This repository contains the implementation of the following paper:

Collaborative Diffusion for Multi-Modal Face Generation and Editing
Ziqi Huang, Kelvin C.K. Chan, Yuming Jiang, Ziwei Liu
IEEE/CVF International Conference on Computer Vision (CVPR), 2023

From MMLab@NTU affiliated with S-Lab, Nanyang Technological University

[Paper] | [Project Page] | [Video]

Overview

We propose Collaborative Diffusion, where users can use multiple modalities to control face generation and editing. (a) Face Generation. Given multi-modal controls, our framework synthesizes high-quality images consistent with the input conditions. (b) Face Editing. Collaborative Diffusion also supports multi-modal editing of real images with promising identity preservation capability.


We use pre-trained uni-modal diffusion models to perform multi-modal guided face generation and editing. At each step of the reverse process (i.e., from timestep t to t − 1), the dynamic diffuser predicts the spatial-varying and temporal-varying influence function to selectively enhance or suppress the contributions of the given modality.

Updates

  • [06/2023] We provide the preprocessed multi-modal annotations here.
  • [05/2023] Training code for Collaborative Diffusion (512x512) released.
  • [04/2023] Project page and video available.
  • [04/2023] Arxiv paper available.
  • [04/2023] Checkpoints for multi-modal face generation (512x512) released.
  • [04/2023] Inference code for multi-modal face generation (512x512) released.

Installation

  1. Clone repo

    git clone https://github.com/ziqihuangg/Collaborative-Diffusion
    cd Collaborative-Diffusion
  2. Create conda environment.
    If you already have an ldm environment installed according to LDM, you do not need to go throught this step (i.e., step 2). You can simply conda activate ldm and jump to step 3.

     conda env create -f environment.yaml
    conda activate codiff
  3. Install dependencies

     pip install transformers==4.19.2 scann kornia==0.6.4 torchmetrics==0.6.0
    conda install -c anaconda git
    pip install git+https://github.com/arogozhnikov/einops.git

Download

Download Checkpoints

  1. Download the pre-trained models from here.

  2. Put the models under pretrained as follows:

    Collaborative-Diffusion
    └── pretrained
    ├── 512_codiff_mask_text.ckpt
    ├── 512_mask.ckpt
    ├── 512_text.ckpt
    └── 512_vae.ckpt
    

Download Datasets

We provide preprocessed data used in this project (see Acknowledgement for data source). You need download them only if you want to reproduce the training of Collaborative Diffusion. You can skip this step if you simply want to use our pre-trained models for inference.

  1. Download the preprocessed training data from here.

  2. Put the datasets under dataset as follows:

    Collaborative-Diffusion
    └── dataset
    ├── image
    | └──image_512_downsampled_from_hq_1024
    ├── text
    | └──captions_hq_beard_and_age_2022-08-19.json
    ├── mask
    | └──CelebAMask-HQ-mask-color-palette_32_nearest_downsampled_from_hq_512_one_hot_2d_tensor
    └── sketch
    └──sketch_1x1024_tensor
    

For more details about the annotations, please refer to CelebA-Dialog.

Generation

You can control face generation using text and segmentation mask.

  1. mask_path is the path to the segmentation mask, and input_text is the text condition.

    python generate_512.py \
    --mask_path test_data/512_masks/27007.png \
    --input_text "This man has beard of medium length. He is in his thirties."
    python generate_512.py \
    --mask_path test_data/512_masks/29980.png \
    --input_text "This woman is in her forties."
  2. You can view different types of intermediate outputs by setting the flags as 1. For example, to view the influence functions, you can set return_influence_function to 1.

    python generate_512.py \
    --mask_path test_data/512_masks/27007.png \
    --input_text "This man has beard of medium length. He is in his thirties." \
    --ddim_steps 10 \
    --batch_size 1 \
    --save_z 1 \
    --return_influence_function 1 \
    --display_x_inter 1 \
    --save_mixed 1

    Note that producing intermediate results might consume a lot of GPU memory, so we suggest setting batch_size to 1, and setting ddim_steps to a smaller value (e.g., 10) to save memory and computation time.

Training

We provide the entire training pipeline, including training the VAE, uni-modal diffusion models, and our proposed dynamic diffusers.

If you are only interested in training dynamic diffusers, you can use our provided checkpoints for VAE and uni-modal diffusion models. Simply skip step 1 and 2 and directly look at step 3.

  1. Train VAE.

    LDM compresses images to the VAE latents to save computational cost, and later train UNet diffusion models on the VAE latents. This step is to reproduce the pretrained/512_vae.ckpt.

    python main.py \
    --logdir 'outputs/512_vae' \
    --base 'configs/512_vae.yaml' \
    -t --gpus 0,1,2,3,
  2. Train the uni-modal diffusion models.

    (1) train text-to-image model. This step is to reproduce the pretrained/512_text.ckpt.

    python main.py \
    --logdir 'outputs/512_text' \
    --base 'configs/512_text.yaml' \
    -t --gpus 0,1,2,3,

    (2) train mask-to-image model. This step is to reproduce the pretrained/512_mask.ckpt.

    python main.py \
    --logdir 'outputs/512_mask' \
    --base 'configs/512_mask.yaml' \
    -t --gpus 0,1,2,3,
  3. Train the dynamic diffusers.

    The dynamic diffusers are the meta-networks that determine how the uni-modal diffusion models collaborate together. This step is to reproduce the pretrained/512_codiff_mask_text.ckpt.

    python main.py \
    --logdir 'outputs/512_codiff_mask_text' \
    --base 'configs/512_codiff_mask_text.yaml' \
    -t --gpus 0,1,2,3,

Citation

If you find our repo useful for your research, please consider citing our paper:

@InProceedings{huang2023collaborative,
author = {Huang, Ziqi and Chan, Kelvin C.K. and Jiang, Yuming and Liu, Ziwei},
title = {Collaborative Diffusion for Multi-Modal Face Generation and Editing},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year = {2023},
}

Acknowledgement

The codebase is maintained by Ziqi Huang.

This project is built on top of LDM. We trained on data provided by CelebA-HQ, CelebA-Dialog, CelebAMask-HQ, and MM-CelebA-HQ-Dataset.

About

Collaborative Diffusion (CVPR 2023)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Collaborative Diffusion (CVPR 2023)

This repository contains the implementation of the following paper:

Collaborative Diffusion for Multi-Modal Face Generation and Editing
Ziqi Huang, Kelvin C.K. Chan, Yuming Jiang, Ziwei Liu
IEEE/CVF International Conference on Computer Vision (CVPR), 2023

From MMLab@NTU affiliated with S-Lab, Nanyang Technological University

[Paper] | [Project Page] | [Video]

Overview

We propose Collaborative Diffusion, where users can use multiple modalities to control face generation and editing. (a) Face Generation. Given multi-modal controls, our framework synthesizes high-quality images consistent with the input conditions. (b) Face Editing. Collaborative Diffusion also supports multi-modal editing of real images with promising identity preservation capability.


We use pre-trained uni-modal diffusion models to perform multi-modal guided face generation and editing. At each step of the reverse process (i.e., from timestep t to t − 1), the dynamic diffuser predicts the spatial-varying and temporal-varying influence function to selectively enhance or suppress the contributions of the given modality.

Updates

  • [06/2023] We provide the preprocessed multi-modal annotations here.
  • [05/2023] Training code for Collaborative Diffusion (512x512) released.
  • [04/2023] Project page and video available.
  • [04/2023] Arxiv paper available.
  • [04/2023] Checkpoints for multi-modal face generation (512x512) released.
  • [04/2023] Inference code for multi-modal face generation (512x512) released.

Installation

  1. Clone repo

    git clone https://github.com/ziqihuangg/Collaborative-Diffusion
    cd Collaborative-Diffusion
  2. Create conda environment.
    If you already have an ldm environment installed according to LDM, you do not need to go throught this step (i.e., step 2). You can simply conda activate ldm and jump to step 3.

     conda env create -f environment.yaml
    conda activate codiff
  3. Install dependencies

     pip install transformers==4.19.2 scann kornia==0.6.4 torchmetrics==0.6.0
    conda install -c anaconda git
    pip install git+https://github.com/arogozhnikov/einops.git

Download

Download Checkpoints

  1. Download the pre-trained models from here.

  2. Put the models under pretrained as follows:

    Collaborative-Diffusion
    └── pretrained
    ├── 512_codiff_mask_text.ckpt
    ├── 512_mask.ckpt
    ├── 512_text.ckpt
    └── 512_vae.ckpt
    

Download Datasets

We provide preprocessed data used in this project (see Acknowledgement for data source). You need download them only if you want to reproduce the training of Collaborative Diffusion. You can skip this step if you simply want to use our pre-trained models for inference.

  1. Download the preprocessed training data from here.

  2. Put the datasets under dataset as follows:

    Collaborative-Diffusion
    └── dataset
    ├── image
    | └──image_512_downsampled_from_hq_1024
    ├── text
    | └──captions_hq_beard_and_age_2022-08-19.json
    ├── mask
    | └──CelebAMask-HQ-mask-color-palette_32_nearest_downsampled_from_hq_512_one_hot_2d_tensor
    └── sketch
    └──sketch_1x1024_tensor
    

For more details about the annotations, please refer to CelebA-Dialog.

Generation

You can control face generation using text and segmentation mask.

  1. mask_path is the path to the segmentation mask, and input_text is the text condition.

    python generate_512.py \
    --mask_path test_data/512_masks/27007.png \
    --input_text "This man has beard of medium length. He is in his thirties."
    python generate_512.py \
    --mask_path test_data/512_masks/29980.png \
    --input_text "This woman is in her forties."
  2. You can view different types of intermediate outputs by setting the flags as 1. For example, to view the influence functions, you can set return_influence_function to 1.

    python generate_512.py \
    --mask_path test_data/512_masks/27007.png \
    --input_text "This man has beard of medium length. He is in his thirties." \
    --ddim_steps 10 \
    --batch_size 1 \
    --save_z 1 \
    --return_influence_function 1 \
    --display_x_inter 1 \
    --save_mixed 1

    Note that producing intermediate results might consume a lot of GPU memory, so we suggest setting batch_size to 1, and setting ddim_steps to a smaller value (e.g., 10) to save memory and computation time.

Training

We provide the entire training pipeline, including training the VAE, uni-modal diffusion models, and our proposed dynamic diffusers.

If you are only interested in training dynamic diffusers, you can use our provided checkpoints for VAE and uni-modal diffusion models. Simply skip step 1 and 2 and directly look at step 3.

  1. Train VAE.

    LDM compresses images to the VAE latents to save computational cost, and later train UNet diffusion models on the VAE latents. This step is to reproduce the pretrained/512_vae.ckpt.

    python main.py \
    --logdir 'outputs/512_vae' \
    --base 'configs/512_vae.yaml' \
    -t --gpus 0,1,2,3,
  2. Train the uni-modal diffusion models.

    (1) train text-to-image model. This step is to reproduce the pretrained/512_text.ckpt.

    python main.py \
    --logdir 'outputs/512_text' \
    --base 'configs/512_text.yaml' \
    -t --gpus 0,1,2,3,

    (2) train mask-to-image model. This step is to reproduce the pretrained/512_mask.ckpt.

    python main.py \
    --logdir 'outputs/512_mask' \
    --base 'configs/512_mask.yaml' \
    -t --gpus 0,1,2,3,
  3. Train the dynamic diffusers.

    The dynamic diffusers are the meta-networks that determine how the uni-modal diffusion models collaborate together. This step is to reproduce the pretrained/512_codiff_mask_text.ckpt.

    python main.py \
    --logdir 'outputs/512_codiff_mask_text' \
    --base 'configs/512_codiff_mask_text.yaml' \
    -t --gpus 0,1,2,3,

Citation

If you find our repo useful for your research, please consider citing our paper:

@InProceedings{huang2023collaborative,
author = {Huang, Ziqi and Chan, Kelvin C.K. and Jiang, Yuming and Liu, Ziwei},
title = {Collaborative Diffusion for Multi-Modal Face Generation and Editing},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year = {2023},
}

Acknowledgement

The codebase is maintained by Ziqi Huang.

This project is built on top of LDM. We trained on data provided by CelebA-HQ, CelebA-Dialog, CelebAMask-HQ, and MM-CelebA-HQ-Dataset.

About

Collaborative Diffusion (CVPR 2023)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Collaborative Diffusion (CVPR 2023)

This repository contains the implementation of the following paper:

Collaborative Diffusion for Multi-Modal Face Generation and Editing
Ziqi Huang, Kelvin C.K. Chan, Yuming Jiang, Ziwei Liu
IEEE/CVF International Conference on Computer Vision (CVPR), 2023

From MMLab@NTU affiliated with S-Lab, Nanyang Technological University

[Paper] | [Project Page] | [Video]

Overview

We propose Collaborative Diffusion, where users can use multiple modalities to control face generation and editing. (a) Face Generation. Given multi-modal controls, our framework synthesizes high-quality images consistent with the input conditions. (b) Face Editing. Collaborative Diffusion also supports multi-modal editing of real images with promising identity preservation capability.


We use pre-trained uni-modal diffusion models to perform multi-modal guided face generation and editing. At each step of the reverse process (i.e., from timestep t to t − 1), the dynamic diffuser predicts the spatial-varying and temporal-varying influence function to selectively enhance or suppress the contributions of the given modality.

Updates

  • [06/2023] We provide the preprocessed multi-modal annotations here.
  • [05/2023] Training code for Collaborative Diffusion (512x512) released.
  • [04/2023] Project page and video available.
  • [04/2023] Arxiv paper available.
  • [04/2023] Checkpoints for multi-modal face generation (512x512) released.
  • [04/2023] Inference code for multi-modal face generation (512x512) released.

Installation

  1. Clone repo

    git clone https://github.com/ziqihuangg/Collaborative-Diffusion
    cd Collaborative-Diffusion
  2. Create conda environment.
    If you already have an ldm environment installed according to LDM, you do not need to go throught this step (i.e., step 2). You can simply conda activate ldm and jump to step 3.

     conda env create -f environment.yaml
    conda activate codiff
  3. Install dependencies

     pip install transformers==4.19.2 scann kornia==0.6.4 torchmetrics==0.6.0
    conda install -c anaconda git
    pip install git+https://github.com/arogozhnikov/einops.git

Download

Download Checkpoints

  1. Download the pre-trained models from here.

  2. Put the models under pretrained as follows:

    Collaborative-Diffusion
    └── pretrained
    ├── 512_codiff_mask_text.ckpt
    ├── 512_mask.ckpt
    ├── 512_text.ckpt
    └── 512_vae.ckpt
    

Download Datasets

We provide preprocessed data used in this project (see Acknowledgement for data source). You need download them only if you want to reproduce the training of Collaborative Diffusion. You can skip this step if you simply want to use our pre-trained models for inference.

  1. Download the preprocessed training data from here.

  2. Put the datasets under dataset as follows:

    Collaborative-Diffusion
    └── dataset
    ├── image
    | └──image_512_downsampled_from_hq_1024
    ├── text
    | └──captions_hq_beard_and_age_2022-08-19.json
    ├── mask
    | └──CelebAMask-HQ-mask-color-palette_32_nearest_downsampled_from_hq_512_one_hot_2d_tensor
    └── sketch
    └──sketch_1x1024_tensor
    

For more details about the annotations, please refer to CelebA-Dialog.

Generation

You can control face generation using text and segmentation mask.

  1. mask_path is the path to the segmentation mask, and input_text is the text condition.

    python generate_512.py \
    --mask_path test_data/512_masks/27007.png \
    --input_text "This man has beard of medium length. He is in his thirties."
    python generate_512.py \
    --mask_path test_data/512_masks/29980.png \
    --input_text "This woman is in her forties."
  2. You can view different types of intermediate outputs by setting the flags as 1. For example, to view the influence functions, you can set return_influence_function to 1.

    python generate_512.py \
    --mask_path test_data/512_masks/27007.png \
    --input_text "This man has beard of medium length. He is in his thirties." \
    --ddim_steps 10 \
    --batch_size 1 \
    --save_z 1 \
    --return_influence_function 1 \
    --display_x_inter 1 \
    --save_mixed 1

    Note that producing intermediate results might consume a lot of GPU memory, so we suggest setting batch_size to 1, and setting ddim_steps to a smaller value (e.g., 10) to save memory and computation time.

Training

We provide the entire training pipeline, including training the VAE, uni-modal diffusion models, and our proposed dynamic diffusers.

If you are only interested in training dynamic diffusers, you can use our provided checkpoints for VAE and uni-modal diffusion models. Simply skip step 1 and 2 and directly look at step 3.

  1. Train VAE.

    LDM compresses images to the VAE latents to save computational cost, and later train UNet diffusion models on the VAE latents. This step is to reproduce the pretrained/512_vae.ckpt.

    python main.py \
    --logdir 'outputs/512_vae' \
    --base 'configs/512_vae.yaml' \
    -t --gpus 0,1,2,3,
  2. Train the uni-modal diffusion models.

    (1) train text-to-image model. This step is to reproduce the pretrained/512_text.ckpt.

    python main.py \
    --logdir 'outputs/512_text' \
    --base 'configs/512_text.yaml' \
    -t --gpus 0,1,2,3,

    (2) train mask-to-image model. This step is to reproduce the pretrained/512_mask.ckpt.

    python main.py \
    --logdir 'outputs/512_mask' \
    --base 'configs/512_mask.yaml' \
    -t --gpus 0,1,2,3,
  3. Train the dynamic diffusers.

    The dynamic diffusers are the meta-networks that determine how the uni-modal diffusion models collaborate together. This step is to reproduce the pretrained/512_codiff_mask_text.ckpt.

    python main.py \
    --logdir 'outputs/512_codiff_mask_text' \
    --base 'configs/512_codiff_mask_text.yaml' \
    -t --gpus 0,1,2,3,

Citation

If you find our repo useful for your research, please consider citing our paper:

@InProceedings{huang2023collaborative,
author = {Huang, Ziqi and Chan, Kelvin C.K. and Jiang, Yuming and Liu, Ziwei},
title = {Collaborative Diffusion for Multi-Modal Face Generation and Editing},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year = {2023},
}

Acknowledgement

The codebase is maintained by Ziqi Huang.

This project is built on top of LDM. We trained on data provided by CelebA-HQ, CelebA-Dialog, CelebAMask-HQ, and MM-CelebA-HQ-Dataset.

About

Collaborative Diffusion (CVPR 2023)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Collaborative Diffusion (CVPR 2023)

This repository contains the implementation of the following paper:

Collaborative Diffusion for Multi-Modal Face Generation and Editing
Ziqi Huang, Kelvin C.K. Chan, Yuming Jiang, Ziwei Liu
IEEE/CVF International Conference on Computer Vision (CVPR), 2023

From MMLab@NTU affiliated with S-Lab, Nanyang Technological University

[Paper] | [Project Page] | [Video]

Overview

We propose Collaborative Diffusion, where users can use multiple modalities to control face generation and editing. (a) Face Generation. Given multi-modal controls, our framework synthesizes high-quality images consistent with the input conditions. (b) Face Editing. Collaborative Diffusion also supports multi-modal editing of real images with promising identity preservation capability.


We use pre-trained uni-modal diffusion models to perform multi-modal guided face generation and editing. At each step of the reverse process (i.e., from timestep t to t − 1), the dynamic diffuser predicts the spatial-varying and temporal-varying influence function to selectively enhance or suppress the contributions of the given modality.

Updates

  • [06/2023] We provide the preprocessed multi-modal annotations here.
  • [05/2023] Training code for Collaborative Diffusion (512x512) released.
  • [04/2023] Project page and video available.
  • [04/2023] Arxiv paper available.
  • [04/2023] Checkpoints for multi-modal face generation (512x512) released.
  • [04/2023] Inference code for multi-modal face generation (512x512) released.

Installation

  1. Clone repo

    git clone https://github.com/ziqihuangg/Collaborative-Diffusion
    cd Collaborative-Diffusion
  2. Create conda environment.
    If you already have an ldm environment installed according to LDM, you do not need to go throught this step (i.e., step 2). You can simply conda activate ldm and jump to step 3.

     conda env create -f environment.yaml
    conda activate codiff
  3. Install dependencies

     pip install transformers==4.19.2 scann kornia==0.6.4 torchmetrics==0.6.0
    conda install -c anaconda git
    pip install git+https://github.com/arogozhnikov/einops.git

Download

Download Checkpoints

  1. Download the pre-trained models from here.

  2. Put the models under pretrained as follows:

    Collaborative-Diffusion
    └── pretrained
    ├── 512_codiff_mask_text.ckpt
    ├── 512_mask.ckpt
    ├── 512_text.ckpt
    └── 512_vae.ckpt
    

Download Datasets

We provide preprocessed data used in this project (see Acknowledgement for data source). You need download them only if you want to reproduce the training of Collaborative Diffusion. You can skip this step if you simply want to use our pre-trained models for inference.

  1. Download the preprocessed training data from here.

  2. Put the datasets under dataset as follows:

    Collaborative-Diffusion
    └── dataset
    ├── image
    | └──image_512_downsampled_from_hq_1024
    ├── text
    | └──captions_hq_beard_and_age_2022-08-19.json
    ├── mask
    | └──CelebAMask-HQ-mask-color-palette_32_nearest_downsampled_from_hq_512_one_hot_2d_tensor
    └── sketch
    └──sketch_1x1024_tensor
    

For more details about the annotations, please refer to CelebA-Dialog.

Generation

You can control face generation using text and segmentation mask.

  1. mask_path is the path to the segmentation mask, and input_text is the text condition.

    python generate_512.py \
    --mask_path test_data/512_masks/27007.png \
    --input_text "This man has beard of medium length. He is in his thirties."
    python generate_512.py \
    --mask_path test_data/512_masks/29980.png \
    --input_text "This woman is in her forties."
  2. You can view different types of intermediate outputs by setting the flags as 1. For example, to view the influence functions, you can set return_influence_function to 1.

    python generate_512.py \
    --mask_path test_data/512_masks/27007.png \
    --input_text "This man has beard of medium length. He is in his thirties." \
    --ddim_steps 10 \
    --batch_size 1 \
    --save_z 1 \
    --return_influence_function 1 \
    --display_x_inter 1 \
    --save_mixed 1

    Note that producing intermediate results might consume a lot of GPU memory, so we suggest setting batch_size to 1, and setting ddim_steps to a smaller value (e.g., 10) to save memory and computation time.

Training

We provide the entire training pipeline, including training the VAE, uni-modal diffusion models, and our proposed dynamic diffusers.

If you are only interested in training dynamic diffusers, you can use our provided checkpoints for VAE and uni-modal diffusion models. Simply skip step 1 and 2 and directly look at step 3.

  1. Train VAE.

    LDM compresses images to the VAE latents to save computational cost, and later train UNet diffusion models on the VAE latents. This step is to reproduce the pretrained/512_vae.ckpt.

    python main.py \
    --logdir 'outputs/512_vae' \
    --base 'configs/512_vae.yaml' \
    -t --gpus 0,1,2,3,
  2. Train the uni-modal diffusion models.

    (1) train text-to-image model. This step is to reproduce the pretrained/512_text.ckpt.

    python main.py \
    --logdir 'outputs/512_text' \
    --base 'configs/512_text.yaml' \
    -t --gpus 0,1,2,3,

    (2) train mask-to-image model. This step is to reproduce the pretrained/512_mask.ckpt.

    python main.py \
    --logdir 'outputs/512_mask' \
    --base 'configs/512_mask.yaml' \
    -t --gpus 0,1,2,3,
  3. Train the dynamic diffusers.

    The dynamic diffusers are the meta-networks that determine how the uni-modal diffusion models collaborate together. This step is to reproduce the pretrained/512_codiff_mask_text.ckpt.

    python main.py \
    --logdir 'outputs/512_codiff_mask_text' \
    --base 'configs/512_codiff_mask_text.yaml' \
    -t --gpus 0,1,2,3,

Citation

If you find our repo useful for your research, please consider citing our paper:

@InProceedings{huang2023collaborative,
author = {Huang, Ziqi and Chan, Kelvin C.K. and Jiang, Yuming and Liu, Ziwei},
title = {Collaborative Diffusion for Multi-Modal Face Generation and Editing},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year = {2023},
}

Acknowledgement

The codebase is maintained by Ziqi Huang.

This project is built on top of LDM. We trained on data provided by CelebA-HQ, CelebA-Dialog, CelebAMask-HQ, and MM-CelebA-HQ-Dataset.

About

Collaborative Diffusion (CVPR 2023)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Collaborative Diffusion (CVPR 2023)

This repository contains the implementation of the following paper:

Collaborative Diffusion for Multi-Modal Face Generation and Editing
Ziqi Huang, Kelvin C.K. Chan, Yuming Jiang, Ziwei Liu
IEEE/CVF International Conference on Computer Vision (CVPR), 2023

From MMLab@NTU affiliated with S-Lab, Nanyang Technological University

[Paper] | [Project Page] | [Video]

Overview

We propose Collaborative Diffusion, where users can use multiple modalities to control face generation and editing. (a) Face Generation. Given multi-modal controls, our framework synthesizes high-quality images consistent with the input conditions. (b) Face Editing. Collaborative Diffusion also supports multi-modal editing of real images with promising identity preservation capability.


We use pre-trained uni-modal diffusion models to perform multi-modal guided face generation and editing. At each step of the reverse process (i.e., from timestep t to t − 1), the dynamic diffuser predicts the spatial-varying and temporal-varying influence function to selectively enhance or suppress the contributions of the given modality.

Updates

  • [06/2023] We provide the preprocessed multi-modal annotations here.
  • [05/2023] Training code for Collaborative Diffusion (512x512) released.
  • [04/2023] Project page and video available.
  • [04/2023] Arxiv paper available.
  • [04/2023] Checkpoints for multi-modal face generation (512x512) released.
  • [04/2023] Inference code for multi-modal face generation (512x512) released.

Installation

  1. Clone repo

    git clone https://github.com/ziqihuangg/Collaborative-Diffusion
    cd Collaborative-Diffusion
  2. Create conda environment.
    If you already have an ldm environment installed according to LDM, you do not need to go throught this step (i.e., step 2). You can simply conda activate ldm and jump to step 3.

     conda env create -f environment.yaml
    conda activate codiff
  3. Install dependencies

     pip install transformers==4.19.2 scann kornia==0.6.4 torchmetrics==0.6.0
    conda install -c anaconda git
    pip install git+https://github.com/arogozhnikov/einops.git

Download

Download Checkpoints

  1. Download the pre-trained models from here.

  2. Put the models under pretrained as follows:

    Collaborative-Diffusion
    └── pretrained
    ├── 512_codiff_mask_text.ckpt
    ├── 512_mask.ckpt
    ├── 512_text.ckpt
    └── 512_vae.ckpt
    

Download Datasets

We provide preprocessed data used in this project (see Acknowledgement for data source). You need download them only if you want to reproduce the training of Collaborative Diffusion. You can skip this step if you simply want to use our pre-trained models for inference.

  1. Download the preprocessed training data from here.

  2. Put the datasets under dataset as follows:

    Collaborative-Diffusion
    └── dataset
    ├── image
    | └──image_512_downsampled_from_hq_1024
    ├── text
    | └──captions_hq_beard_and_age_2022-08-19.json
    ├── mask
    | └──CelebAMask-HQ-mask-color-palette_32_nearest_downsampled_from_hq_512_one_hot_2d_tensor
    └── sketch
    └──sketch_1x1024_tensor
    

For more details about the annotations, please refer to CelebA-Dialog.

Generation

You can control face generation using text and segmentation mask.

  1. mask_path is the path to the segmentation mask, and input_text is the text condition.

    python generate_512.py \
    --mask_path test_data/512_masks/27007.png \
    --input_text "This man has beard of medium length. He is in his thirties."
    python generate_512.py \
    --mask_path test_data/512_masks/29980.png \
    --input_text "This woman is in her forties."
  2. You can view different types of intermediate outputs by setting the flags as 1. For example, to view the influence functions, you can set return_influence_function to 1.

    python generate_512.py \
    --mask_path test_data/512_masks/27007.png \
    --input_text "This man has beard of medium length. He is in his thirties." \
    --ddim_steps 10 \
    --batch_size 1 \
    --save_z 1 \
    --return_influence_function 1 \
    --display_x_inter 1 \
    --save_mixed 1

    Note that producing intermediate results might consume a lot of GPU memory, so we suggest setting batch_size to 1, and setting ddim_steps to a smaller value (e.g., 10) to save memory and computation time.

Training

We provide the entire training pipeline, including training the VAE, uni-modal diffusion models, and our proposed dynamic diffusers.

If you are only interested in training dynamic diffusers, you can use our provided checkpoints for VAE and uni-modal diffusion models. Simply skip step 1 and 2 and directly look at step 3.

  1. Train VAE.

    LDM compresses images to the VAE latents to save computational cost, and later train UNet diffusion models on the VAE latents. This step is to reproduce the pretrained/512_vae.ckpt.

    python main.py \
    --logdir 'outputs/512_vae' \
    --base 'configs/512_vae.yaml' \
    -t --gpus 0,1,2,3,
  2. Train the uni-modal diffusion models.

    (1) train text-to-image model. This step is to reproduce the pretrained/512_text.ckpt.

    python main.py \
    --logdir 'outputs/512_text' \
    --base 'configs/512_text.yaml' \
    -t --gpus 0,1,2,3,

    (2) train mask-to-image model. This step is to reproduce the pretrained/512_mask.ckpt.

    python main.py \
    --logdir 'outputs/512_mask' \
    --base 'configs/512_mask.yaml' \
    -t --gpus 0,1,2,3,
  3. Train the dynamic diffusers.

    The dynamic diffusers are the meta-networks that determine how the uni-modal diffusion models collaborate together. This step is to reproduce the pretrained/512_codiff_mask_text.ckpt.

    python main.py \
    --logdir 'outputs/512_codiff_mask_text' \
    --base 'configs/512_codiff_mask_text.yaml' \
    -t --gpus 0,1,2,3,

Citation

If you find our repo useful for your research, please consider citing our paper:

@InProceedings{huang2023collaborative,
author = {Huang, Ziqi and Chan, Kelvin C.K. and Jiang, Yuming and Liu, Ziwei},
title = {Collaborative Diffusion for Multi-Modal Face Generation and Editing},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year = {2023},
}

Acknowledgement

The codebase is maintained by Ziqi Huang.

This project is built on top of LDM. We trained on data provided by CelebA-HQ, CelebA-Dialog, CelebAMask-HQ, and MM-CelebA-HQ-Dataset.

About

Collaborative Diffusion (CVPR 2023)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Collaborative Diffusion (CVPR 2023)

This repository contains the implementation of the following paper:

Collaborative Diffusion for Multi-Modal Face Generation and Editing
Ziqi Huang, Kelvin C.K. Chan, Yuming Jiang, Ziwei Liu
IEEE/CVF International Conference on Computer Vision (CVPR), 2023

From MMLab@NTU affiliated with S-Lab, Nanyang Technological University

[Paper] | [Project Page] | [Video]

Overview

We propose Collaborative Diffusion, where users can use multiple modalities to control face generation and editing. (a) Face Generation. Given multi-modal controls, our framework synthesizes high-quality images consistent with the input conditions. (b) Face Editing. Collaborative Diffusion also supports multi-modal editing of real images with promising identity preservation capability.


We use pre-trained uni-modal diffusion models to perform multi-modal guided face generation and editing. At each step of the reverse process (i.e., from timestep t to t − 1), the dynamic diffuser predicts the spatial-varying and temporal-varying influence function to selectively enhance or suppress the contributions of the given modality.

Updates

  • [06/2023] We provide the preprocessed multi-modal annotations here.
  • [05/2023] Training code for Collaborative Diffusion (512x512) released.
  • [04/2023] Project page and video available.
  • [04/2023] Arxiv paper available.
  • [04/2023] Checkpoints for multi-modal face generation (512x512) released.
  • [04/2023] Inference code for multi-modal face generation (512x512) released.

Installation

  1. Clone repo

    git clone https://github.com/ziqihuangg/Collaborative-Diffusion
    cd Collaborative-Diffusion
  2. Create conda environment.
    If you already have an ldm environment installed according to LDM, you do not need to go throught this step (i.e., step 2). You can simply conda activate ldm and jump to step 3.

     conda env create -f environment.yaml
    conda activate codiff
  3. Install dependencies

     pip install transformers==4.19.2 scann kornia==0.6.4 torchmetrics==0.6.0
    conda install -c anaconda git
    pip install git+https://github.com/arogozhnikov/einops.git

Download

Download Checkpoints

  1. Download the pre-trained models from here.

  2. Put the models under pretrained as follows:

    Collaborative-Diffusion
    └── pretrained
    ├── 512_codiff_mask_text.ckpt
    ├── 512_mask.ckpt
    ├── 512_text.ckpt
    └── 512_vae.ckpt
    

Download Datasets

We provide preprocessed data used in this project (see Acknowledgement for data source). You need download them only if you want to reproduce the training of Collaborative Diffusion. You can skip this step if you simply want to use our pre-trained models for inference.

  1. Download the preprocessed training data from here.

  2. Put the datasets under dataset as follows:

    Collaborative-Diffusion
    └── dataset
    ├── image
    | └──image_512_downsampled_from_hq_1024
    ├── text
    | └──captions_hq_beard_and_age_2022-08-19.json
    ├── mask
    | └──CelebAMask-HQ-mask-color-palette_32_nearest_downsampled_from_hq_512_one_hot_2d_tensor
    └── sketch
    └──sketch_1x1024_tensor
    

For more details about the annotations, please refer to CelebA-Dialog.

Generation

You can control face generation using text and segmentation mask.

  1. mask_path is the path to the segmentation mask, and input_text is the text condition.

    python generate_512.py \
    --mask_path test_data/512_masks/27007.png \
    --input_text "This man has beard of medium length. He is in his thirties."
    python generate_512.py \
    --mask_path test_data/512_masks/29980.png \
    --input_text "This woman is in her forties."
  2. You can view different types of intermediate outputs by setting the flags as 1. For example, to view the influence functions, you can set return_influence_function to 1.

    python generate_512.py \
    --mask_path test_data/512_masks/27007.png \
    --input_text "This man has beard of medium length. He is in his thirties." \
    --ddim_steps 10 \
    --batch_size 1 \
    --save_z 1 \
    --return_influence_function 1 \
    --display_x_inter 1 \
    --save_mixed 1

    Note that producing intermediate results might consume a lot of GPU memory, so we suggest setting batch_size to 1, and setting ddim_steps to a smaller value (e.g., 10) to save memory and computation time.

Training

We provide the entire training pipeline, including training the VAE, uni-modal diffusion models, and our proposed dynamic diffusers.

If you are only interested in training dynamic diffusers, you can use our provided checkpoints for VAE and uni-modal diffusion models. Simply skip step 1 and 2 and directly look at step 3.

  1. Train VAE.

    LDM compresses images to the VAE latents to save computational cost, and later train UNet diffusion models on the VAE latents. This step is to reproduce the pretrained/512_vae.ckpt.

    python main.py \
    --logdir 'outputs/512_vae' \
    --base 'configs/512_vae.yaml' \
    -t --gpus 0,1,2,3,
  2. Train the uni-modal diffusion models.

    (1) train text-to-image model. This step is to reproduce the pretrained/512_text.ckpt.

    python main.py \
    --logdir 'outputs/512_text' \
    --base 'configs/512_text.yaml' \
    -t --gpus 0,1,2,3,

    (2) train mask-to-image model. This step is to reproduce the pretrained/512_mask.ckpt.

    python main.py \
    --logdir 'outputs/512_mask' \
    --base 'configs/512_mask.yaml' \
    -t --gpus 0,1,2,3,
  3. Train the dynamic diffusers.

    The dynamic diffusers are the meta-networks that determine how the uni-modal diffusion models collaborate together. This step is to reproduce the pretrained/512_codiff_mask_text.ckpt.

    python main.py \
    --logdir 'outputs/512_codiff_mask_text' \
    --base 'configs/512_codiff_mask_text.yaml' \
    -t --gpus 0,1,2,3,

Citation

If you find our repo useful for your research, please consider citing our paper:

@InProceedings{huang2023collaborative,
author = {Huang, Ziqi and Chan, Kelvin C.K. and Jiang, Yuming and Liu, Ziwei},
title = {Collaborative Diffusion for Multi-Modal Face Generation and Editing},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year = {2023},
}

Acknowledgement

The codebase is maintained by Ziqi Huang.

This project is built on top of LDM. We trained on data provided by CelebA-HQ, CelebA-Dialog, CelebAMask-HQ, and MM-CelebA-HQ-Dataset.

About

Collaborative Diffusion (CVPR 2023)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Collaborative Diffusion (CVPR 2023)

This repository contains the implementation of the following paper:

Collaborative Diffusion for Multi-Modal Face Generation and Editing
Ziqi Huang, Kelvin C.K. Chan, Yuming Jiang, Ziwei Liu
IEEE/CVF International Conference on Computer Vision (CVPR), 2023

From MMLab@NTU affiliated with S-Lab, Nanyang Technological University

[Paper] | [Project Page] | [Video]

Overview

We propose Collaborative Diffusion, where users can use multiple modalities to control face generation and editing. (a) Face Generation. Given multi-modal controls, our framework synthesizes high-quality images consistent with the input conditions. (b) Face Editing. Collaborative Diffusion also supports multi-modal editing of real images with promising identity preservation capability.


We use pre-trained uni-modal diffusion models to perform multi-modal guided face generation and editing. At each step of the reverse process (i.e., from timestep t to t − 1), the dynamic diffuser predicts the spatial-varying and temporal-varying influence function to selectively enhance or suppress the contributions of the given modality.

Updates

  • [06/2023] We provide the preprocessed multi-modal annotations here.
  • [05/2023] Training code for Collaborative Diffusion (512x512) released.
  • [04/2023] Project page and video available.
  • [04/2023] Arxiv paper available.
  • [04/2023] Checkpoints for multi-modal face generation (512x512) released.
  • [04/2023] Inference code for multi-modal face generation (512x512) released.

Installation

  1. Clone repo

    git clone https://github.com/ziqihuangg/Collaborative-Diffusion
    cd Collaborative-Diffusion
  2. Create conda environment.
    If you already have an ldm environment installed according to LDM, you do not need to go throught this step (i.e., step 2). You can simply conda activate ldm and jump to step 3.

     conda env create -f environment.yaml
    conda activate codiff
  3. Install dependencies

     pip install transformers==4.19.2 scann kornia==0.6.4 torchmetrics==0.6.0
    conda install -c anaconda git
    pip install git+https://github.com/arogozhnikov/einops.git

Download

Download Checkpoints

  1. Download the pre-trained models from here.

  2. Put the models under pretrained as follows:

    Collaborative-Diffusion
    └── pretrained
    ├── 512_codiff_mask_text.ckpt
    ├── 512_mask.ckpt
    ├── 512_text.ckpt
    └── 512_vae.ckpt
    

Download Datasets

We provide preprocessed data used in this project (see Acknowledgement for data source). You need download them only if you want to reproduce the training of Collaborative Diffusion. You can skip this step if you simply want to use our pre-trained models for inference.

  1. Download the preprocessed training data from here.

  2. Put the datasets under dataset as follows:

    Collaborative-Diffusion
    └── dataset
    ├── image
    | └──image_512_downsampled_from_hq_1024
    ├── text
    | └──captions_hq_beard_and_age_2022-08-19.json
    ├── mask
    | └──CelebAMask-HQ-mask-color-palette_32_nearest_downsampled_from_hq_512_one_hot_2d_tensor
    └── sketch
    └──sketch_1x1024_tensor
    

For more details about the annotations, please refer to CelebA-Dialog.

Generation

You can control face generation using text and segmentation mask.

  1. mask_path is the path to the segmentation mask, and input_text is the text condition.

    python generate_512.py \
    --mask_path test_data/512_masks/27007.png \
    --input_text "This man has beard of medium length. He is in his thirties."
    python generate_512.py \
    --mask_path test_data/512_masks/29980.png \
    --input_text "This woman is in her forties."
  2. You can view different types of intermediate outputs by setting the flags as 1. For example, to view the influence functions, you can set return_influence_function to 1.

    python generate_512.py \
    --mask_path test_data/512_masks/27007.png \
    --input_text "This man has beard of medium length. He is in his thirties." \
    --ddim_steps 10 \
    --batch_size 1 \
    --save_z 1 \
    --return_influence_function 1 \
    --display_x_inter 1 \
    --save_mixed 1

    Note that producing intermediate results might consume a lot of GPU memory, so we suggest setting batch_size to 1, and setting ddim_steps to a smaller value (e.g., 10) to save memory and computation time.

Training

We provide the entire training pipeline, including training the VAE, uni-modal diffusion models, and our proposed dynamic diffusers.

If you are only interested in training dynamic diffusers, you can use our provided checkpoints for VAE and uni-modal diffusion models. Simply skip step 1 and 2 and directly look at step 3.

  1. Train VAE.

    LDM compresses images to the VAE latents to save computational cost, and later train UNet diffusion models on the VAE latents. This step is to reproduce the pretrained/512_vae.ckpt.

    python main.py \
    --logdir 'outputs/512_vae' \
    --base 'configs/512_vae.yaml' \
    -t --gpus 0,1,2,3,
  2. Train the uni-modal diffusion models.

    (1) train text-to-image model. This step is to reproduce the pretrained/512_text.ckpt.

    python main.py \
    --logdir 'outputs/512_text' \
    --base 'configs/512_text.yaml' \
    -t --gpus 0,1,2,3,

    (2) train mask-to-image model. This step is to reproduce the pretrained/512_mask.ckpt.

    python main.py \
    --logdir 'outputs/512_mask' \
    --base 'configs/512_mask.yaml' \
    -t --gpus 0,1,2,3,
  3. Train the dynamic diffusers.

    The dynamic diffusers are the meta-networks that determine how the uni-modal diffusion models collaborate together. This step is to reproduce the pretrained/512_codiff_mask_text.ckpt.

    python main.py \
    --logdir 'outputs/512_codiff_mask_text' \
    --base 'configs/512_codiff_mask_text.yaml' \
    -t --gpus 0,1,2,3,

Citation

If you find our repo useful for your research, please consider citing our paper:

@InProceedings{huang2023collaborative,
author = {Huang, Ziqi and Chan, Kelvin C.K. and Jiang, Yuming and Liu, Ziwei},
title = {Collaborative Diffusion for Multi-Modal Face Generation and Editing},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year = {2023},
}

Acknowledgement

The codebase is maintained by Ziqi Huang.

This project is built on top of LDM. We trained on data provided by CelebA-HQ, CelebA-Dialog, CelebAMask-HQ, and MM-CelebA-HQ-Dataset.

About

Collaborative Diffusion (CVPR 2023)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages