Repository files navigation

Contrastive Mixture of Posteriors (CoMP) for Counterfactual Inference, Data Integration and Fairness

This is the PyTorch implementation of Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and Fairness ICML 2022.

CoMP Illustration

Requirements

System requirements

The following environment was used for the experiments, based on the nvidia/cuda:11.0.3-cudnn8-devel-ubuntu20.04 docker image from dockerhub:

  • Ubuntu 20.04
  • CUDA 11.0.3, cudnn 8
  • Python 3.8

Python environment

Running in a virtualenv or container is recommended. To install requirements:

pip install -U pip wheel
pip install -r requirements.txt

Datasets

The model expects two files -- a data file named features.tsv and a file with labels named metadata.tsv, with sample indices as the first column in both files, and with headers as the first row. We set the conditional variable $c$ to have column name "type" in the metadata file.

The data is prepared in two steps.

Step 1: Download source datasets

First the source datasets need to be downloaded from the following sources and placed into a folder <data-dir>.

Tumour / Cell Line Dataset

There are four files to download and prepare, as follows:

  1. Metadata file: The Celligner_info.csv file is taken from here. Save the file as <data-dir>/Celligner_info.csv.
  2. HGNC gene names: The file hgnc_complete_set_7.24.2018.txt can be downloaded from here. Save the file as <data-dir>/hgnc_complete_set_7.24.2018.txt.
  3. Tumour gene expression data: The TPM expression values can be downloaded from here. Rename this file to <data-dir>/TCGA_mat.tsv.
  4. Cell Line gene expression data: The data is downloaded from DepMap Public 19Q4 file: CCLE_expression_full.csv here. Rename this file to <data-dir>/CCLE_mat.csv.

Single cell PBMCs

Data was downloaded from the theislab/trVAE_reproducibility repository. Download kang_count.h5ad from the Google Drive link in the Getting Started section of the README to <data-dir>/kang_count.h5ad.

UCI Adult Income

Data was downloaded from the UCI Machine Learning Repository. Download all adult.{data,names,test} files from the data directory to the <data-dir> directory.

Step 2: Process data

Next we process the data prior to training the model. Specify a separate output folder <processed-data-x> for each of the datasets and run the following:

# Tumour / Cell line data
python3 data/process_celligner_data.py --input-dir <data-dir> --output-dir <processed-data-tcl> --top-var-number 8000
# Kang et al. PBMC scRNA-seq under INFb stimulation
python3 data/process_kang.py --input-dir <data-dir> --output-dir <processed-data-kang> --top-var-number 2000
# UCI Income data
python3 data/process_uci_income.py --input-dir <data-dir> --output-dir <processed-data-uci>

Training

Evaluation metrics are computed at the end of the training loop for the best checkpoint (by validation loss). In all the commands below the <processed-data> below should be the path the directory containing the files listed in the previous section. <output-dir> should be the desired directory in which to save the outputs. It will be created if it does not exist.

To run the code using only the CPU, omit the --use-cuda flag in the arguments to the run script.

Tumor / cell-line

python3 run.py --data-dir <processed-data-tcl> --output-dir <output-dir> --dataset tumour_cl --model comp --hidden-dim 512 --latent-dim 16 --num-layers 3 --use-batchnorm 1 --batch-size 5500 --num-epochs 4000 --learning-rate 0.0001 --penalty-scale 0.5 --kl-beta 1e-07 --seed 80244971 --use-cuda

Single cell PBMCs

python3 run.py --data-dir <processed-data-kang> --output-dir <output-dir> --dataset kang --model comp --hidden-dim 512 --latent-dim 40 --num-layers 3 --use-batchnorm 1 --batch-size 512 --num-epochs 10000 --learning-rate 1e-06 --penalty-scale 1.0 --kl-beta 1e-07 --seed 196117 --use-cuda

UCI Adult Income

python3 run.py --data-dir <processed-data-uci> --output-dir <output-dir> --dataset uci-income --model comp --hidden-dim 64 --latent-dim 16 --num-layers 2 --use-batchnorm 1 --batch-size 4096 --num-epochs 10000 --learning-rate 0.0001 --kl-beta 1.0 --penalty-scale 0.5 --seed 116983357 --use-cuda

Evaluation

An example notebook to evaluate the summary metrics is given in evaluation.ipynb.

Results

Our model achieves the following performance on :

Tumour / Cell Line Dataset

Modelsilhouettekbetmean-silhouettemean-kbet
VAE0.6580.9740.8030.581
CVAE0.5540.9310.6840.571
VFAE0.1680.2580.1980.188
trVAE0.0960.1630.1380.123
Celligner0.0820.5250.5680.226
CoMP (ours)0.0230.1600.0940.101

UCI Adult Income

ModelGender Acc.Income Acc.silhouettekbet
Original data0.7960.8490.0670.786
VAE0.7640.8120.0540.748
CVAE0.7780.8190.0540.724
VFAE-sampled0.6800.815--
VFAE-mean0.7890.8050.0460.571
trVAE0.6980.8080.0660.731
CoMP (ours)0.6790.8050.0110.451

Reference

If you find this code useful, do cite the following paper in your publication:

@inproceedings{foster2022contrastive,
title={Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and Fairness},
author={Foster, Adam and Vez{\'e}r, {\'A}rpi and Glastonbury, Craig A and Creed, P{\'a}id{\'\i} and Abujudeh, Samer and Sim, Aaron},
booktitle={International Conference on Machine Learning},
pages={6578--6621},
year={2022},
organization={PMLR}
}

About

CoMP: Contrastive Mixture of Posteriors

Resources

Stars

10 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Contrastive Mixture of Posteriors (CoMP) for Counterfactual Inference, Data Integration and Fairness

This is the PyTorch implementation of Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and Fairness ICML 2022.

CoMP Illustration

Requirements

System requirements

The following environment was used for the experiments, based on the nvidia/cuda:11.0.3-cudnn8-devel-ubuntu20.04 docker image from dockerhub:

  • Ubuntu 20.04
  • CUDA 11.0.3, cudnn 8
  • Python 3.8

Python environment

Running in a virtualenv or container is recommended. To install requirements:

pip install -U pip wheel
pip install -r requirements.txt

Datasets

The model expects two files -- a data file named features.tsv and a file with labels named metadata.tsv, with sample indices as the first column in both files, and with headers as the first row. We set the conditional variable $c$ to have column name "type" in the metadata file.

The data is prepared in two steps.

Step 1: Download source datasets

First the source datasets need to be downloaded from the following sources and placed into a folder <data-dir>.

Tumour / Cell Line Dataset

There are four files to download and prepare, as follows:

  1. Metadata file: The Celligner_info.csv file is taken from here. Save the file as <data-dir>/Celligner_info.csv.
  2. HGNC gene names: The file hgnc_complete_set_7.24.2018.txt can be downloaded from here. Save the file as <data-dir>/hgnc_complete_set_7.24.2018.txt.
  3. Tumour gene expression data: The TPM expression values can be downloaded from here. Rename this file to <data-dir>/TCGA_mat.tsv.
  4. Cell Line gene expression data: The data is downloaded from DepMap Public 19Q4 file: CCLE_expression_full.csv here. Rename this file to <data-dir>/CCLE_mat.csv.

Single cell PBMCs

Data was downloaded from the theislab/trVAE_reproducibility repository. Download kang_count.h5ad from the Google Drive link in the Getting Started section of the README to <data-dir>/kang_count.h5ad.

UCI Adult Income

Data was downloaded from the UCI Machine Learning Repository. Download all adult.{data,names,test} files from the data directory to the <data-dir> directory.

Step 2: Process data

Next we process the data prior to training the model. Specify a separate output folder <processed-data-x> for each of the datasets and run the following:

# Tumour / Cell line data
python3 data/process_celligner_data.py --input-dir <data-dir> --output-dir <processed-data-tcl> --top-var-number 8000
# Kang et al. PBMC scRNA-seq under INFb stimulation
python3 data/process_kang.py --input-dir <data-dir> --output-dir <processed-data-kang> --top-var-number 2000
# UCI Income data
python3 data/process_uci_income.py --input-dir <data-dir> --output-dir <processed-data-uci>

Training

Evaluation metrics are computed at the end of the training loop for the best checkpoint (by validation loss). In all the commands below the <processed-data> below should be the path the directory containing the files listed in the previous section. <output-dir> should be the desired directory in which to save the outputs. It will be created if it does not exist.

To run the code using only the CPU, omit the --use-cuda flag in the arguments to the run script.

Tumor / cell-line

python3 run.py --data-dir <processed-data-tcl> --output-dir <output-dir> --dataset tumour_cl --model comp --hidden-dim 512 --latent-dim 16 --num-layers 3 --use-batchnorm 1 --batch-size 5500 --num-epochs 4000 --learning-rate 0.0001 --penalty-scale 0.5 --kl-beta 1e-07 --seed 80244971 --use-cuda

Single cell PBMCs

python3 run.py --data-dir <processed-data-kang> --output-dir <output-dir> --dataset kang --model comp --hidden-dim 512 --latent-dim 40 --num-layers 3 --use-batchnorm 1 --batch-size 512 --num-epochs 10000 --learning-rate 1e-06 --penalty-scale 1.0 --kl-beta 1e-07 --seed 196117 --use-cuda

UCI Adult Income

python3 run.py --data-dir <processed-data-uci> --output-dir <output-dir> --dataset uci-income --model comp --hidden-dim 64 --latent-dim 16 --num-layers 2 --use-batchnorm 1 --batch-size 4096 --num-epochs 10000 --learning-rate 0.0001 --kl-beta 1.0 --penalty-scale 0.5 --seed 116983357 --use-cuda

Evaluation

An example notebook to evaluate the summary metrics is given in evaluation.ipynb.

Results

Our model achieves the following performance on :

Tumour / Cell Line Dataset

Modelsilhouettekbetmean-silhouettemean-kbet
VAE0.6580.9740.8030.581
CVAE0.5540.9310.6840.571
VFAE0.1680.2580.1980.188
trVAE0.0960.1630.1380.123
Celligner0.0820.5250.5680.226
CoMP (ours)0.0230.1600.0940.101

UCI Adult Income

ModelGender Acc.Income Acc.silhouettekbet
Original data0.7960.8490.0670.786
VAE0.7640.8120.0540.748
CVAE0.7780.8190.0540.724
VFAE-sampled0.6800.815--
VFAE-mean0.7890.8050.0460.571
trVAE0.6980.8080.0660.731
CoMP (ours)0.6790.8050.0110.451

Reference

If you find this code useful, do cite the following paper in your publication:

@inproceedings{foster2022contrastive,
title={Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and Fairness},
author={Foster, Adam and Vez{\'e}r, {\'A}rpi and Glastonbury, Craig A and Creed, P{\'a}id{\'\i} and Abujudeh, Samer and Sim, Aaron},
booktitle={International Conference on Machine Learning},
pages={6578--6621},
year={2022},
organization={PMLR}
}

About

CoMP: Contrastive Mixture of Posteriors

Resources

Stars

10 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Contrastive Mixture of Posteriors (CoMP) for Counterfactual Inference, Data Integration and Fairness

This is the PyTorch implementation of Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and Fairness ICML 2022.

CoMP Illustration

Requirements

System requirements

The following environment was used for the experiments, based on the nvidia/cuda:11.0.3-cudnn8-devel-ubuntu20.04 docker image from dockerhub:

  • Ubuntu 20.04
  • CUDA 11.0.3, cudnn 8
  • Python 3.8

Python environment

Running in a virtualenv or container is recommended. To install requirements:

pip install -U pip wheel
pip install -r requirements.txt

Datasets

The model expects two files -- a data file named features.tsv and a file with labels named metadata.tsv, with sample indices as the first column in both files, and with headers as the first row. We set the conditional variable $c$ to have column name "type" in the metadata file.

The data is prepared in two steps.

Step 1: Download source datasets

First the source datasets need to be downloaded from the following sources and placed into a folder <data-dir>.

Tumour / Cell Line Dataset

There are four files to download and prepare, as follows:

  1. Metadata file: The Celligner_info.csv file is taken from here. Save the file as <data-dir>/Celligner_info.csv.
  2. HGNC gene names: The file hgnc_complete_set_7.24.2018.txt can be downloaded from here. Save the file as <data-dir>/hgnc_complete_set_7.24.2018.txt.
  3. Tumour gene expression data: The TPM expression values can be downloaded from here. Rename this file to <data-dir>/TCGA_mat.tsv.
  4. Cell Line gene expression data: The data is downloaded from DepMap Public 19Q4 file: CCLE_expression_full.csv here. Rename this file to <data-dir>/CCLE_mat.csv.

Single cell PBMCs

Data was downloaded from the theislab/trVAE_reproducibility repository. Download kang_count.h5ad from the Google Drive link in the Getting Started section of the README to <data-dir>/kang_count.h5ad.

UCI Adult Income

Data was downloaded from the UCI Machine Learning Repository. Download all adult.{data,names,test} files from the data directory to the <data-dir> directory.

Step 2: Process data

Next we process the data prior to training the model. Specify a separate output folder <processed-data-x> for each of the datasets and run the following:

# Tumour / Cell line data
python3 data/process_celligner_data.py --input-dir <data-dir> --output-dir <processed-data-tcl> --top-var-number 8000
# Kang et al. PBMC scRNA-seq under INFb stimulation
python3 data/process_kang.py --input-dir <data-dir> --output-dir <processed-data-kang> --top-var-number 2000
# UCI Income data
python3 data/process_uci_income.py --input-dir <data-dir> --output-dir <processed-data-uci>

Training

Evaluation metrics are computed at the end of the training loop for the best checkpoint (by validation loss). In all the commands below the <processed-data> below should be the path the directory containing the files listed in the previous section. <output-dir> should be the desired directory in which to save the outputs. It will be created if it does not exist.

To run the code using only the CPU, omit the --use-cuda flag in the arguments to the run script.

Tumor / cell-line

python3 run.py --data-dir <processed-data-tcl> --output-dir <output-dir> --dataset tumour_cl --model comp --hidden-dim 512 --latent-dim 16 --num-layers 3 --use-batchnorm 1 --batch-size 5500 --num-epochs 4000 --learning-rate 0.0001 --penalty-scale 0.5 --kl-beta 1e-07 --seed 80244971 --use-cuda

Single cell PBMCs

python3 run.py --data-dir <processed-data-kang> --output-dir <output-dir> --dataset kang --model comp --hidden-dim 512 --latent-dim 40 --num-layers 3 --use-batchnorm 1 --batch-size 512 --num-epochs 10000 --learning-rate 1e-06 --penalty-scale 1.0 --kl-beta 1e-07 --seed 196117 --use-cuda

UCI Adult Income

python3 run.py --data-dir <processed-data-uci> --output-dir <output-dir> --dataset uci-income --model comp --hidden-dim 64 --latent-dim 16 --num-layers 2 --use-batchnorm 1 --batch-size 4096 --num-epochs 10000 --learning-rate 0.0001 --kl-beta 1.0 --penalty-scale 0.5 --seed 116983357 --use-cuda

Evaluation

An example notebook to evaluate the summary metrics is given in evaluation.ipynb.

Results

Our model achieves the following performance on :

Tumour / Cell Line Dataset

Modelsilhouettekbetmean-silhouettemean-kbet
VAE0.6580.9740.8030.581
CVAE0.5540.9310.6840.571
VFAE0.1680.2580.1980.188
trVAE0.0960.1630.1380.123
Celligner0.0820.5250.5680.226
CoMP (ours)0.0230.1600.0940.101

UCI Adult Income

ModelGender Acc.Income Acc.silhouettekbet
Original data0.7960.8490.0670.786
VAE0.7640.8120.0540.748
CVAE0.7780.8190.0540.724
VFAE-sampled0.6800.815--
VFAE-mean0.7890.8050.0460.571
trVAE0.6980.8080.0660.731
CoMP (ours)0.6790.8050.0110.451

Reference

If you find this code useful, do cite the following paper in your publication:

@inproceedings{foster2022contrastive,
title={Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and Fairness},
author={Foster, Adam and Vez{\'e}r, {\'A}rpi and Glastonbury, Craig A and Creed, P{\'a}id{\'\i} and Abujudeh, Samer and Sim, Aaron},
booktitle={International Conference on Machine Learning},
pages={6578--6621},
year={2022},
organization={PMLR}
}

About

CoMP: Contrastive Mixture of Posteriors

Resources

Stars

10 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Contrastive Mixture of Posteriors (CoMP) for Counterfactual Inference, Data Integration and Fairness

This is the PyTorch implementation of Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and Fairness ICML 2022.

CoMP Illustration

Requirements

System requirements

The following environment was used for the experiments, based on the nvidia/cuda:11.0.3-cudnn8-devel-ubuntu20.04 docker image from dockerhub:

  • Ubuntu 20.04
  • CUDA 11.0.3, cudnn 8
  • Python 3.8

Python environment

Running in a virtualenv or container is recommended. To install requirements:

pip install -U pip wheel
pip install -r requirements.txt

Datasets

The model expects two files -- a data file named features.tsv and a file with labels named metadata.tsv, with sample indices as the first column in both files, and with headers as the first row. We set the conditional variable $c$ to have column name "type" in the metadata file.

The data is prepared in two steps.

Step 1: Download source datasets

First the source datasets need to be downloaded from the following sources and placed into a folder <data-dir>.

Tumour / Cell Line Dataset

There are four files to download and prepare, as follows:

  1. Metadata file: The Celligner_info.csv file is taken from here. Save the file as <data-dir>/Celligner_info.csv.
  2. HGNC gene names: The file hgnc_complete_set_7.24.2018.txt can be downloaded from here. Save the file as <data-dir>/hgnc_complete_set_7.24.2018.txt.
  3. Tumour gene expression data: The TPM expression values can be downloaded from here. Rename this file to <data-dir>/TCGA_mat.tsv.
  4. Cell Line gene expression data: The data is downloaded from DepMap Public 19Q4 file: CCLE_expression_full.csv here. Rename this file to <data-dir>/CCLE_mat.csv.

Single cell PBMCs

Data was downloaded from the theislab/trVAE_reproducibility repository. Download kang_count.h5ad from the Google Drive link in the Getting Started section of the README to <data-dir>/kang_count.h5ad.

UCI Adult Income

Data was downloaded from the UCI Machine Learning Repository. Download all adult.{data,names,test} files from the data directory to the <data-dir> directory.

Step 2: Process data

Next we process the data prior to training the model. Specify a separate output folder <processed-data-x> for each of the datasets and run the following:

# Tumour / Cell line data
python3 data/process_celligner_data.py --input-dir <data-dir> --output-dir <processed-data-tcl> --top-var-number 8000
# Kang et al. PBMC scRNA-seq under INFb stimulation
python3 data/process_kang.py --input-dir <data-dir> --output-dir <processed-data-kang> --top-var-number 2000
# UCI Income data
python3 data/process_uci_income.py --input-dir <data-dir> --output-dir <processed-data-uci>

Training

Evaluation metrics are computed at the end of the training loop for the best checkpoint (by validation loss). In all the commands below the <processed-data> below should be the path the directory containing the files listed in the previous section. <output-dir> should be the desired directory in which to save the outputs. It will be created if it does not exist.

To run the code using only the CPU, omit the --use-cuda flag in the arguments to the run script.

Tumor / cell-line

python3 run.py --data-dir <processed-data-tcl> --output-dir <output-dir> --dataset tumour_cl --model comp --hidden-dim 512 --latent-dim 16 --num-layers 3 --use-batchnorm 1 --batch-size 5500 --num-epochs 4000 --learning-rate 0.0001 --penalty-scale 0.5 --kl-beta 1e-07 --seed 80244971 --use-cuda

Single cell PBMCs

python3 run.py --data-dir <processed-data-kang> --output-dir <output-dir> --dataset kang --model comp --hidden-dim 512 --latent-dim 40 --num-layers 3 --use-batchnorm 1 --batch-size 512 --num-epochs 10000 --learning-rate 1e-06 --penalty-scale 1.0 --kl-beta 1e-07 --seed 196117 --use-cuda

UCI Adult Income

python3 run.py --data-dir <processed-data-uci> --output-dir <output-dir> --dataset uci-income --model comp --hidden-dim 64 --latent-dim 16 --num-layers 2 --use-batchnorm 1 --batch-size 4096 --num-epochs 10000 --learning-rate 0.0001 --kl-beta 1.0 --penalty-scale 0.5 --seed 116983357 --use-cuda

Evaluation

An example notebook to evaluate the summary metrics is given in evaluation.ipynb.

Results

Our model achieves the following performance on :

Tumour / Cell Line Dataset

Modelsilhouettekbetmean-silhouettemean-kbet
VAE0.6580.9740.8030.581
CVAE0.5540.9310.6840.571
VFAE0.1680.2580.1980.188
trVAE0.0960.1630.1380.123
Celligner0.0820.5250.5680.226
CoMP (ours)0.0230.1600.0940.101

UCI Adult Income

ModelGender Acc.Income Acc.silhouettekbet
Original data0.7960.8490.0670.786
VAE0.7640.8120.0540.748
CVAE0.7780.8190.0540.724
VFAE-sampled0.6800.815--
VFAE-mean0.7890.8050.0460.571
trVAE0.6980.8080.0660.731
CoMP (ours)0.6790.8050.0110.451

Reference

If you find this code useful, do cite the following paper in your publication:

@inproceedings{foster2022contrastive,
title={Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and Fairness},
author={Foster, Adam and Vez{\'e}r, {\'A}rpi and Glastonbury, Craig A and Creed, P{\'a}id{\'\i} and Abujudeh, Samer and Sim, Aaron},
booktitle={International Conference on Machine Learning},
pages={6578--6621},
year={2022},
organization={PMLR}
}

About

CoMP: Contrastive Mixture of Posteriors

Resources

Stars

10 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Contrastive Mixture of Posteriors (CoMP) for Counterfactual Inference, Data Integration and Fairness

This is the PyTorch implementation of Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and Fairness ICML 2022.

CoMP Illustration

Requirements

System requirements

The following environment was used for the experiments, based on the nvidia/cuda:11.0.3-cudnn8-devel-ubuntu20.04 docker image from dockerhub:

  • Ubuntu 20.04
  • CUDA 11.0.3, cudnn 8
  • Python 3.8

Python environment

Running in a virtualenv or container is recommended. To install requirements:

pip install -U pip wheel
pip install -r requirements.txt

Datasets

The model expects two files -- a data file named features.tsv and a file with labels named metadata.tsv, with sample indices as the first column in both files, and with headers as the first row. We set the conditional variable $c$ to have column name "type" in the metadata file.

The data is prepared in two steps.

Step 1: Download source datasets

First the source datasets need to be downloaded from the following sources and placed into a folder <data-dir>.

Tumour / Cell Line Dataset

There are four files to download and prepare, as follows:

  1. Metadata file: The Celligner_info.csv file is taken from here. Save the file as <data-dir>/Celligner_info.csv.
  2. HGNC gene names: The file hgnc_complete_set_7.24.2018.txt can be downloaded from here. Save the file as <data-dir>/hgnc_complete_set_7.24.2018.txt.
  3. Tumour gene expression data: The TPM expression values can be downloaded from here. Rename this file to <data-dir>/TCGA_mat.tsv.
  4. Cell Line gene expression data: The data is downloaded from DepMap Public 19Q4 file: CCLE_expression_full.csv here. Rename this file to <data-dir>/CCLE_mat.csv.

Single cell PBMCs

Data was downloaded from the theislab/trVAE_reproducibility repository. Download kang_count.h5ad from the Google Drive link in the Getting Started section of the README to <data-dir>/kang_count.h5ad.

UCI Adult Income

Data was downloaded from the UCI Machine Learning Repository. Download all adult.{data,names,test} files from the data directory to the <data-dir> directory.

Step 2: Process data

Next we process the data prior to training the model. Specify a separate output folder <processed-data-x> for each of the datasets and run the following:

# Tumour / Cell line data
python3 data/process_celligner_data.py --input-dir <data-dir> --output-dir <processed-data-tcl> --top-var-number 8000
# Kang et al. PBMC scRNA-seq under INFb stimulation
python3 data/process_kang.py --input-dir <data-dir> --output-dir <processed-data-kang> --top-var-number 2000
# UCI Income data
python3 data/process_uci_income.py --input-dir <data-dir> --output-dir <processed-data-uci>

Training

Evaluation metrics are computed at the end of the training loop for the best checkpoint (by validation loss). In all the commands below the <processed-data> below should be the path the directory containing the files listed in the previous section. <output-dir> should be the desired directory in which to save the outputs. It will be created if it does not exist.

To run the code using only the CPU, omit the --use-cuda flag in the arguments to the run script.

Tumor / cell-line

python3 run.py --data-dir <processed-data-tcl> --output-dir <output-dir> --dataset tumour_cl --model comp --hidden-dim 512 --latent-dim 16 --num-layers 3 --use-batchnorm 1 --batch-size 5500 --num-epochs 4000 --learning-rate 0.0001 --penalty-scale 0.5 --kl-beta 1e-07 --seed 80244971 --use-cuda

Single cell PBMCs

python3 run.py --data-dir <processed-data-kang> --output-dir <output-dir> --dataset kang --model comp --hidden-dim 512 --latent-dim 40 --num-layers 3 --use-batchnorm 1 --batch-size 512 --num-epochs 10000 --learning-rate 1e-06 --penalty-scale 1.0 --kl-beta 1e-07 --seed 196117 --use-cuda

UCI Adult Income

python3 run.py --data-dir <processed-data-uci> --output-dir <output-dir> --dataset uci-income --model comp --hidden-dim 64 --latent-dim 16 --num-layers 2 --use-batchnorm 1 --batch-size 4096 --num-epochs 10000 --learning-rate 0.0001 --kl-beta 1.0 --penalty-scale 0.5 --seed 116983357 --use-cuda

Evaluation

An example notebook to evaluate the summary metrics is given in evaluation.ipynb.

Results

Our model achieves the following performance on :

Tumour / Cell Line Dataset

Modelsilhouettekbetmean-silhouettemean-kbet
VAE0.6580.9740.8030.581
CVAE0.5540.9310.6840.571
VFAE0.1680.2580.1980.188
trVAE0.0960.1630.1380.123
Celligner0.0820.5250.5680.226
CoMP (ours)0.0230.1600.0940.101

UCI Adult Income

ModelGender Acc.Income Acc.silhouettekbet
Original data0.7960.8490.0670.786
VAE0.7640.8120.0540.748
CVAE0.7780.8190.0540.724
VFAE-sampled0.6800.815--
VFAE-mean0.7890.8050.0460.571
trVAE0.6980.8080.0660.731
CoMP (ours)0.6790.8050.0110.451

Reference

If you find this code useful, do cite the following paper in your publication:

@inproceedings{foster2022contrastive,
title={Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and Fairness},
author={Foster, Adam and Vez{\'e}r, {\'A}rpi and Glastonbury, Craig A and Creed, P{\'a}id{\'\i} and Abujudeh, Samer and Sim, Aaron},
booktitle={International Conference on Machine Learning},
pages={6578--6621},
year={2022},
organization={PMLR}
}

About

CoMP: Contrastive Mixture of Posteriors

Resources

Stars

10 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Contrastive Mixture of Posteriors (CoMP) for Counterfactual Inference, Data Integration and Fairness

This is the PyTorch implementation of Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and Fairness ICML 2022.

CoMP Illustration

Requirements

System requirements

The following environment was used for the experiments, based on the nvidia/cuda:11.0.3-cudnn8-devel-ubuntu20.04 docker image from dockerhub:

  • Ubuntu 20.04
  • CUDA 11.0.3, cudnn 8
  • Python 3.8

Python environment

Running in a virtualenv or container is recommended. To install requirements:

pip install -U pip wheel
pip install -r requirements.txt

Datasets

The model expects two files -- a data file named features.tsv and a file with labels named metadata.tsv, with sample indices as the first column in both files, and with headers as the first row. We set the conditional variable $c$ to have column name "type" in the metadata file.

The data is prepared in two steps.

Step 1: Download source datasets

First the source datasets need to be downloaded from the following sources and placed into a folder <data-dir>.

Tumour / Cell Line Dataset

There are four files to download and prepare, as follows:

  1. Metadata file: The Celligner_info.csv file is taken from here. Save the file as <data-dir>/Celligner_info.csv.
  2. HGNC gene names: The file hgnc_complete_set_7.24.2018.txt can be downloaded from here. Save the file as <data-dir>/hgnc_complete_set_7.24.2018.txt.
  3. Tumour gene expression data: The TPM expression values can be downloaded from here. Rename this file to <data-dir>/TCGA_mat.tsv.
  4. Cell Line gene expression data: The data is downloaded from DepMap Public 19Q4 file: CCLE_expression_full.csv here. Rename this file to <data-dir>/CCLE_mat.csv.

Single cell PBMCs

Data was downloaded from the theislab/trVAE_reproducibility repository. Download kang_count.h5ad from the Google Drive link in the Getting Started section of the README to <data-dir>/kang_count.h5ad.

UCI Adult Income

Data was downloaded from the UCI Machine Learning Repository. Download all adult.{data,names,test} files from the data directory to the <data-dir> directory.

Step 2: Process data

Next we process the data prior to training the model. Specify a separate output folder <processed-data-x> for each of the datasets and run the following:

# Tumour / Cell line data
python3 data/process_celligner_data.py --input-dir <data-dir> --output-dir <processed-data-tcl> --top-var-number 8000
# Kang et al. PBMC scRNA-seq under INFb stimulation
python3 data/process_kang.py --input-dir <data-dir> --output-dir <processed-data-kang> --top-var-number 2000
# UCI Income data
python3 data/process_uci_income.py --input-dir <data-dir> --output-dir <processed-data-uci>

Training

Evaluation metrics are computed at the end of the training loop for the best checkpoint (by validation loss). In all the commands below the <processed-data> below should be the path the directory containing the files listed in the previous section. <output-dir> should be the desired directory in which to save the outputs. It will be created if it does not exist.

To run the code using only the CPU, omit the --use-cuda flag in the arguments to the run script.

Tumor / cell-line

python3 run.py --data-dir <processed-data-tcl> --output-dir <output-dir> --dataset tumour_cl --model comp --hidden-dim 512 --latent-dim 16 --num-layers 3 --use-batchnorm 1 --batch-size 5500 --num-epochs 4000 --learning-rate 0.0001 --penalty-scale 0.5 --kl-beta 1e-07 --seed 80244971 --use-cuda

Single cell PBMCs

python3 run.py --data-dir <processed-data-kang> --output-dir <output-dir> --dataset kang --model comp --hidden-dim 512 --latent-dim 40 --num-layers 3 --use-batchnorm 1 --batch-size 512 --num-epochs 10000 --learning-rate 1e-06 --penalty-scale 1.0 --kl-beta 1e-07 --seed 196117 --use-cuda

UCI Adult Income

python3 run.py --data-dir <processed-data-uci> --output-dir <output-dir> --dataset uci-income --model comp --hidden-dim 64 --latent-dim 16 --num-layers 2 --use-batchnorm 1 --batch-size 4096 --num-epochs 10000 --learning-rate 0.0001 --kl-beta 1.0 --penalty-scale 0.5 --seed 116983357 --use-cuda

Evaluation

An example notebook to evaluate the summary metrics is given in evaluation.ipynb.

Results

Our model achieves the following performance on :

Tumour / Cell Line Dataset

Modelsilhouettekbetmean-silhouettemean-kbet
VAE0.6580.9740.8030.581
CVAE0.5540.9310.6840.571
VFAE0.1680.2580.1980.188
trVAE0.0960.1630.1380.123
Celligner0.0820.5250.5680.226
CoMP (ours)0.0230.1600.0940.101

UCI Adult Income

ModelGender Acc.Income Acc.silhouettekbet
Original data0.7960.8490.0670.786
VAE0.7640.8120.0540.748
CVAE0.7780.8190.0540.724
VFAE-sampled0.6800.815--
VFAE-mean0.7890.8050.0460.571
trVAE0.6980.8080.0660.731
CoMP (ours)0.6790.8050.0110.451

Reference

If you find this code useful, do cite the following paper in your publication:

@inproceedings{foster2022contrastive,
title={Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and Fairness},
author={Foster, Adam and Vez{\'e}r, {\'A}rpi and Glastonbury, Craig A and Creed, P{\'a}id{\'\i} and Abujudeh, Samer and Sim, Aaron},
booktitle={International Conference on Machine Learning},
pages={6578--6621},
year={2022},
organization={PMLR}
}

About

CoMP: Contrastive Mixture of Posteriors

Resources

Stars

10 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Contrastive Mixture of Posteriors (CoMP) for Counterfactual Inference, Data Integration and Fairness

This is the PyTorch implementation of Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and Fairness ICML 2022.

CoMP Illustration

Requirements

System requirements

The following environment was used for the experiments, based on the nvidia/cuda:11.0.3-cudnn8-devel-ubuntu20.04 docker image from dockerhub:

  • Ubuntu 20.04
  • CUDA 11.0.3, cudnn 8
  • Python 3.8

Python environment

Running in a virtualenv or container is recommended. To install requirements:

pip install -U pip wheel
pip install -r requirements.txt

Datasets

The model expects two files -- a data file named features.tsv and a file with labels named metadata.tsv, with sample indices as the first column in both files, and with headers as the first row. We set the conditional variable $c$ to have column name "type" in the metadata file.

The data is prepared in two steps.

Step 1: Download source datasets

First the source datasets need to be downloaded from the following sources and placed into a folder <data-dir>.

Tumour / Cell Line Dataset

There are four files to download and prepare, as follows:

  1. Metadata file: The Celligner_info.csv file is taken from here. Save the file as <data-dir>/Celligner_info.csv.
  2. HGNC gene names: The file hgnc_complete_set_7.24.2018.txt can be downloaded from here. Save the file as <data-dir>/hgnc_complete_set_7.24.2018.txt.
  3. Tumour gene expression data: The TPM expression values can be downloaded from here. Rename this file to <data-dir>/TCGA_mat.tsv.
  4. Cell Line gene expression data: The data is downloaded from DepMap Public 19Q4 file: CCLE_expression_full.csv here. Rename this file to <data-dir>/CCLE_mat.csv.

Single cell PBMCs

Data was downloaded from the theislab/trVAE_reproducibility repository. Download kang_count.h5ad from the Google Drive link in the Getting Started section of the README to <data-dir>/kang_count.h5ad.

UCI Adult Income

Data was downloaded from the UCI Machine Learning Repository. Download all adult.{data,names,test} files from the data directory to the <data-dir> directory.

Step 2: Process data

Next we process the data prior to training the model. Specify a separate output folder <processed-data-x> for each of the datasets and run the following:

# Tumour / Cell line data
python3 data/process_celligner_data.py --input-dir <data-dir> --output-dir <processed-data-tcl> --top-var-number 8000
# Kang et al. PBMC scRNA-seq under INFb stimulation
python3 data/process_kang.py --input-dir <data-dir> --output-dir <processed-data-kang> --top-var-number 2000
# UCI Income data
python3 data/process_uci_income.py --input-dir <data-dir> --output-dir <processed-data-uci>

Training

Evaluation metrics are computed at the end of the training loop for the best checkpoint (by validation loss). In all the commands below the <processed-data> below should be the path the directory containing the files listed in the previous section. <output-dir> should be the desired directory in which to save the outputs. It will be created if it does not exist.

To run the code using only the CPU, omit the --use-cuda flag in the arguments to the run script.

Tumor / cell-line

python3 run.py --data-dir <processed-data-tcl> --output-dir <output-dir> --dataset tumour_cl --model comp --hidden-dim 512 --latent-dim 16 --num-layers 3 --use-batchnorm 1 --batch-size 5500 --num-epochs 4000 --learning-rate 0.0001 --penalty-scale 0.5 --kl-beta 1e-07 --seed 80244971 --use-cuda

Single cell PBMCs

python3 run.py --data-dir <processed-data-kang> --output-dir <output-dir> --dataset kang --model comp --hidden-dim 512 --latent-dim 40 --num-layers 3 --use-batchnorm 1 --batch-size 512 --num-epochs 10000 --learning-rate 1e-06 --penalty-scale 1.0 --kl-beta 1e-07 --seed 196117 --use-cuda

UCI Adult Income

python3 run.py --data-dir <processed-data-uci> --output-dir <output-dir> --dataset uci-income --model comp --hidden-dim 64 --latent-dim 16 --num-layers 2 --use-batchnorm 1 --batch-size 4096 --num-epochs 10000 --learning-rate 0.0001 --kl-beta 1.0 --penalty-scale 0.5 --seed 116983357 --use-cuda

Evaluation

An example notebook to evaluate the summary metrics is given in evaluation.ipynb.

Results

Our model achieves the following performance on :

Tumour / Cell Line Dataset

Modelsilhouettekbetmean-silhouettemean-kbet
VAE0.6580.9740.8030.581
CVAE0.5540.9310.6840.571
VFAE0.1680.2580.1980.188
trVAE0.0960.1630.1380.123
Celligner0.0820.5250.5680.226
CoMP (ours)0.0230.1600.0940.101

UCI Adult Income

ModelGender Acc.Income Acc.silhouettekbet
Original data0.7960.8490.0670.786
VAE0.7640.8120.0540.748
CVAE0.7780.8190.0540.724
VFAE-sampled0.6800.815--
VFAE-mean0.7890.8050.0460.571
trVAE0.6980.8080.0660.731
CoMP (ours)0.6790.8050.0110.451

Reference

If you find this code useful, do cite the following paper in your publication:

@inproceedings{foster2022contrastive,
title={Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and Fairness},
author={Foster, Adam and Vez{\'e}r, {\'A}rpi and Glastonbury, Craig A and Creed, P{\'a}id{\'\i} and Abujudeh, Samer and Sim, Aaron},
booktitle={International Conference on Machine Learning},
pages={6578--6621},
year={2022},
organization={PMLR}
}

About

CoMP: Contrastive Mixture of Posteriors

Resources

Stars

10 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Contrastive Mixture of Posteriors (CoMP) for Counterfactual Inference, Data Integration and Fairness

This is the PyTorch implementation of Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and Fairness ICML 2022.

CoMP Illustration

Requirements

System requirements

The following environment was used for the experiments, based on the nvidia/cuda:11.0.3-cudnn8-devel-ubuntu20.04 docker image from dockerhub:

  • Ubuntu 20.04
  • CUDA 11.0.3, cudnn 8
  • Python 3.8

Python environment

Running in a virtualenv or container is recommended. To install requirements:

pip install -U pip wheel
pip install -r requirements.txt

Datasets

The model expects two files -- a data file named features.tsv and a file with labels named metadata.tsv, with sample indices as the first column in both files, and with headers as the first row. We set the conditional variable $c$ to have column name "type" in the metadata file.

The data is prepared in two steps.

Step 1: Download source datasets

First the source datasets need to be downloaded from the following sources and placed into a folder <data-dir>.

Tumour / Cell Line Dataset

There are four files to download and prepare, as follows:

  1. Metadata file: The Celligner_info.csv file is taken from here. Save the file as <data-dir>/Celligner_info.csv.
  2. HGNC gene names: The file hgnc_complete_set_7.24.2018.txt can be downloaded from here. Save the file as <data-dir>/hgnc_complete_set_7.24.2018.txt.
  3. Tumour gene expression data: The TPM expression values can be downloaded from here. Rename this file to <data-dir>/TCGA_mat.tsv.
  4. Cell Line gene expression data: The data is downloaded from DepMap Public 19Q4 file: CCLE_expression_full.csv here. Rename this file to <data-dir>/CCLE_mat.csv.

Single cell PBMCs

Data was downloaded from the theislab/trVAE_reproducibility repository. Download kang_count.h5ad from the Google Drive link in the Getting Started section of the README to <data-dir>/kang_count.h5ad.

UCI Adult Income

Data was downloaded from the UCI Machine Learning Repository. Download all adult.{data,names,test} files from the data directory to the <data-dir> directory.

Step 2: Process data

Next we process the data prior to training the model. Specify a separate output folder <processed-data-x> for each of the datasets and run the following:

# Tumour / Cell line data
python3 data/process_celligner_data.py --input-dir <data-dir> --output-dir <processed-data-tcl> --top-var-number 8000
# Kang et al. PBMC scRNA-seq under INFb stimulation
python3 data/process_kang.py --input-dir <data-dir> --output-dir <processed-data-kang> --top-var-number 2000
# UCI Income data
python3 data/process_uci_income.py --input-dir <data-dir> --output-dir <processed-data-uci>

Training

Evaluation metrics are computed at the end of the training loop for the best checkpoint (by validation loss). In all the commands below the <processed-data> below should be the path the directory containing the files listed in the previous section. <output-dir> should be the desired directory in which to save the outputs. It will be created if it does not exist.

To run the code using only the CPU, omit the --use-cuda flag in the arguments to the run script.

Tumor / cell-line

python3 run.py --data-dir <processed-data-tcl> --output-dir <output-dir> --dataset tumour_cl --model comp --hidden-dim 512 --latent-dim 16 --num-layers 3 --use-batchnorm 1 --batch-size 5500 --num-epochs 4000 --learning-rate 0.0001 --penalty-scale 0.5 --kl-beta 1e-07 --seed 80244971 --use-cuda

Single cell PBMCs

python3 run.py --data-dir <processed-data-kang> --output-dir <output-dir> --dataset kang --model comp --hidden-dim 512 --latent-dim 40 --num-layers 3 --use-batchnorm 1 --batch-size 512 --num-epochs 10000 --learning-rate 1e-06 --penalty-scale 1.0 --kl-beta 1e-07 --seed 196117 --use-cuda

UCI Adult Income

python3 run.py --data-dir <processed-data-uci> --output-dir <output-dir> --dataset uci-income --model comp --hidden-dim 64 --latent-dim 16 --num-layers 2 --use-batchnorm 1 --batch-size 4096 --num-epochs 10000 --learning-rate 0.0001 --kl-beta 1.0 --penalty-scale 0.5 --seed 116983357 --use-cuda

Evaluation

An example notebook to evaluate the summary metrics is given in evaluation.ipynb.

Results

Our model achieves the following performance on :

Tumour / Cell Line Dataset

Modelsilhouettekbetmean-silhouettemean-kbet
VAE0.6580.9740.8030.581
CVAE0.5540.9310.6840.571
VFAE0.1680.2580.1980.188
trVAE0.0960.1630.1380.123
Celligner0.0820.5250.5680.226
CoMP (ours)0.0230.1600.0940.101

UCI Adult Income

ModelGender Acc.Income Acc.silhouettekbet
Original data0.7960.8490.0670.786
VAE0.7640.8120.0540.748
CVAE0.7780.8190.0540.724
VFAE-sampled0.6800.815--
VFAE-mean0.7890.8050.0460.571
trVAE0.6980.8080.0660.731
CoMP (ours)0.6790.8050.0110.451

Reference

If you find this code useful, do cite the following paper in your publication:

@inproceedings{foster2022contrastive,
title={Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and Fairness},
author={Foster, Adam and Vez{\'e}r, {\'A}rpi and Glastonbury, Craig A and Creed, P{\'a}id{\'\i} and Abujudeh, Samer and Sim, Aaron},
booktitle={International Conference on Machine Learning},
pages={6578--6621},
year={2022},
organization={PMLR}
}

About

CoMP: Contrastive Mixture of Posteriors

Resources

Stars

10 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages