Latest commit

History

63 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

contrastive

A python library for performing unsupervised machine learning on datasets with learning (e.g. PCA) in contrastive settings, where one is interested in patterns (e.g. clusters or clines) that exist one dataset, but not the other.

Applications include dicovering subgroups in biological and medical data. Here are basic installation and usage instructions, written for Python 3 (in which the library has been developed and tested, although it should work in Python 2 as well).

For more details, see the accompanying paper: "Exploring Patterns Enriched in a Dataset with Contrastive Principal Component Analysis", Nature Communications (2018), and please use the citation below.

@article{abid2018exploring,
title={Exploring patterns enriched in a dataset with contrastive principal component analysis},
author={Abid, Abubakar and Zhang, Martin J and Bagaria, Vivek K and Zou, James},
journal={Nature communications},
volume={9},
number={1},
pages={2134},
year={2018},
}

This repository also includes experiments to reproduce most of the figures in the paper. Please see the python notebooks in the experiments folder.

Installation

$ pip3 install contrastive

Basic Usage

The basic functions enabled by this library are shown below. Generally speaking, we have two datasets, one is a dataset that we can label as foreground_data, which is the dataset in which we are discovering patterns and directions, and another dataset called background_data, which is the dataset that does not have the patterns or directions we are interested in discovering. In some cases, both datasets may contain the signal of interest, but the foreground dataset may have the pattern enriched relative to the background. In these analyses, there is a contrast parameter, known as alpha, which can be thought of as a hyperparameter.

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data)
#returns a set of 2-dimensional projections of the foreground data stored in the list 'projected_data', for several different values of 'alpha' that are automatically chosen (by default, 4 values of alpha are chosen)

Note that foreground_data and background_data should be 2D numpy arrays that have the second dimension (which represents the number of features). In other words, foreground_data.shape[1]==background_data.shape[1] should return True.

Built-in plotting: to quickly see the results of contrastive PCA, simply enable the plot parameter to true:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, plot=True)

images/plot_true.png

Interactive GUI: if you are running these analyses inside a jupyter notebook, you can easily launch an interactive GUI as shown here:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, gui=True)

images/gui_true.png

Using the slider, you can see how the your data points move as you change the value of the contrast parameter. These animations can reveal groups in the data and other insights:

images/animation.gif

Quick Test

To ensure that the library is working, here is a quick script that will allow you to test the code on synthetic data. Simply run the following commands:

importnumpyasnpfromcontrastiveimportCPCAN=400; D=30; gap=3# In B, all the data pts are from the same distribution, which has different variances in three subspaces.B=np.zeros((N, D))
B[:,0:10] =np.random.normal(0,10,(N,10))
B[:,10:20] =np.random.normal(0,3,(N,10))
B[:,20:30] =np.random.normal(0,1,(N,10))
# In A there are four clusters.A=np.zeros((N, D))
A[:,0:10] =np.random.normal(0,10,(N,10))
# group 1A[0:100, 10:20] =np.random.normal(0,1,(100,10))
A[0:100, 20:30] =np.random.normal(0,1,(100,10))
# group 2A[100:200, 10:20] =np.random.normal(0,1,(100,10))
A[100:200, 20:30] =np.random.normal(gap,1,(100,10))
# group 3A[200:300, 10:20] =np.random.normal(2*gap,1,(100,10))
A[200:300, 20:30] =np.random.normal(0,1,(100,10))
# group 4A[300:400, 10:20] =np.random.normal(2*gap,1,(100,10))
A[300:400, 20:30] =np.random.normal(gap,1,(100,10))
A_labels= [0]*100+[1]*100+[2]*100+[3]*100cpca=CPCA(standardize=False)
cpca.fit_transform(A, B, plot=True, active_labels=A_labels)

You should see a series of plots that looks something like this:

images/plot_example.png

Optional Parameters

Labels for foreground data (plot/gui mode): In the examples above, the data points are colored according to labels known ahead of time. You can supply these labels using the active_labels parameter, as shown here:

fromcontrastiveimportCPCAmdl=CPCA()
#labels = [0, 1, 0, 1, 1 ... 1, 0]projected_data=mdl.fit_transform(foreground_data, background_data, plot=True, active_labels=labels)

Additional # of components: Sometimes, you'd like to project your data on more than the top 2 contrastive principal components (cPCs). Specify the number of cPCs when you instantiate your model using the n_components parameter:

fromcontrastiveimportCPCAmdl=CPCA(n_components=3) #the top 3 components will be returnedprojected_data=mdl.fit_transform(foreground_data, background_data)

However, note that only when n_components=2 can the data be plotted or visualized through the GUI.

How values of alpha are chosen: So far, we've always plotted the data when the values of alpha have been chosen automatically with default parameters. However, the values of alpha can be customized. For example, if you'd like to still choose the values of alpha automatically, but change the range or number of alphas considered, you can use the n_alphas and max_log_alpha parameters. The former sets the number of alphas that are analyzed, and the latter sets the upper bound on the highest value of log (base 10) alpha. (The minimum value of alpha, besides alpha = 0, is always alpha = 0.1). Finally, you can change the number of values of alpha that are returned using the n_alphas_to_return parameter.

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, n_alphas=10, max_log_alpha=2, n_alphas_to_return=1) #search through 10 logarithmically spaced values of alpha from 0.1 to 100 and return the PCs for only 1 of them.

You can also decide to set the value of alpha to a particular value of alpha manually by changing the alpha_selection and alpha_value parameters as follows:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, alpha_selection='manual', alpha_value=2.0)

Or you can decide to plot or return the data for _all_ values of alpha in the given range. In this case, you can still choose to set the n_alphas and max_log_alpha parameters:

fromcontrastiveimportCPCAmdl=CPCA() #the top 3 components will be returnedprojected_data=mdl.fit_transform(foreground_data, background_data, n_alphas=10, max_log_alpha=2, alpha_selection='all') #search through 10 logarithmically spaced values of alpha from 0.1 to 100 and return the PCs for all of them!

Whether to standardize your data: By default, before performing contrastive PCA, the data are standardized so that each column or dimension has unit variance. You can turn this off by doing the following:

fromcontrastiveimportCPCAmdl=CPCA(standardize=False)
projected_data=mdl.fit_transform(foreground_data, background_data)

Custom colors (plot/gui mode): As a stylistic touch, you can also customize which colors are used to label the points when the data is plotted by using the colors argument. Here's an example:

fromcontrastiveimportCPCAmdl=CPCA(standardize=False)
projected_data=mdl.fit_transform(foreground_data, background_data, gui=True, colors=['r','b','k','c'])

will produce something along the lines of:

images/gui_colors.png

About

a faster Contrastive PCA with ability to only output correction vectors.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

63 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

contrastive

A python library for performing unsupervised machine learning on datasets with learning (e.g. PCA) in contrastive settings, where one is interested in patterns (e.g. clusters or clines) that exist one dataset, but not the other.

Applications include dicovering subgroups in biological and medical data. Here are basic installation and usage instructions, written for Python 3 (in which the library has been developed and tested, although it should work in Python 2 as well).

For more details, see the accompanying paper: "Exploring Patterns Enriched in a Dataset with Contrastive Principal Component Analysis", Nature Communications (2018), and please use the citation below.

@article{abid2018exploring,
title={Exploring patterns enriched in a dataset with contrastive principal component analysis},
author={Abid, Abubakar and Zhang, Martin J and Bagaria, Vivek K and Zou, James},
journal={Nature communications},
volume={9},
number={1},
pages={2134},
year={2018},
}

This repository also includes experiments to reproduce most of the figures in the paper. Please see the python notebooks in the experiments folder.

Installation

$ pip3 install contrastive

Basic Usage

The basic functions enabled by this library are shown below. Generally speaking, we have two datasets, one is a dataset that we can label as foreground_data, which is the dataset in which we are discovering patterns and directions, and another dataset called background_data, which is the dataset that does not have the patterns or directions we are interested in discovering. In some cases, both datasets may contain the signal of interest, but the foreground dataset may have the pattern enriched relative to the background. In these analyses, there is a contrast parameter, known as alpha, which can be thought of as a hyperparameter.

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data)
#returns a set of 2-dimensional projections of the foreground data stored in the list 'projected_data', for several different values of 'alpha' that are automatically chosen (by default, 4 values of alpha are chosen)

Note that foreground_data and background_data should be 2D numpy arrays that have the second dimension (which represents the number of features). In other words, foreground_data.shape[1]==background_data.shape[1] should return True.

Built-in plotting: to quickly see the results of contrastive PCA, simply enable the plot parameter to true:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, plot=True)

images/plot_true.png

Interactive GUI: if you are running these analyses inside a jupyter notebook, you can easily launch an interactive GUI as shown here:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, gui=True)

images/gui_true.png

Using the slider, you can see how the your data points move as you change the value of the contrast parameter. These animations can reveal groups in the data and other insights:

images/animation.gif

Quick Test

To ensure that the library is working, here is a quick script that will allow you to test the code on synthetic data. Simply run the following commands:

importnumpyasnpfromcontrastiveimportCPCAN=400; D=30; gap=3# In B, all the data pts are from the same distribution, which has different variances in three subspaces.B=np.zeros((N, D))
B[:,0:10] =np.random.normal(0,10,(N,10))
B[:,10:20] =np.random.normal(0,3,(N,10))
B[:,20:30] =np.random.normal(0,1,(N,10))
# In A there are four clusters.A=np.zeros((N, D))
A[:,0:10] =np.random.normal(0,10,(N,10))
# group 1A[0:100, 10:20] =np.random.normal(0,1,(100,10))
A[0:100, 20:30] =np.random.normal(0,1,(100,10))
# group 2A[100:200, 10:20] =np.random.normal(0,1,(100,10))
A[100:200, 20:30] =np.random.normal(gap,1,(100,10))
# group 3A[200:300, 10:20] =np.random.normal(2*gap,1,(100,10))
A[200:300, 20:30] =np.random.normal(0,1,(100,10))
# group 4A[300:400, 10:20] =np.random.normal(2*gap,1,(100,10))
A[300:400, 20:30] =np.random.normal(gap,1,(100,10))
A_labels= [0]*100+[1]*100+[2]*100+[3]*100cpca=CPCA(standardize=False)
cpca.fit_transform(A, B, plot=True, active_labels=A_labels)

You should see a series of plots that looks something like this:

images/plot_example.png

Optional Parameters

Labels for foreground data (plot/gui mode): In the examples above, the data points are colored according to labels known ahead of time. You can supply these labels using the active_labels parameter, as shown here:

fromcontrastiveimportCPCAmdl=CPCA()
#labels = [0, 1, 0, 1, 1 ... 1, 0]projected_data=mdl.fit_transform(foreground_data, background_data, plot=True, active_labels=labels)

Additional # of components: Sometimes, you'd like to project your data on more than the top 2 contrastive principal components (cPCs). Specify the number of cPCs when you instantiate your model using the n_components parameter:

fromcontrastiveimportCPCAmdl=CPCA(n_components=3) #the top 3 components will be returnedprojected_data=mdl.fit_transform(foreground_data, background_data)

However, note that only when n_components=2 can the data be plotted or visualized through the GUI.

How values of alpha are chosen: So far, we've always plotted the data when the values of alpha have been chosen automatically with default parameters. However, the values of alpha can be customized. For example, if you'd like to still choose the values of alpha automatically, but change the range or number of alphas considered, you can use the n_alphas and max_log_alpha parameters. The former sets the number of alphas that are analyzed, and the latter sets the upper bound on the highest value of log (base 10) alpha. (The minimum value of alpha, besides alpha = 0, is always alpha = 0.1). Finally, you can change the number of values of alpha that are returned using the n_alphas_to_return parameter.

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, n_alphas=10, max_log_alpha=2, n_alphas_to_return=1) #search through 10 logarithmically spaced values of alpha from 0.1 to 100 and return the PCs for only 1 of them.

You can also decide to set the value of alpha to a particular value of alpha manually by changing the alpha_selection and alpha_value parameters as follows:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, alpha_selection='manual', alpha_value=2.0)

Or you can decide to plot or return the data for _all_ values of alpha in the given range. In this case, you can still choose to set the n_alphas and max_log_alpha parameters:

fromcontrastiveimportCPCAmdl=CPCA() #the top 3 components will be returnedprojected_data=mdl.fit_transform(foreground_data, background_data, n_alphas=10, max_log_alpha=2, alpha_selection='all') #search through 10 logarithmically spaced values of alpha from 0.1 to 100 and return the PCs for all of them!

Whether to standardize your data: By default, before performing contrastive PCA, the data are standardized so that each column or dimension has unit variance. You can turn this off by doing the following:

fromcontrastiveimportCPCAmdl=CPCA(standardize=False)
projected_data=mdl.fit_transform(foreground_data, background_data)

Custom colors (plot/gui mode): As a stylistic touch, you can also customize which colors are used to label the points when the data is plotted by using the colors argument. Here's an example:

fromcontrastiveimportCPCAmdl=CPCA(standardize=False)
projected_data=mdl.fit_transform(foreground_data, background_data, gui=True, colors=['r','b','k','c'])

will produce something along the lines of:

images/gui_colors.png

About

a faster Contrastive PCA with ability to only output correction vectors.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

63 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

contrastive

A python library for performing unsupervised machine learning on datasets with learning (e.g. PCA) in contrastive settings, where one is interested in patterns (e.g. clusters or clines) that exist one dataset, but not the other.

Applications include dicovering subgroups in biological and medical data. Here are basic installation and usage instructions, written for Python 3 (in which the library has been developed and tested, although it should work in Python 2 as well).

For more details, see the accompanying paper: "Exploring Patterns Enriched in a Dataset with Contrastive Principal Component Analysis", Nature Communications (2018), and please use the citation below.

@article{abid2018exploring,
title={Exploring patterns enriched in a dataset with contrastive principal component analysis},
author={Abid, Abubakar and Zhang, Martin J and Bagaria, Vivek K and Zou, James},
journal={Nature communications},
volume={9},
number={1},
pages={2134},
year={2018},
}

This repository also includes experiments to reproduce most of the figures in the paper. Please see the python notebooks in the experiments folder.

Installation

$ pip3 install contrastive

Basic Usage

The basic functions enabled by this library are shown below. Generally speaking, we have two datasets, one is a dataset that we can label as foreground_data, which is the dataset in which we are discovering patterns and directions, and another dataset called background_data, which is the dataset that does not have the patterns or directions we are interested in discovering. In some cases, both datasets may contain the signal of interest, but the foreground dataset may have the pattern enriched relative to the background. In these analyses, there is a contrast parameter, known as alpha, which can be thought of as a hyperparameter.

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data)
#returns a set of 2-dimensional projections of the foreground data stored in the list 'projected_data', for several different values of 'alpha' that are automatically chosen (by default, 4 values of alpha are chosen)

Note that foreground_data and background_data should be 2D numpy arrays that have the second dimension (which represents the number of features). In other words, foreground_data.shape[1]==background_data.shape[1] should return True.

Built-in plotting: to quickly see the results of contrastive PCA, simply enable the plot parameter to true:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, plot=True)

images/plot_true.png

Interactive GUI: if you are running these analyses inside a jupyter notebook, you can easily launch an interactive GUI as shown here:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, gui=True)

images/gui_true.png

Using the slider, you can see how the your data points move as you change the value of the contrast parameter. These animations can reveal groups in the data and other insights:

images/animation.gif

Quick Test

To ensure that the library is working, here is a quick script that will allow you to test the code on synthetic data. Simply run the following commands:

importnumpyasnpfromcontrastiveimportCPCAN=400; D=30; gap=3# In B, all the data pts are from the same distribution, which has different variances in three subspaces.B=np.zeros((N, D))
B[:,0:10] =np.random.normal(0,10,(N,10))
B[:,10:20] =np.random.normal(0,3,(N,10))
B[:,20:30] =np.random.normal(0,1,(N,10))
# In A there are four clusters.A=np.zeros((N, D))
A[:,0:10] =np.random.normal(0,10,(N,10))
# group 1A[0:100, 10:20] =np.random.normal(0,1,(100,10))
A[0:100, 20:30] =np.random.normal(0,1,(100,10))
# group 2A[100:200, 10:20] =np.random.normal(0,1,(100,10))
A[100:200, 20:30] =np.random.normal(gap,1,(100,10))
# group 3A[200:300, 10:20] =np.random.normal(2*gap,1,(100,10))
A[200:300, 20:30] =np.random.normal(0,1,(100,10))
# group 4A[300:400, 10:20] =np.random.normal(2*gap,1,(100,10))
A[300:400, 20:30] =np.random.normal(gap,1,(100,10))
A_labels= [0]*100+[1]*100+[2]*100+[3]*100cpca=CPCA(standardize=False)
cpca.fit_transform(A, B, plot=True, active_labels=A_labels)

You should see a series of plots that looks something like this:

images/plot_example.png

Optional Parameters

Labels for foreground data (plot/gui mode): In the examples above, the data points are colored according to labels known ahead of time. You can supply these labels using the active_labels parameter, as shown here:

fromcontrastiveimportCPCAmdl=CPCA()
#labels = [0, 1, 0, 1, 1 ... 1, 0]projected_data=mdl.fit_transform(foreground_data, background_data, plot=True, active_labels=labels)

Additional # of components: Sometimes, you'd like to project your data on more than the top 2 contrastive principal components (cPCs). Specify the number of cPCs when you instantiate your model using the n_components parameter:

fromcontrastiveimportCPCAmdl=CPCA(n_components=3) #the top 3 components will be returnedprojected_data=mdl.fit_transform(foreground_data, background_data)

However, note that only when n_components=2 can the data be plotted or visualized through the GUI.

How values of alpha are chosen: So far, we've always plotted the data when the values of alpha have been chosen automatically with default parameters. However, the values of alpha can be customized. For example, if you'd like to still choose the values of alpha automatically, but change the range or number of alphas considered, you can use the n_alphas and max_log_alpha parameters. The former sets the number of alphas that are analyzed, and the latter sets the upper bound on the highest value of log (base 10) alpha. (The minimum value of alpha, besides alpha = 0, is always alpha = 0.1). Finally, you can change the number of values of alpha that are returned using the n_alphas_to_return parameter.

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, n_alphas=10, max_log_alpha=2, n_alphas_to_return=1) #search through 10 logarithmically spaced values of alpha from 0.1 to 100 and return the PCs for only 1 of them.

You can also decide to set the value of alpha to a particular value of alpha manually by changing the alpha_selection and alpha_value parameters as follows:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, alpha_selection='manual', alpha_value=2.0)

Or you can decide to plot or return the data for _all_ values of alpha in the given range. In this case, you can still choose to set the n_alphas and max_log_alpha parameters:

fromcontrastiveimportCPCAmdl=CPCA() #the top 3 components will be returnedprojected_data=mdl.fit_transform(foreground_data, background_data, n_alphas=10, max_log_alpha=2, alpha_selection='all') #search through 10 logarithmically spaced values of alpha from 0.1 to 100 and return the PCs for all of them!

Whether to standardize your data: By default, before performing contrastive PCA, the data are standardized so that each column or dimension has unit variance. You can turn this off by doing the following:

fromcontrastiveimportCPCAmdl=CPCA(standardize=False)
projected_data=mdl.fit_transform(foreground_data, background_data)

Custom colors (plot/gui mode): As a stylistic touch, you can also customize which colors are used to label the points when the data is plotted by using the colors argument. Here's an example:

fromcontrastiveimportCPCAmdl=CPCA(standardize=False)
projected_data=mdl.fit_transform(foreground_data, background_data, gui=True, colors=['r','b','k','c'])

will produce something along the lines of:

images/gui_colors.png

About

a faster Contrastive PCA with ability to only output correction vectors.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

63 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

contrastive

A python library for performing unsupervised machine learning on datasets with learning (e.g. PCA) in contrastive settings, where one is interested in patterns (e.g. clusters or clines) that exist one dataset, but not the other.

Applications include dicovering subgroups in biological and medical data. Here are basic installation and usage instructions, written for Python 3 (in which the library has been developed and tested, although it should work in Python 2 as well).

For more details, see the accompanying paper: "Exploring Patterns Enriched in a Dataset with Contrastive Principal Component Analysis", Nature Communications (2018), and please use the citation below.

@article{abid2018exploring,
title={Exploring patterns enriched in a dataset with contrastive principal component analysis},
author={Abid, Abubakar and Zhang, Martin J and Bagaria, Vivek K and Zou, James},
journal={Nature communications},
volume={9},
number={1},
pages={2134},
year={2018},
}

This repository also includes experiments to reproduce most of the figures in the paper. Please see the python notebooks in the experiments folder.

Installation

$ pip3 install contrastive

Basic Usage

The basic functions enabled by this library are shown below. Generally speaking, we have two datasets, one is a dataset that we can label as foreground_data, which is the dataset in which we are discovering patterns and directions, and another dataset called background_data, which is the dataset that does not have the patterns or directions we are interested in discovering. In some cases, both datasets may contain the signal of interest, but the foreground dataset may have the pattern enriched relative to the background. In these analyses, there is a contrast parameter, known as alpha, which can be thought of as a hyperparameter.

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data)
#returns a set of 2-dimensional projections of the foreground data stored in the list 'projected_data', for several different values of 'alpha' that are automatically chosen (by default, 4 values of alpha are chosen)

Note that foreground_data and background_data should be 2D numpy arrays that have the second dimension (which represents the number of features). In other words, foreground_data.shape[1]==background_data.shape[1] should return True.

Built-in plotting: to quickly see the results of contrastive PCA, simply enable the plot parameter to true:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, plot=True)

images/plot_true.png

Interactive GUI: if you are running these analyses inside a jupyter notebook, you can easily launch an interactive GUI as shown here:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, gui=True)

images/gui_true.png

Using the slider, you can see how the your data points move as you change the value of the contrast parameter. These animations can reveal groups in the data and other insights:

images/animation.gif

Quick Test

To ensure that the library is working, here is a quick script that will allow you to test the code on synthetic data. Simply run the following commands:

importnumpyasnpfromcontrastiveimportCPCAN=400; D=30; gap=3# In B, all the data pts are from the same distribution, which has different variances in three subspaces.B=np.zeros((N, D))
B[:,0:10] =np.random.normal(0,10,(N,10))
B[:,10:20] =np.random.normal(0,3,(N,10))
B[:,20:30] =np.random.normal(0,1,(N,10))
# In A there are four clusters.A=np.zeros((N, D))
A[:,0:10] =np.random.normal(0,10,(N,10))
# group 1A[0:100, 10:20] =np.random.normal(0,1,(100,10))
A[0:100, 20:30] =np.random.normal(0,1,(100,10))
# group 2A[100:200, 10:20] =np.random.normal(0,1,(100,10))
A[100:200, 20:30] =np.random.normal(gap,1,(100,10))
# group 3A[200:300, 10:20] =np.random.normal(2*gap,1,(100,10))
A[200:300, 20:30] =np.random.normal(0,1,(100,10))
# group 4A[300:400, 10:20] =np.random.normal(2*gap,1,(100,10))
A[300:400, 20:30] =np.random.normal(gap,1,(100,10))
A_labels= [0]*100+[1]*100+[2]*100+[3]*100cpca=CPCA(standardize=False)
cpca.fit_transform(A, B, plot=True, active_labels=A_labels)

You should see a series of plots that looks something like this:

images/plot_example.png

Optional Parameters

Labels for foreground data (plot/gui mode): In the examples above, the data points are colored according to labels known ahead of time. You can supply these labels using the active_labels parameter, as shown here:

fromcontrastiveimportCPCAmdl=CPCA()
#labels = [0, 1, 0, 1, 1 ... 1, 0]projected_data=mdl.fit_transform(foreground_data, background_data, plot=True, active_labels=labels)

Additional # of components: Sometimes, you'd like to project your data on more than the top 2 contrastive principal components (cPCs). Specify the number of cPCs when you instantiate your model using the n_components parameter:

fromcontrastiveimportCPCAmdl=CPCA(n_components=3) #the top 3 components will be returnedprojected_data=mdl.fit_transform(foreground_data, background_data)

However, note that only when n_components=2 can the data be plotted or visualized through the GUI.

How values of alpha are chosen: So far, we've always plotted the data when the values of alpha have been chosen automatically with default parameters. However, the values of alpha can be customized. For example, if you'd like to still choose the values of alpha automatically, but change the range or number of alphas considered, you can use the n_alphas and max_log_alpha parameters. The former sets the number of alphas that are analyzed, and the latter sets the upper bound on the highest value of log (base 10) alpha. (The minimum value of alpha, besides alpha = 0, is always alpha = 0.1). Finally, you can change the number of values of alpha that are returned using the n_alphas_to_return parameter.

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, n_alphas=10, max_log_alpha=2, n_alphas_to_return=1) #search through 10 logarithmically spaced values of alpha from 0.1 to 100 and return the PCs for only 1 of them.

You can also decide to set the value of alpha to a particular value of alpha manually by changing the alpha_selection and alpha_value parameters as follows:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, alpha_selection='manual', alpha_value=2.0)

Or you can decide to plot or return the data for _all_ values of alpha in the given range. In this case, you can still choose to set the n_alphas and max_log_alpha parameters:

fromcontrastiveimportCPCAmdl=CPCA() #the top 3 components will be returnedprojected_data=mdl.fit_transform(foreground_data, background_data, n_alphas=10, max_log_alpha=2, alpha_selection='all') #search through 10 logarithmically spaced values of alpha from 0.1 to 100 and return the PCs for all of them!

Whether to standardize your data: By default, before performing contrastive PCA, the data are standardized so that each column or dimension has unit variance. You can turn this off by doing the following:

fromcontrastiveimportCPCAmdl=CPCA(standardize=False)
projected_data=mdl.fit_transform(foreground_data, background_data)

Custom colors (plot/gui mode): As a stylistic touch, you can also customize which colors are used to label the points when the data is plotted by using the colors argument. Here's an example:

fromcontrastiveimportCPCAmdl=CPCA(standardize=False)
projected_data=mdl.fit_transform(foreground_data, background_data, gui=True, colors=['r','b','k','c'])

will produce something along the lines of:

images/gui_colors.png

About

a faster Contrastive PCA with ability to only output correction vectors.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

63 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

contrastive

A python library for performing unsupervised machine learning on datasets with learning (e.g. PCA) in contrastive settings, where one is interested in patterns (e.g. clusters or clines) that exist one dataset, but not the other.

Applications include dicovering subgroups in biological and medical data. Here are basic installation and usage instructions, written for Python 3 (in which the library has been developed and tested, although it should work in Python 2 as well).

For more details, see the accompanying paper: "Exploring Patterns Enriched in a Dataset with Contrastive Principal Component Analysis", Nature Communications (2018), and please use the citation below.

@article{abid2018exploring,
title={Exploring patterns enriched in a dataset with contrastive principal component analysis},
author={Abid, Abubakar and Zhang, Martin J and Bagaria, Vivek K and Zou, James},
journal={Nature communications},
volume={9},
number={1},
pages={2134},
year={2018},
}

This repository also includes experiments to reproduce most of the figures in the paper. Please see the python notebooks in the experiments folder.

Installation

$ pip3 install contrastive

Basic Usage

The basic functions enabled by this library are shown below. Generally speaking, we have two datasets, one is a dataset that we can label as foreground_data, which is the dataset in which we are discovering patterns and directions, and another dataset called background_data, which is the dataset that does not have the patterns or directions we are interested in discovering. In some cases, both datasets may contain the signal of interest, but the foreground dataset may have the pattern enriched relative to the background. In these analyses, there is a contrast parameter, known as alpha, which can be thought of as a hyperparameter.

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data)
#returns a set of 2-dimensional projections of the foreground data stored in the list 'projected_data', for several different values of 'alpha' that are automatically chosen (by default, 4 values of alpha are chosen)

Note that foreground_data and background_data should be 2D numpy arrays that have the second dimension (which represents the number of features). In other words, foreground_data.shape[1]==background_data.shape[1] should return True.

Built-in plotting: to quickly see the results of contrastive PCA, simply enable the plot parameter to true:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, plot=True)

images/plot_true.png

Interactive GUI: if you are running these analyses inside a jupyter notebook, you can easily launch an interactive GUI as shown here:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, gui=True)

images/gui_true.png

Using the slider, you can see how the your data points move as you change the value of the contrast parameter. These animations can reveal groups in the data and other insights:

images/animation.gif

Quick Test

To ensure that the library is working, here is a quick script that will allow you to test the code on synthetic data. Simply run the following commands:

importnumpyasnpfromcontrastiveimportCPCAN=400; D=30; gap=3# In B, all the data pts are from the same distribution, which has different variances in three subspaces.B=np.zeros((N, D))
B[:,0:10] =np.random.normal(0,10,(N,10))
B[:,10:20] =np.random.normal(0,3,(N,10))
B[:,20:30] =np.random.normal(0,1,(N,10))
# In A there are four clusters.A=np.zeros((N, D))
A[:,0:10] =np.random.normal(0,10,(N,10))
# group 1A[0:100, 10:20] =np.random.normal(0,1,(100,10))
A[0:100, 20:30] =np.random.normal(0,1,(100,10))
# group 2A[100:200, 10:20] =np.random.normal(0,1,(100,10))
A[100:200, 20:30] =np.random.normal(gap,1,(100,10))
# group 3A[200:300, 10:20] =np.random.normal(2*gap,1,(100,10))
A[200:300, 20:30] =np.random.normal(0,1,(100,10))
# group 4A[300:400, 10:20] =np.random.normal(2*gap,1,(100,10))
A[300:400, 20:30] =np.random.normal(gap,1,(100,10))
A_labels= [0]*100+[1]*100+[2]*100+[3]*100cpca=CPCA(standardize=False)
cpca.fit_transform(A, B, plot=True, active_labels=A_labels)

You should see a series of plots that looks something like this:

images/plot_example.png

Optional Parameters

Labels for foreground data (plot/gui mode): In the examples above, the data points are colored according to labels known ahead of time. You can supply these labels using the active_labels parameter, as shown here:

fromcontrastiveimportCPCAmdl=CPCA()
#labels = [0, 1, 0, 1, 1 ... 1, 0]projected_data=mdl.fit_transform(foreground_data, background_data, plot=True, active_labels=labels)

Additional # of components: Sometimes, you'd like to project your data on more than the top 2 contrastive principal components (cPCs). Specify the number of cPCs when you instantiate your model using the n_components parameter:

fromcontrastiveimportCPCAmdl=CPCA(n_components=3) #the top 3 components will be returnedprojected_data=mdl.fit_transform(foreground_data, background_data)

However, note that only when n_components=2 can the data be plotted or visualized through the GUI.

How values of alpha are chosen: So far, we've always plotted the data when the values of alpha have been chosen automatically with default parameters. However, the values of alpha can be customized. For example, if you'd like to still choose the values of alpha automatically, but change the range or number of alphas considered, you can use the n_alphas and max_log_alpha parameters. The former sets the number of alphas that are analyzed, and the latter sets the upper bound on the highest value of log (base 10) alpha. (The minimum value of alpha, besides alpha = 0, is always alpha = 0.1). Finally, you can change the number of values of alpha that are returned using the n_alphas_to_return parameter.

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, n_alphas=10, max_log_alpha=2, n_alphas_to_return=1) #search through 10 logarithmically spaced values of alpha from 0.1 to 100 and return the PCs for only 1 of them.

You can also decide to set the value of alpha to a particular value of alpha manually by changing the alpha_selection and alpha_value parameters as follows:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, alpha_selection='manual', alpha_value=2.0)

Or you can decide to plot or return the data for _all_ values of alpha in the given range. In this case, you can still choose to set the n_alphas and max_log_alpha parameters:

fromcontrastiveimportCPCAmdl=CPCA() #the top 3 components will be returnedprojected_data=mdl.fit_transform(foreground_data, background_data, n_alphas=10, max_log_alpha=2, alpha_selection='all') #search through 10 logarithmically spaced values of alpha from 0.1 to 100 and return the PCs for all of them!

Whether to standardize your data: By default, before performing contrastive PCA, the data are standardized so that each column or dimension has unit variance. You can turn this off by doing the following:

fromcontrastiveimportCPCAmdl=CPCA(standardize=False)
projected_data=mdl.fit_transform(foreground_data, background_data)

Custom colors (plot/gui mode): As a stylistic touch, you can also customize which colors are used to label the points when the data is plotted by using the colors argument. Here's an example:

fromcontrastiveimportCPCAmdl=CPCA(standardize=False)
projected_data=mdl.fit_transform(foreground_data, background_data, gui=True, colors=['r','b','k','c'])

will produce something along the lines of:

images/gui_colors.png

About

a faster Contrastive PCA with ability to only output correction vectors.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

63 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

contrastive

A python library for performing unsupervised machine learning on datasets with learning (e.g. PCA) in contrastive settings, where one is interested in patterns (e.g. clusters or clines) that exist one dataset, but not the other.

Applications include dicovering subgroups in biological and medical data. Here are basic installation and usage instructions, written for Python 3 (in which the library has been developed and tested, although it should work in Python 2 as well).

For more details, see the accompanying paper: "Exploring Patterns Enriched in a Dataset with Contrastive Principal Component Analysis", Nature Communications (2018), and please use the citation below.

@article{abid2018exploring,
title={Exploring patterns enriched in a dataset with contrastive principal component analysis},
author={Abid, Abubakar and Zhang, Martin J and Bagaria, Vivek K and Zou, James},
journal={Nature communications},
volume={9},
number={1},
pages={2134},
year={2018},
}

This repository also includes experiments to reproduce most of the figures in the paper. Please see the python notebooks in the experiments folder.

Installation

$ pip3 install contrastive

Basic Usage

The basic functions enabled by this library are shown below. Generally speaking, we have two datasets, one is a dataset that we can label as foreground_data, which is the dataset in which we are discovering patterns and directions, and another dataset called background_data, which is the dataset that does not have the patterns or directions we are interested in discovering. In some cases, both datasets may contain the signal of interest, but the foreground dataset may have the pattern enriched relative to the background. In these analyses, there is a contrast parameter, known as alpha, which can be thought of as a hyperparameter.

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data)
#returns a set of 2-dimensional projections of the foreground data stored in the list 'projected_data', for several different values of 'alpha' that are automatically chosen (by default, 4 values of alpha are chosen)

Note that foreground_data and background_data should be 2D numpy arrays that have the second dimension (which represents the number of features). In other words, foreground_data.shape[1]==background_data.shape[1] should return True.

Built-in plotting: to quickly see the results of contrastive PCA, simply enable the plot parameter to true:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, plot=True)

images/plot_true.png

Interactive GUI: if you are running these analyses inside a jupyter notebook, you can easily launch an interactive GUI as shown here:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, gui=True)

images/gui_true.png

Using the slider, you can see how the your data points move as you change the value of the contrast parameter. These animations can reveal groups in the data and other insights:

images/animation.gif

Quick Test

To ensure that the library is working, here is a quick script that will allow you to test the code on synthetic data. Simply run the following commands:

importnumpyasnpfromcontrastiveimportCPCAN=400; D=30; gap=3# In B, all the data pts are from the same distribution, which has different variances in three subspaces.B=np.zeros((N, D))
B[:,0:10] =np.random.normal(0,10,(N,10))
B[:,10:20] =np.random.normal(0,3,(N,10))
B[:,20:30] =np.random.normal(0,1,(N,10))
# In A there are four clusters.A=np.zeros((N, D))
A[:,0:10] =np.random.normal(0,10,(N,10))
# group 1A[0:100, 10:20] =np.random.normal(0,1,(100,10))
A[0:100, 20:30] =np.random.normal(0,1,(100,10))
# group 2A[100:200, 10:20] =np.random.normal(0,1,(100,10))
A[100:200, 20:30] =np.random.normal(gap,1,(100,10))
# group 3A[200:300, 10:20] =np.random.normal(2*gap,1,(100,10))
A[200:300, 20:30] =np.random.normal(0,1,(100,10))
# group 4A[300:400, 10:20] =np.random.normal(2*gap,1,(100,10))
A[300:400, 20:30] =np.random.normal(gap,1,(100,10))
A_labels= [0]*100+[1]*100+[2]*100+[3]*100cpca=CPCA(standardize=False)
cpca.fit_transform(A, B, plot=True, active_labels=A_labels)

You should see a series of plots that looks something like this:

images/plot_example.png

Optional Parameters

Labels for foreground data (plot/gui mode): In the examples above, the data points are colored according to labels known ahead of time. You can supply these labels using the active_labels parameter, as shown here:

fromcontrastiveimportCPCAmdl=CPCA()
#labels = [0, 1, 0, 1, 1 ... 1, 0]projected_data=mdl.fit_transform(foreground_data, background_data, plot=True, active_labels=labels)

Additional # of components: Sometimes, you'd like to project your data on more than the top 2 contrastive principal components (cPCs). Specify the number of cPCs when you instantiate your model using the n_components parameter:

fromcontrastiveimportCPCAmdl=CPCA(n_components=3) #the top 3 components will be returnedprojected_data=mdl.fit_transform(foreground_data, background_data)

However, note that only when n_components=2 can the data be plotted or visualized through the GUI.

How values of alpha are chosen: So far, we've always plotted the data when the values of alpha have been chosen automatically with default parameters. However, the values of alpha can be customized. For example, if you'd like to still choose the values of alpha automatically, but change the range or number of alphas considered, you can use the n_alphas and max_log_alpha parameters. The former sets the number of alphas that are analyzed, and the latter sets the upper bound on the highest value of log (base 10) alpha. (The minimum value of alpha, besides alpha = 0, is always alpha = 0.1). Finally, you can change the number of values of alpha that are returned using the n_alphas_to_return parameter.

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, n_alphas=10, max_log_alpha=2, n_alphas_to_return=1) #search through 10 logarithmically spaced values of alpha from 0.1 to 100 and return the PCs for only 1 of them.

You can also decide to set the value of alpha to a particular value of alpha manually by changing the alpha_selection and alpha_value parameters as follows:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, alpha_selection='manual', alpha_value=2.0)

Or you can decide to plot or return the data for _all_ values of alpha in the given range. In this case, you can still choose to set the n_alphas and max_log_alpha parameters:

fromcontrastiveimportCPCAmdl=CPCA() #the top 3 components will be returnedprojected_data=mdl.fit_transform(foreground_data, background_data, n_alphas=10, max_log_alpha=2, alpha_selection='all') #search through 10 logarithmically spaced values of alpha from 0.1 to 100 and return the PCs for all of them!

Whether to standardize your data: By default, before performing contrastive PCA, the data are standardized so that each column or dimension has unit variance. You can turn this off by doing the following:

fromcontrastiveimportCPCAmdl=CPCA(standardize=False)
projected_data=mdl.fit_transform(foreground_data, background_data)

Custom colors (plot/gui mode): As a stylistic touch, you can also customize which colors are used to label the points when the data is plotted by using the colors argument. Here's an example:

fromcontrastiveimportCPCAmdl=CPCA(standardize=False)
projected_data=mdl.fit_transform(foreground_data, background_data, gui=True, colors=['r','b','k','c'])

will produce something along the lines of:

images/gui_colors.png

About

a faster Contrastive PCA with ability to only output correction vectors.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

63 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

contrastive

A python library for performing unsupervised machine learning on datasets with learning (e.g. PCA) in contrastive settings, where one is interested in patterns (e.g. clusters or clines) that exist one dataset, but not the other.

Applications include dicovering subgroups in biological and medical data. Here are basic installation and usage instructions, written for Python 3 (in which the library has been developed and tested, although it should work in Python 2 as well).

For more details, see the accompanying paper: "Exploring Patterns Enriched in a Dataset with Contrastive Principal Component Analysis", Nature Communications (2018), and please use the citation below.

@article{abid2018exploring,
title={Exploring patterns enriched in a dataset with contrastive principal component analysis},
author={Abid, Abubakar and Zhang, Martin J and Bagaria, Vivek K and Zou, James},
journal={Nature communications},
volume={9},
number={1},
pages={2134},
year={2018},
}

This repository also includes experiments to reproduce most of the figures in the paper. Please see the python notebooks in the experiments folder.

Installation

$ pip3 install contrastive

Basic Usage

The basic functions enabled by this library are shown below. Generally speaking, we have two datasets, one is a dataset that we can label as foreground_data, which is the dataset in which we are discovering patterns and directions, and another dataset called background_data, which is the dataset that does not have the patterns or directions we are interested in discovering. In some cases, both datasets may contain the signal of interest, but the foreground dataset may have the pattern enriched relative to the background. In these analyses, there is a contrast parameter, known as alpha, which can be thought of as a hyperparameter.

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data)
#returns a set of 2-dimensional projections of the foreground data stored in the list 'projected_data', for several different values of 'alpha' that are automatically chosen (by default, 4 values of alpha are chosen)

Note that foreground_data and background_data should be 2D numpy arrays that have the second dimension (which represents the number of features). In other words, foreground_data.shape[1]==background_data.shape[1] should return True.

Built-in plotting: to quickly see the results of contrastive PCA, simply enable the plot parameter to true:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, plot=True)

images/plot_true.png

Interactive GUI: if you are running these analyses inside a jupyter notebook, you can easily launch an interactive GUI as shown here:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, gui=True)

images/gui_true.png

Using the slider, you can see how the your data points move as you change the value of the contrast parameter. These animations can reveal groups in the data and other insights:

images/animation.gif

Quick Test

To ensure that the library is working, here is a quick script that will allow you to test the code on synthetic data. Simply run the following commands:

importnumpyasnpfromcontrastiveimportCPCAN=400; D=30; gap=3# In B, all the data pts are from the same distribution, which has different variances in three subspaces.B=np.zeros((N, D))
B[:,0:10] =np.random.normal(0,10,(N,10))
B[:,10:20] =np.random.normal(0,3,(N,10))
B[:,20:30] =np.random.normal(0,1,(N,10))
# In A there are four clusters.A=np.zeros((N, D))
A[:,0:10] =np.random.normal(0,10,(N,10))
# group 1A[0:100, 10:20] =np.random.normal(0,1,(100,10))
A[0:100, 20:30] =np.random.normal(0,1,(100,10))
# group 2A[100:200, 10:20] =np.random.normal(0,1,(100,10))
A[100:200, 20:30] =np.random.normal(gap,1,(100,10))
# group 3A[200:300, 10:20] =np.random.normal(2*gap,1,(100,10))
A[200:300, 20:30] =np.random.normal(0,1,(100,10))
# group 4A[300:400, 10:20] =np.random.normal(2*gap,1,(100,10))
A[300:400, 20:30] =np.random.normal(gap,1,(100,10))
A_labels= [0]*100+[1]*100+[2]*100+[3]*100cpca=CPCA(standardize=False)
cpca.fit_transform(A, B, plot=True, active_labels=A_labels)

You should see a series of plots that looks something like this:

images/plot_example.png

Optional Parameters

Labels for foreground data (plot/gui mode): In the examples above, the data points are colored according to labels known ahead of time. You can supply these labels using the active_labels parameter, as shown here:

fromcontrastiveimportCPCAmdl=CPCA()
#labels = [0, 1, 0, 1, 1 ... 1, 0]projected_data=mdl.fit_transform(foreground_data, background_data, plot=True, active_labels=labels)

Additional # of components: Sometimes, you'd like to project your data on more than the top 2 contrastive principal components (cPCs). Specify the number of cPCs when you instantiate your model using the n_components parameter:

fromcontrastiveimportCPCAmdl=CPCA(n_components=3) #the top 3 components will be returnedprojected_data=mdl.fit_transform(foreground_data, background_data)

However, note that only when n_components=2 can the data be plotted or visualized through the GUI.

How values of alpha are chosen: So far, we've always plotted the data when the values of alpha have been chosen automatically with default parameters. However, the values of alpha can be customized. For example, if you'd like to still choose the values of alpha automatically, but change the range or number of alphas considered, you can use the n_alphas and max_log_alpha parameters. The former sets the number of alphas that are analyzed, and the latter sets the upper bound on the highest value of log (base 10) alpha. (The minimum value of alpha, besides alpha = 0, is always alpha = 0.1). Finally, you can change the number of values of alpha that are returned using the n_alphas_to_return parameter.

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, n_alphas=10, max_log_alpha=2, n_alphas_to_return=1) #search through 10 logarithmically spaced values of alpha from 0.1 to 100 and return the PCs for only 1 of them.

You can also decide to set the value of alpha to a particular value of alpha manually by changing the alpha_selection and alpha_value parameters as follows:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, alpha_selection='manual', alpha_value=2.0)

Or you can decide to plot or return the data for _all_ values of alpha in the given range. In this case, you can still choose to set the n_alphas and max_log_alpha parameters:

fromcontrastiveimportCPCAmdl=CPCA() #the top 3 components will be returnedprojected_data=mdl.fit_transform(foreground_data, background_data, n_alphas=10, max_log_alpha=2, alpha_selection='all') #search through 10 logarithmically spaced values of alpha from 0.1 to 100 and return the PCs for all of them!

Whether to standardize your data: By default, before performing contrastive PCA, the data are standardized so that each column or dimension has unit variance. You can turn this off by doing the following:

fromcontrastiveimportCPCAmdl=CPCA(standardize=False)
projected_data=mdl.fit_transform(foreground_data, background_data)

Custom colors (plot/gui mode): As a stylistic touch, you can also customize which colors are used to label the points when the data is plotted by using the colors argument. Here's an example:

fromcontrastiveimportCPCAmdl=CPCA(standardize=False)
projected_data=mdl.fit_transform(foreground_data, background_data, gui=True, colors=['r','b','k','c'])

will produce something along the lines of:

images/gui_colors.png

About

a faster Contrastive PCA with ability to only output correction vectors.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

63 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

contrastive

A python library for performing unsupervised machine learning on datasets with learning (e.g. PCA) in contrastive settings, where one is interested in patterns (e.g. clusters or clines) that exist one dataset, but not the other.

Applications include dicovering subgroups in biological and medical data. Here are basic installation and usage instructions, written for Python 3 (in which the library has been developed and tested, although it should work in Python 2 as well).

For more details, see the accompanying paper: "Exploring Patterns Enriched in a Dataset with Contrastive Principal Component Analysis", Nature Communications (2018), and please use the citation below.

@article{abid2018exploring,
title={Exploring patterns enriched in a dataset with contrastive principal component analysis},
author={Abid, Abubakar and Zhang, Martin J and Bagaria, Vivek K and Zou, James},
journal={Nature communications},
volume={9},
number={1},
pages={2134},
year={2018},
}

This repository also includes experiments to reproduce most of the figures in the paper. Please see the python notebooks in the experiments folder.

Installation

$ pip3 install contrastive

Basic Usage

The basic functions enabled by this library are shown below. Generally speaking, we have two datasets, one is a dataset that we can label as foreground_data, which is the dataset in which we are discovering patterns and directions, and another dataset called background_data, which is the dataset that does not have the patterns or directions we are interested in discovering. In some cases, both datasets may contain the signal of interest, but the foreground dataset may have the pattern enriched relative to the background. In these analyses, there is a contrast parameter, known as alpha, which can be thought of as a hyperparameter.

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data)
#returns a set of 2-dimensional projections of the foreground data stored in the list 'projected_data', for several different values of 'alpha' that are automatically chosen (by default, 4 values of alpha are chosen)

Note that foreground_data and background_data should be 2D numpy arrays that have the second dimension (which represents the number of features). In other words, foreground_data.shape[1]==background_data.shape[1] should return True.

Built-in plotting: to quickly see the results of contrastive PCA, simply enable the plot parameter to true:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, plot=True)

images/plot_true.png

Interactive GUI: if you are running these analyses inside a jupyter notebook, you can easily launch an interactive GUI as shown here:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, gui=True)

images/gui_true.png

Using the slider, you can see how the your data points move as you change the value of the contrast parameter. These animations can reveal groups in the data and other insights:

images/animation.gif

Quick Test

To ensure that the library is working, here is a quick script that will allow you to test the code on synthetic data. Simply run the following commands:

importnumpyasnpfromcontrastiveimportCPCAN=400; D=30; gap=3# In B, all the data pts are from the same distribution, which has different variances in three subspaces.B=np.zeros((N, D))
B[:,0:10] =np.random.normal(0,10,(N,10))
B[:,10:20] =np.random.normal(0,3,(N,10))
B[:,20:30] =np.random.normal(0,1,(N,10))
# In A there are four clusters.A=np.zeros((N, D))
A[:,0:10] =np.random.normal(0,10,(N,10))
# group 1A[0:100, 10:20] =np.random.normal(0,1,(100,10))
A[0:100, 20:30] =np.random.normal(0,1,(100,10))
# group 2A[100:200, 10:20] =np.random.normal(0,1,(100,10))
A[100:200, 20:30] =np.random.normal(gap,1,(100,10))
# group 3A[200:300, 10:20] =np.random.normal(2*gap,1,(100,10))
A[200:300, 20:30] =np.random.normal(0,1,(100,10))
# group 4A[300:400, 10:20] =np.random.normal(2*gap,1,(100,10))
A[300:400, 20:30] =np.random.normal(gap,1,(100,10))
A_labels= [0]*100+[1]*100+[2]*100+[3]*100cpca=CPCA(standardize=False)
cpca.fit_transform(A, B, plot=True, active_labels=A_labels)

You should see a series of plots that looks something like this:

images/plot_example.png

Optional Parameters

Labels for foreground data (plot/gui mode): In the examples above, the data points are colored according to labels known ahead of time. You can supply these labels using the active_labels parameter, as shown here:

fromcontrastiveimportCPCAmdl=CPCA()
#labels = [0, 1, 0, 1, 1 ... 1, 0]projected_data=mdl.fit_transform(foreground_data, background_data, plot=True, active_labels=labels)

Additional # of components: Sometimes, you'd like to project your data on more than the top 2 contrastive principal components (cPCs). Specify the number of cPCs when you instantiate your model using the n_components parameter:

fromcontrastiveimportCPCAmdl=CPCA(n_components=3) #the top 3 components will be returnedprojected_data=mdl.fit_transform(foreground_data, background_data)

However, note that only when n_components=2 can the data be plotted or visualized through the GUI.

How values of alpha are chosen: So far, we've always plotted the data when the values of alpha have been chosen automatically with default parameters. However, the values of alpha can be customized. For example, if you'd like to still choose the values of alpha automatically, but change the range or number of alphas considered, you can use the n_alphas and max_log_alpha parameters. The former sets the number of alphas that are analyzed, and the latter sets the upper bound on the highest value of log (base 10) alpha. (The minimum value of alpha, besides alpha = 0, is always alpha = 0.1). Finally, you can change the number of values of alpha that are returned using the n_alphas_to_return parameter.

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, n_alphas=10, max_log_alpha=2, n_alphas_to_return=1) #search through 10 logarithmically spaced values of alpha from 0.1 to 100 and return the PCs for only 1 of them.

You can also decide to set the value of alpha to a particular value of alpha manually by changing the alpha_selection and alpha_value parameters as follows:

fromcontrastiveimportCPCAmdl=CPCA()
projected_data=mdl.fit_transform(foreground_data, background_data, alpha_selection='manual', alpha_value=2.0)

Or you can decide to plot or return the data for _all_ values of alpha in the given range. In this case, you can still choose to set the n_alphas and max_log_alpha parameters:

fromcontrastiveimportCPCAmdl=CPCA() #the top 3 components will be returnedprojected_data=mdl.fit_transform(foreground_data, background_data, n_alphas=10, max_log_alpha=2, alpha_selection='all') #search through 10 logarithmically spaced values of alpha from 0.1 to 100 and return the PCs for all of them!

Whether to standardize your data: By default, before performing contrastive PCA, the data are standardized so that each column or dimension has unit variance. You can turn this off by doing the following:

fromcontrastiveimportCPCAmdl=CPCA(standardize=False)
projected_data=mdl.fit_transform(foreground_data, background_data)

Custom colors (plot/gui mode): As a stylistic touch, you can also customize which colors are used to label the points when the data is plotted by using the colors argument. Here's an example:

fromcontrastiveimportCPCAmdl=CPCA(standardize=False)
projected_data=mdl.fit_transform(foreground_data, background_data, gui=True, colors=['r','b','k','c'])

will produce something along the lines of:

images/gui_colors.png

About

a faster Contrastive PCA with ability to only output correction vectors.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages