Repository files navigation

Documentation StatusPyPIDownloadsGitHub stars

Synthetic Data: Utility, Regulatory compliance, and Ethical privacy

The SURE package is an open-source Python library intended to be used for the assessment of the utility and privacy performance of any tabular synthetic dataset.

The SURE library features multiple Python modules that can be easily imported and seamlessly integrated into any Python script after installing the library.

Warning

This is a beta version of the library and only runs on Linux and MacOS for the moment.

Important

Requires Python >= 3.10

Installation

To install the library run the following command in your terminal:

$ pip install clearbox-sure

Modules overview

The SURE library features the following modules:

  1. Preprocessor
  2. Statistical similarity metrics
  3. Model garden
  4. ML utility metrics
  5. Distance metrics
  6. Privacy attack sandbox
  7. Report generator

Preprocessor

The input datasets undergo manipulation by the preprocessor module, tailored to conform to the standard structure utilized across the subsequent processes. The Polars library used in the preprocessor makes this operation significantly faster compared to the use of other data processing libraries.

Utility

The statistical similarity metrics, the ML utility metrics and the model garden modules constitute the data utility evaluation part.

The statistical similarity module and the distance metrics module take as input the pre-processed datasets and carry out the operation to assess the statistical similarity between the datasets and how different the content of the synthetic dataset is from the one of the original dataset. In particular, The real and synthetic input datasets are used in the statistical similarity metrics module to assess how close the two datasets are in terms of statistical properties, such as mean, correlation, distribution.

The model garden executes a classification or regression task on the given dataset with multiple machine learning models, returning the performance metrics of each of the models tested on the given task and dataset.

The model garden module’s best performing models are employed in the machine learning utility metrics module to compute the usefulness of the synthetic data on a given ML task (classification or regression).

Privacy

The distance metrics and the privacy attack sandbox make up the synthetic data privacy assessment modules.

The distance metrics module computes the Gower distance between the two input datasets and the distance to the closest record for each line of the first dataset.

The ML privacy attack sandbox allows to simulate a Membership Inference Attack for re-identification of vulnerable records identified with the distance metrics module and evaluate how exposed the synthetic dataset is to this kind of assault.

Report

Eventually, the report generator provides a summary of the utility and privacy metrics computed in the previous modules, providing a visual digest with charts and tables of the results.

This following diagram serves as a visual representation of how each module contributes to the utility-privacy assessment process and highlights the seamless interconnection and synergy between individual blocks.

drawing

Usage

The library leverages Polars, which ensures faster computations compared to other data manipulation libraries. It supports both Polars and Pandas dataframes.

The user must provide both the original real training dataset (which was used to train the generative model that produced the synthetic dataset), the real holdout dataset (which was NOT used to train the generative model that produced the synthetic dataset) and the corresponding synthetic dataset to enable the library's modules to perform the necessary computations for evaluation.

Below is a code snippet example for the usage of the library:

# Import the necessary modules from the SURE libraryfromsureimportPreprocessor, reportfromsure.utilityimport (compute_statistical_metrics, compute_mutual_info,
compute_utility_metrics_class,
detection,
query_power)
fromsure.privacyimport (distance_to_closest_record, dcr_stats, number_of_dcr_equal_to_zero, validation_dcr_test, adversary_dataset, membership_inference_test)
# Assuming real_data, valid_data and synth_data are three pandas DataFrames# Preprocessor initialization and query execution on the real, synthetic and validation datasetspreprocessor=Preprocessor(real_data, num_fill_null='forward', scaling='standardize')
real_data_preprocessed=preprocessor.transform(real_data)
valid_data_preprocessed=preprocessor.transform(valid_data)
synth_data_preprocessed=preprocessor.transform(synth_data)
# Statistical properties and mutual informationnum_features_stats, cat_features_stats, temporal_feat_stats=compute_statistical_metrics(real_data, synth_data)
corr_real, corr_synth, corr_difference=compute_mutual_info(real_data_preprocessed, synth_data_preprocessed)
# ML utility: TSTR - Train on Synthetic, Test on RealX_train=real_data_preprocessed.drop("label", axis=1) # Assuming the datasets have a “label” column for the machine learning task they are intended fory_train=real_data_preprocessed["label"]
X_synth=synth_data_preprocessed.drop("label", axis=1)
y_synth=synth_data_preprocessed["label"]
X_test=valid_data_preprocessed.drop("label", axis=1).limit(10000) # Test the trained models on a portion of the original real dataset (first 10k rows)y_test=valid_data_preprocessed["label"].limit(10000)
TSTR_metrics=compute_utility_metrics_class(X_train, X_synth, X_test, y_train, y_synth, y_test)
# Distance to closest recorddcr_synth_train=distance_to_closest_record("synth_train", synth_data, real_data)
dcr_synth_valid=distance_to_closest_record("synth_val", synth_data, valid_data)
dcr_stats_synth_train=dcr_stats("synth_train", dcr_synth_train)
dcr_stats_synth_valid=dcr_stats("synth_val", dcr_synth_valid)
dcr_zero_synth_train=number_of_dcr_equal_to_zero("synth_train", dcr_synth_train)
dcr_zero_synth_valid=number_of_dcr_equal_to_zero("synth_val", dcr_synth_valid)
share=validation_dcr_test(dcr_synth_train, dcr_synth_valid)
# Detection Scoredetection_score=detection(real_data, synth_data, preprocessor)
# Query Powerquery_power_score=query_power(real_data, synth_data, preprocessor)
# ML privacy attack sandbox initialization and simulationadversary_df=adversary_dataset(real_data, valid_data)
# The function adversary_dataset adds a column "privacy_test_is_training" to the adversary dataset, indicating whether the record was part of the training set or notadversary_guesses_ground_truth=adversary_df["privacy_test_is_training"] MIA=membership_inference_test(adversary_df, synth_data, adversary_guesses_ground_truth)
# Report generation as HTML pagereport(real_data, synth_data)

Follow the step-by-step guide to test the library.

About

An open-source Python library for the assessment of utility and privacy performance of any tabular synthetic dataset.

Resources

Stars

23 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Documentation StatusPyPIDownloadsGitHub stars

Synthetic Data: Utility, Regulatory compliance, and Ethical privacy

The SURE package is an open-source Python library intended to be used for the assessment of the utility and privacy performance of any tabular synthetic dataset.

The SURE library features multiple Python modules that can be easily imported and seamlessly integrated into any Python script after installing the library.

Warning

This is a beta version of the library and only runs on Linux and MacOS for the moment.

Important

Requires Python >= 3.10

Installation

To install the library run the following command in your terminal:

$ pip install clearbox-sure

Modules overview

The SURE library features the following modules:

  1. Preprocessor
  2. Statistical similarity metrics
  3. Model garden
  4. ML utility metrics
  5. Distance metrics
  6. Privacy attack sandbox
  7. Report generator

Preprocessor

The input datasets undergo manipulation by the preprocessor module, tailored to conform to the standard structure utilized across the subsequent processes. The Polars library used in the preprocessor makes this operation significantly faster compared to the use of other data processing libraries.

Utility

The statistical similarity metrics, the ML utility metrics and the model garden modules constitute the data utility evaluation part.

The statistical similarity module and the distance metrics module take as input the pre-processed datasets and carry out the operation to assess the statistical similarity between the datasets and how different the content of the synthetic dataset is from the one of the original dataset. In particular, The real and synthetic input datasets are used in the statistical similarity metrics module to assess how close the two datasets are in terms of statistical properties, such as mean, correlation, distribution.

The model garden executes a classification or regression task on the given dataset with multiple machine learning models, returning the performance metrics of each of the models tested on the given task and dataset.

The model garden module’s best performing models are employed in the machine learning utility metrics module to compute the usefulness of the synthetic data on a given ML task (classification or regression).

Privacy

The distance metrics and the privacy attack sandbox make up the synthetic data privacy assessment modules.

The distance metrics module computes the Gower distance between the two input datasets and the distance to the closest record for each line of the first dataset.

The ML privacy attack sandbox allows to simulate a Membership Inference Attack for re-identification of vulnerable records identified with the distance metrics module and evaluate how exposed the synthetic dataset is to this kind of assault.

Report

Eventually, the report generator provides a summary of the utility and privacy metrics computed in the previous modules, providing a visual digest with charts and tables of the results.

This following diagram serves as a visual representation of how each module contributes to the utility-privacy assessment process and highlights the seamless interconnection and synergy between individual blocks.

drawing

Usage

The library leverages Polars, which ensures faster computations compared to other data manipulation libraries. It supports both Polars and Pandas dataframes.

The user must provide both the original real training dataset (which was used to train the generative model that produced the synthetic dataset), the real holdout dataset (which was NOT used to train the generative model that produced the synthetic dataset) and the corresponding synthetic dataset to enable the library's modules to perform the necessary computations for evaluation.

Below is a code snippet example for the usage of the library:

# Import the necessary modules from the SURE libraryfromsureimportPreprocessor, reportfromsure.utilityimport (compute_statistical_metrics, compute_mutual_info,
compute_utility_metrics_class,
detection,
query_power)
fromsure.privacyimport (distance_to_closest_record, dcr_stats, number_of_dcr_equal_to_zero, validation_dcr_test, adversary_dataset, membership_inference_test)
# Assuming real_data, valid_data and synth_data are three pandas DataFrames# Preprocessor initialization and query execution on the real, synthetic and validation datasetspreprocessor=Preprocessor(real_data, num_fill_null='forward', scaling='standardize')
real_data_preprocessed=preprocessor.transform(real_data)
valid_data_preprocessed=preprocessor.transform(valid_data)
synth_data_preprocessed=preprocessor.transform(synth_data)
# Statistical properties and mutual informationnum_features_stats, cat_features_stats, temporal_feat_stats=compute_statistical_metrics(real_data, synth_data)
corr_real, corr_synth, corr_difference=compute_mutual_info(real_data_preprocessed, synth_data_preprocessed)
# ML utility: TSTR - Train on Synthetic, Test on RealX_train=real_data_preprocessed.drop("label", axis=1) # Assuming the datasets have a “label” column for the machine learning task they are intended fory_train=real_data_preprocessed["label"]
X_synth=synth_data_preprocessed.drop("label", axis=1)
y_synth=synth_data_preprocessed["label"]
X_test=valid_data_preprocessed.drop("label", axis=1).limit(10000) # Test the trained models on a portion of the original real dataset (first 10k rows)y_test=valid_data_preprocessed["label"].limit(10000)
TSTR_metrics=compute_utility_metrics_class(X_train, X_synth, X_test, y_train, y_synth, y_test)
# Distance to closest recorddcr_synth_train=distance_to_closest_record("synth_train", synth_data, real_data)
dcr_synth_valid=distance_to_closest_record("synth_val", synth_data, valid_data)
dcr_stats_synth_train=dcr_stats("synth_train", dcr_synth_train)
dcr_stats_synth_valid=dcr_stats("synth_val", dcr_synth_valid)
dcr_zero_synth_train=number_of_dcr_equal_to_zero("synth_train", dcr_synth_train)
dcr_zero_synth_valid=number_of_dcr_equal_to_zero("synth_val", dcr_synth_valid)
share=validation_dcr_test(dcr_synth_train, dcr_synth_valid)
# Detection Scoredetection_score=detection(real_data, synth_data, preprocessor)
# Query Powerquery_power_score=query_power(real_data, synth_data, preprocessor)
# ML privacy attack sandbox initialization and simulationadversary_df=adversary_dataset(real_data, valid_data)
# The function adversary_dataset adds a column "privacy_test_is_training" to the adversary dataset, indicating whether the record was part of the training set or notadversary_guesses_ground_truth=adversary_df["privacy_test_is_training"] MIA=membership_inference_test(adversary_df, synth_data, adversary_guesses_ground_truth)
# Report generation as HTML pagereport(real_data, synth_data)

Follow the step-by-step guide to test the library.

About

An open-source Python library for the assessment of utility and privacy performance of any tabular synthetic dataset.

Resources

Stars

23 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Documentation StatusPyPIDownloadsGitHub stars

Synthetic Data: Utility, Regulatory compliance, and Ethical privacy

The SURE package is an open-source Python library intended to be used for the assessment of the utility and privacy performance of any tabular synthetic dataset.

The SURE library features multiple Python modules that can be easily imported and seamlessly integrated into any Python script after installing the library.

Warning

This is a beta version of the library and only runs on Linux and MacOS for the moment.

Important

Requires Python >= 3.10

Installation

To install the library run the following command in your terminal:

$ pip install clearbox-sure

Modules overview

The SURE library features the following modules:

  1. Preprocessor
  2. Statistical similarity metrics
  3. Model garden
  4. ML utility metrics
  5. Distance metrics
  6. Privacy attack sandbox
  7. Report generator

Preprocessor

The input datasets undergo manipulation by the preprocessor module, tailored to conform to the standard structure utilized across the subsequent processes. The Polars library used in the preprocessor makes this operation significantly faster compared to the use of other data processing libraries.

Utility

The statistical similarity metrics, the ML utility metrics and the model garden modules constitute the data utility evaluation part.

The statistical similarity module and the distance metrics module take as input the pre-processed datasets and carry out the operation to assess the statistical similarity between the datasets and how different the content of the synthetic dataset is from the one of the original dataset. In particular, The real and synthetic input datasets are used in the statistical similarity metrics module to assess how close the two datasets are in terms of statistical properties, such as mean, correlation, distribution.

The model garden executes a classification or regression task on the given dataset with multiple machine learning models, returning the performance metrics of each of the models tested on the given task and dataset.

The model garden module’s best performing models are employed in the machine learning utility metrics module to compute the usefulness of the synthetic data on a given ML task (classification or regression).

Privacy

The distance metrics and the privacy attack sandbox make up the synthetic data privacy assessment modules.

The distance metrics module computes the Gower distance between the two input datasets and the distance to the closest record for each line of the first dataset.

The ML privacy attack sandbox allows to simulate a Membership Inference Attack for re-identification of vulnerable records identified with the distance metrics module and evaluate how exposed the synthetic dataset is to this kind of assault.

Report

Eventually, the report generator provides a summary of the utility and privacy metrics computed in the previous modules, providing a visual digest with charts and tables of the results.

This following diagram serves as a visual representation of how each module contributes to the utility-privacy assessment process and highlights the seamless interconnection and synergy between individual blocks.

drawing

Usage

The library leverages Polars, which ensures faster computations compared to other data manipulation libraries. It supports both Polars and Pandas dataframes.

The user must provide both the original real training dataset (which was used to train the generative model that produced the synthetic dataset), the real holdout dataset (which was NOT used to train the generative model that produced the synthetic dataset) and the corresponding synthetic dataset to enable the library's modules to perform the necessary computations for evaluation.

Below is a code snippet example for the usage of the library:

# Import the necessary modules from the SURE libraryfromsureimportPreprocessor, reportfromsure.utilityimport (compute_statistical_metrics, compute_mutual_info,
compute_utility_metrics_class,
detection,
query_power)
fromsure.privacyimport (distance_to_closest_record, dcr_stats, number_of_dcr_equal_to_zero, validation_dcr_test, adversary_dataset, membership_inference_test)
# Assuming real_data, valid_data and synth_data are three pandas DataFrames# Preprocessor initialization and query execution on the real, synthetic and validation datasetspreprocessor=Preprocessor(real_data, num_fill_null='forward', scaling='standardize')
real_data_preprocessed=preprocessor.transform(real_data)
valid_data_preprocessed=preprocessor.transform(valid_data)
synth_data_preprocessed=preprocessor.transform(synth_data)
# Statistical properties and mutual informationnum_features_stats, cat_features_stats, temporal_feat_stats=compute_statistical_metrics(real_data, synth_data)
corr_real, corr_synth, corr_difference=compute_mutual_info(real_data_preprocessed, synth_data_preprocessed)
# ML utility: TSTR - Train on Synthetic, Test on RealX_train=real_data_preprocessed.drop("label", axis=1) # Assuming the datasets have a “label” column for the machine learning task they are intended fory_train=real_data_preprocessed["label"]
X_synth=synth_data_preprocessed.drop("label", axis=1)
y_synth=synth_data_preprocessed["label"]
X_test=valid_data_preprocessed.drop("label", axis=1).limit(10000) # Test the trained models on a portion of the original real dataset (first 10k rows)y_test=valid_data_preprocessed["label"].limit(10000)
TSTR_metrics=compute_utility_metrics_class(X_train, X_synth, X_test, y_train, y_synth, y_test)
# Distance to closest recorddcr_synth_train=distance_to_closest_record("synth_train", synth_data, real_data)
dcr_synth_valid=distance_to_closest_record("synth_val", synth_data, valid_data)
dcr_stats_synth_train=dcr_stats("synth_train", dcr_synth_train)
dcr_stats_synth_valid=dcr_stats("synth_val", dcr_synth_valid)
dcr_zero_synth_train=number_of_dcr_equal_to_zero("synth_train", dcr_synth_train)
dcr_zero_synth_valid=number_of_dcr_equal_to_zero("synth_val", dcr_synth_valid)
share=validation_dcr_test(dcr_synth_train, dcr_synth_valid)
# Detection Scoredetection_score=detection(real_data, synth_data, preprocessor)
# Query Powerquery_power_score=query_power(real_data, synth_data, preprocessor)
# ML privacy attack sandbox initialization and simulationadversary_df=adversary_dataset(real_data, valid_data)
# The function adversary_dataset adds a column "privacy_test_is_training" to the adversary dataset, indicating whether the record was part of the training set or notadversary_guesses_ground_truth=adversary_df["privacy_test_is_training"] MIA=membership_inference_test(adversary_df, synth_data, adversary_guesses_ground_truth)
# Report generation as HTML pagereport(real_data, synth_data)

Follow the step-by-step guide to test the library.

About

An open-source Python library for the assessment of utility and privacy performance of any tabular synthetic dataset.

Resources

Stars

23 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Documentation StatusPyPIDownloadsGitHub stars

Synthetic Data: Utility, Regulatory compliance, and Ethical privacy

The SURE package is an open-source Python library intended to be used for the assessment of the utility and privacy performance of any tabular synthetic dataset.

The SURE library features multiple Python modules that can be easily imported and seamlessly integrated into any Python script after installing the library.

Warning

This is a beta version of the library and only runs on Linux and MacOS for the moment.

Important

Requires Python >= 3.10

Installation

To install the library run the following command in your terminal:

$ pip install clearbox-sure

Modules overview

The SURE library features the following modules:

  1. Preprocessor
  2. Statistical similarity metrics
  3. Model garden
  4. ML utility metrics
  5. Distance metrics
  6. Privacy attack sandbox
  7. Report generator

Preprocessor

The input datasets undergo manipulation by the preprocessor module, tailored to conform to the standard structure utilized across the subsequent processes. The Polars library used in the preprocessor makes this operation significantly faster compared to the use of other data processing libraries.

Utility

The statistical similarity metrics, the ML utility metrics and the model garden modules constitute the data utility evaluation part.

The statistical similarity module and the distance metrics module take as input the pre-processed datasets and carry out the operation to assess the statistical similarity between the datasets and how different the content of the synthetic dataset is from the one of the original dataset. In particular, The real and synthetic input datasets are used in the statistical similarity metrics module to assess how close the two datasets are in terms of statistical properties, such as mean, correlation, distribution.

The model garden executes a classification or regression task on the given dataset with multiple machine learning models, returning the performance metrics of each of the models tested on the given task and dataset.

The model garden module’s best performing models are employed in the machine learning utility metrics module to compute the usefulness of the synthetic data on a given ML task (classification or regression).

Privacy

The distance metrics and the privacy attack sandbox make up the synthetic data privacy assessment modules.

The distance metrics module computes the Gower distance between the two input datasets and the distance to the closest record for each line of the first dataset.

The ML privacy attack sandbox allows to simulate a Membership Inference Attack for re-identification of vulnerable records identified with the distance metrics module and evaluate how exposed the synthetic dataset is to this kind of assault.

Report

Eventually, the report generator provides a summary of the utility and privacy metrics computed in the previous modules, providing a visual digest with charts and tables of the results.

This following diagram serves as a visual representation of how each module contributes to the utility-privacy assessment process and highlights the seamless interconnection and synergy between individual blocks.

drawing

Usage

The library leverages Polars, which ensures faster computations compared to other data manipulation libraries. It supports both Polars and Pandas dataframes.

The user must provide both the original real training dataset (which was used to train the generative model that produced the synthetic dataset), the real holdout dataset (which was NOT used to train the generative model that produced the synthetic dataset) and the corresponding synthetic dataset to enable the library's modules to perform the necessary computations for evaluation.

Below is a code snippet example for the usage of the library:

# Import the necessary modules from the SURE libraryfromsureimportPreprocessor, reportfromsure.utilityimport (compute_statistical_metrics, compute_mutual_info,
compute_utility_metrics_class,
detection,
query_power)
fromsure.privacyimport (distance_to_closest_record, dcr_stats, number_of_dcr_equal_to_zero, validation_dcr_test, adversary_dataset, membership_inference_test)
# Assuming real_data, valid_data and synth_data are three pandas DataFrames# Preprocessor initialization and query execution on the real, synthetic and validation datasetspreprocessor=Preprocessor(real_data, num_fill_null='forward', scaling='standardize')
real_data_preprocessed=preprocessor.transform(real_data)
valid_data_preprocessed=preprocessor.transform(valid_data)
synth_data_preprocessed=preprocessor.transform(synth_data)
# Statistical properties and mutual informationnum_features_stats, cat_features_stats, temporal_feat_stats=compute_statistical_metrics(real_data, synth_data)
corr_real, corr_synth, corr_difference=compute_mutual_info(real_data_preprocessed, synth_data_preprocessed)
# ML utility: TSTR - Train on Synthetic, Test on RealX_train=real_data_preprocessed.drop("label", axis=1) # Assuming the datasets have a “label” column for the machine learning task they are intended fory_train=real_data_preprocessed["label"]
X_synth=synth_data_preprocessed.drop("label", axis=1)
y_synth=synth_data_preprocessed["label"]
X_test=valid_data_preprocessed.drop("label", axis=1).limit(10000) # Test the trained models on a portion of the original real dataset (first 10k rows)y_test=valid_data_preprocessed["label"].limit(10000)
TSTR_metrics=compute_utility_metrics_class(X_train, X_synth, X_test, y_train, y_synth, y_test)
# Distance to closest recorddcr_synth_train=distance_to_closest_record("synth_train", synth_data, real_data)
dcr_synth_valid=distance_to_closest_record("synth_val", synth_data, valid_data)
dcr_stats_synth_train=dcr_stats("synth_train", dcr_synth_train)
dcr_stats_synth_valid=dcr_stats("synth_val", dcr_synth_valid)
dcr_zero_synth_train=number_of_dcr_equal_to_zero("synth_train", dcr_synth_train)
dcr_zero_synth_valid=number_of_dcr_equal_to_zero("synth_val", dcr_synth_valid)
share=validation_dcr_test(dcr_synth_train, dcr_synth_valid)
# Detection Scoredetection_score=detection(real_data, synth_data, preprocessor)
# Query Powerquery_power_score=query_power(real_data, synth_data, preprocessor)
# ML privacy attack sandbox initialization and simulationadversary_df=adversary_dataset(real_data, valid_data)
# The function adversary_dataset adds a column "privacy_test_is_training" to the adversary dataset, indicating whether the record was part of the training set or notadversary_guesses_ground_truth=adversary_df["privacy_test_is_training"] MIA=membership_inference_test(adversary_df, synth_data, adversary_guesses_ground_truth)
# Report generation as HTML pagereport(real_data, synth_data)

Follow the step-by-step guide to test the library.

About

An open-source Python library for the assessment of utility and privacy performance of any tabular synthetic dataset.

Resources

Stars

23 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Documentation StatusPyPIDownloadsGitHub stars

Synthetic Data: Utility, Regulatory compliance, and Ethical privacy

The SURE package is an open-source Python library intended to be used for the assessment of the utility and privacy performance of any tabular synthetic dataset.

The SURE library features multiple Python modules that can be easily imported and seamlessly integrated into any Python script after installing the library.

Warning

This is a beta version of the library and only runs on Linux and MacOS for the moment.

Important

Requires Python >= 3.10

Installation

To install the library run the following command in your terminal:

$ pip install clearbox-sure

Modules overview

The SURE library features the following modules:

  1. Preprocessor
  2. Statistical similarity metrics
  3. Model garden
  4. ML utility metrics
  5. Distance metrics
  6. Privacy attack sandbox
  7. Report generator

Preprocessor

The input datasets undergo manipulation by the preprocessor module, tailored to conform to the standard structure utilized across the subsequent processes. The Polars library used in the preprocessor makes this operation significantly faster compared to the use of other data processing libraries.

Utility

The statistical similarity metrics, the ML utility metrics and the model garden modules constitute the data utility evaluation part.

The statistical similarity module and the distance metrics module take as input the pre-processed datasets and carry out the operation to assess the statistical similarity between the datasets and how different the content of the synthetic dataset is from the one of the original dataset. In particular, The real and synthetic input datasets are used in the statistical similarity metrics module to assess how close the two datasets are in terms of statistical properties, such as mean, correlation, distribution.

The model garden executes a classification or regression task on the given dataset with multiple machine learning models, returning the performance metrics of each of the models tested on the given task and dataset.

The model garden module’s best performing models are employed in the machine learning utility metrics module to compute the usefulness of the synthetic data on a given ML task (classification or regression).

Privacy

The distance metrics and the privacy attack sandbox make up the synthetic data privacy assessment modules.

The distance metrics module computes the Gower distance between the two input datasets and the distance to the closest record for each line of the first dataset.

The ML privacy attack sandbox allows to simulate a Membership Inference Attack for re-identification of vulnerable records identified with the distance metrics module and evaluate how exposed the synthetic dataset is to this kind of assault.

Report

Eventually, the report generator provides a summary of the utility and privacy metrics computed in the previous modules, providing a visual digest with charts and tables of the results.

This following diagram serves as a visual representation of how each module contributes to the utility-privacy assessment process and highlights the seamless interconnection and synergy between individual blocks.

drawing

Usage

The library leverages Polars, which ensures faster computations compared to other data manipulation libraries. It supports both Polars and Pandas dataframes.

The user must provide both the original real training dataset (which was used to train the generative model that produced the synthetic dataset), the real holdout dataset (which was NOT used to train the generative model that produced the synthetic dataset) and the corresponding synthetic dataset to enable the library's modules to perform the necessary computations for evaluation.

Below is a code snippet example for the usage of the library:

# Import the necessary modules from the SURE libraryfromsureimportPreprocessor, reportfromsure.utilityimport (compute_statistical_metrics, compute_mutual_info,
compute_utility_metrics_class,
detection,
query_power)
fromsure.privacyimport (distance_to_closest_record, dcr_stats, number_of_dcr_equal_to_zero, validation_dcr_test, adversary_dataset, membership_inference_test)
# Assuming real_data, valid_data and synth_data are three pandas DataFrames# Preprocessor initialization and query execution on the real, synthetic and validation datasetspreprocessor=Preprocessor(real_data, num_fill_null='forward', scaling='standardize')
real_data_preprocessed=preprocessor.transform(real_data)
valid_data_preprocessed=preprocessor.transform(valid_data)
synth_data_preprocessed=preprocessor.transform(synth_data)
# Statistical properties and mutual informationnum_features_stats, cat_features_stats, temporal_feat_stats=compute_statistical_metrics(real_data, synth_data)
corr_real, corr_synth, corr_difference=compute_mutual_info(real_data_preprocessed, synth_data_preprocessed)
# ML utility: TSTR - Train on Synthetic, Test on RealX_train=real_data_preprocessed.drop("label", axis=1) # Assuming the datasets have a “label” column for the machine learning task they are intended fory_train=real_data_preprocessed["label"]
X_synth=synth_data_preprocessed.drop("label", axis=1)
y_synth=synth_data_preprocessed["label"]
X_test=valid_data_preprocessed.drop("label", axis=1).limit(10000) # Test the trained models on a portion of the original real dataset (first 10k rows)y_test=valid_data_preprocessed["label"].limit(10000)
TSTR_metrics=compute_utility_metrics_class(X_train, X_synth, X_test, y_train, y_synth, y_test)
# Distance to closest recorddcr_synth_train=distance_to_closest_record("synth_train", synth_data, real_data)
dcr_synth_valid=distance_to_closest_record("synth_val", synth_data, valid_data)
dcr_stats_synth_train=dcr_stats("synth_train", dcr_synth_train)
dcr_stats_synth_valid=dcr_stats("synth_val", dcr_synth_valid)
dcr_zero_synth_train=number_of_dcr_equal_to_zero("synth_train", dcr_synth_train)
dcr_zero_synth_valid=number_of_dcr_equal_to_zero("synth_val", dcr_synth_valid)
share=validation_dcr_test(dcr_synth_train, dcr_synth_valid)
# Detection Scoredetection_score=detection(real_data, synth_data, preprocessor)
# Query Powerquery_power_score=query_power(real_data, synth_data, preprocessor)
# ML privacy attack sandbox initialization and simulationadversary_df=adversary_dataset(real_data, valid_data)
# The function adversary_dataset adds a column "privacy_test_is_training" to the adversary dataset, indicating whether the record was part of the training set or notadversary_guesses_ground_truth=adversary_df["privacy_test_is_training"] MIA=membership_inference_test(adversary_df, synth_data, adversary_guesses_ground_truth)
# Report generation as HTML pagereport(real_data, synth_data)

Follow the step-by-step guide to test the library.

About

An open-source Python library for the assessment of utility and privacy performance of any tabular synthetic dataset.

Resources

Stars

23 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Documentation StatusPyPIDownloadsGitHub stars

Synthetic Data: Utility, Regulatory compliance, and Ethical privacy

The SURE package is an open-source Python library intended to be used for the assessment of the utility and privacy performance of any tabular synthetic dataset.

The SURE library features multiple Python modules that can be easily imported and seamlessly integrated into any Python script after installing the library.

Warning

This is a beta version of the library and only runs on Linux and MacOS for the moment.

Important

Requires Python >= 3.10

Installation

To install the library run the following command in your terminal:

$ pip install clearbox-sure

Modules overview

The SURE library features the following modules:

  1. Preprocessor
  2. Statistical similarity metrics
  3. Model garden
  4. ML utility metrics
  5. Distance metrics
  6. Privacy attack sandbox
  7. Report generator

Preprocessor

The input datasets undergo manipulation by the preprocessor module, tailored to conform to the standard structure utilized across the subsequent processes. The Polars library used in the preprocessor makes this operation significantly faster compared to the use of other data processing libraries.

Utility

The statistical similarity metrics, the ML utility metrics and the model garden modules constitute the data utility evaluation part.

The statistical similarity module and the distance metrics module take as input the pre-processed datasets and carry out the operation to assess the statistical similarity between the datasets and how different the content of the synthetic dataset is from the one of the original dataset. In particular, The real and synthetic input datasets are used in the statistical similarity metrics module to assess how close the two datasets are in terms of statistical properties, such as mean, correlation, distribution.

The model garden executes a classification or regression task on the given dataset with multiple machine learning models, returning the performance metrics of each of the models tested on the given task and dataset.

The model garden module’s best performing models are employed in the machine learning utility metrics module to compute the usefulness of the synthetic data on a given ML task (classification or regression).

Privacy

The distance metrics and the privacy attack sandbox make up the synthetic data privacy assessment modules.

The distance metrics module computes the Gower distance between the two input datasets and the distance to the closest record for each line of the first dataset.

The ML privacy attack sandbox allows to simulate a Membership Inference Attack for re-identification of vulnerable records identified with the distance metrics module and evaluate how exposed the synthetic dataset is to this kind of assault.

Report

Eventually, the report generator provides a summary of the utility and privacy metrics computed in the previous modules, providing a visual digest with charts and tables of the results.

This following diagram serves as a visual representation of how each module contributes to the utility-privacy assessment process and highlights the seamless interconnection and synergy between individual blocks.

drawing

Usage

The library leverages Polars, which ensures faster computations compared to other data manipulation libraries. It supports both Polars and Pandas dataframes.

The user must provide both the original real training dataset (which was used to train the generative model that produced the synthetic dataset), the real holdout dataset (which was NOT used to train the generative model that produced the synthetic dataset) and the corresponding synthetic dataset to enable the library's modules to perform the necessary computations for evaluation.

Below is a code snippet example for the usage of the library:

# Import the necessary modules from the SURE libraryfromsureimportPreprocessor, reportfromsure.utilityimport (compute_statistical_metrics, compute_mutual_info,
compute_utility_metrics_class,
detection,
query_power)
fromsure.privacyimport (distance_to_closest_record, dcr_stats, number_of_dcr_equal_to_zero, validation_dcr_test, adversary_dataset, membership_inference_test)
# Assuming real_data, valid_data and synth_data are three pandas DataFrames# Preprocessor initialization and query execution on the real, synthetic and validation datasetspreprocessor=Preprocessor(real_data, num_fill_null='forward', scaling='standardize')
real_data_preprocessed=preprocessor.transform(real_data)
valid_data_preprocessed=preprocessor.transform(valid_data)
synth_data_preprocessed=preprocessor.transform(synth_data)
# Statistical properties and mutual informationnum_features_stats, cat_features_stats, temporal_feat_stats=compute_statistical_metrics(real_data, synth_data)
corr_real, corr_synth, corr_difference=compute_mutual_info(real_data_preprocessed, synth_data_preprocessed)
# ML utility: TSTR - Train on Synthetic, Test on RealX_train=real_data_preprocessed.drop("label", axis=1) # Assuming the datasets have a “label” column for the machine learning task they are intended fory_train=real_data_preprocessed["label"]
X_synth=synth_data_preprocessed.drop("label", axis=1)
y_synth=synth_data_preprocessed["label"]
X_test=valid_data_preprocessed.drop("label", axis=1).limit(10000) # Test the trained models on a portion of the original real dataset (first 10k rows)y_test=valid_data_preprocessed["label"].limit(10000)
TSTR_metrics=compute_utility_metrics_class(X_train, X_synth, X_test, y_train, y_synth, y_test)
# Distance to closest recorddcr_synth_train=distance_to_closest_record("synth_train", synth_data, real_data)
dcr_synth_valid=distance_to_closest_record("synth_val", synth_data, valid_data)
dcr_stats_synth_train=dcr_stats("synth_train", dcr_synth_train)
dcr_stats_synth_valid=dcr_stats("synth_val", dcr_synth_valid)
dcr_zero_synth_train=number_of_dcr_equal_to_zero("synth_train", dcr_synth_train)
dcr_zero_synth_valid=number_of_dcr_equal_to_zero("synth_val", dcr_synth_valid)
share=validation_dcr_test(dcr_synth_train, dcr_synth_valid)
# Detection Scoredetection_score=detection(real_data, synth_data, preprocessor)
# Query Powerquery_power_score=query_power(real_data, synth_data, preprocessor)
# ML privacy attack sandbox initialization and simulationadversary_df=adversary_dataset(real_data, valid_data)
# The function adversary_dataset adds a column "privacy_test_is_training" to the adversary dataset, indicating whether the record was part of the training set or notadversary_guesses_ground_truth=adversary_df["privacy_test_is_training"] MIA=membership_inference_test(adversary_df, synth_data, adversary_guesses_ground_truth)
# Report generation as HTML pagereport(real_data, synth_data)

Follow the step-by-step guide to test the library.

About

An open-source Python library for the assessment of utility and privacy performance of any tabular synthetic dataset.

Resources

Stars

23 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Documentation StatusPyPIDownloadsGitHub stars

Synthetic Data: Utility, Regulatory compliance, and Ethical privacy

The SURE package is an open-source Python library intended to be used for the assessment of the utility and privacy performance of any tabular synthetic dataset.

The SURE library features multiple Python modules that can be easily imported and seamlessly integrated into any Python script after installing the library.

Warning

This is a beta version of the library and only runs on Linux and MacOS for the moment.

Important

Requires Python >= 3.10

Installation

To install the library run the following command in your terminal:

$ pip install clearbox-sure

Modules overview

The SURE library features the following modules:

  1. Preprocessor
  2. Statistical similarity metrics
  3. Model garden
  4. ML utility metrics
  5. Distance metrics
  6. Privacy attack sandbox
  7. Report generator

Preprocessor

The input datasets undergo manipulation by the preprocessor module, tailored to conform to the standard structure utilized across the subsequent processes. The Polars library used in the preprocessor makes this operation significantly faster compared to the use of other data processing libraries.

Utility

The statistical similarity metrics, the ML utility metrics and the model garden modules constitute the data utility evaluation part.

The statistical similarity module and the distance metrics module take as input the pre-processed datasets and carry out the operation to assess the statistical similarity between the datasets and how different the content of the synthetic dataset is from the one of the original dataset. In particular, The real and synthetic input datasets are used in the statistical similarity metrics module to assess how close the two datasets are in terms of statistical properties, such as mean, correlation, distribution.

The model garden executes a classification or regression task on the given dataset with multiple machine learning models, returning the performance metrics of each of the models tested on the given task and dataset.

The model garden module’s best performing models are employed in the machine learning utility metrics module to compute the usefulness of the synthetic data on a given ML task (classification or regression).

Privacy

The distance metrics and the privacy attack sandbox make up the synthetic data privacy assessment modules.

The distance metrics module computes the Gower distance between the two input datasets and the distance to the closest record for each line of the first dataset.

The ML privacy attack sandbox allows to simulate a Membership Inference Attack for re-identification of vulnerable records identified with the distance metrics module and evaluate how exposed the synthetic dataset is to this kind of assault.

Report

Eventually, the report generator provides a summary of the utility and privacy metrics computed in the previous modules, providing a visual digest with charts and tables of the results.

This following diagram serves as a visual representation of how each module contributes to the utility-privacy assessment process and highlights the seamless interconnection and synergy between individual blocks.

drawing

Usage

The library leverages Polars, which ensures faster computations compared to other data manipulation libraries. It supports both Polars and Pandas dataframes.

The user must provide both the original real training dataset (which was used to train the generative model that produced the synthetic dataset), the real holdout dataset (which was NOT used to train the generative model that produced the synthetic dataset) and the corresponding synthetic dataset to enable the library's modules to perform the necessary computations for evaluation.

Below is a code snippet example for the usage of the library:

# Import the necessary modules from the SURE libraryfromsureimportPreprocessor, reportfromsure.utilityimport (compute_statistical_metrics, compute_mutual_info,
compute_utility_metrics_class,
detection,
query_power)
fromsure.privacyimport (distance_to_closest_record, dcr_stats, number_of_dcr_equal_to_zero, validation_dcr_test, adversary_dataset, membership_inference_test)
# Assuming real_data, valid_data and synth_data are three pandas DataFrames# Preprocessor initialization and query execution on the real, synthetic and validation datasetspreprocessor=Preprocessor(real_data, num_fill_null='forward', scaling='standardize')
real_data_preprocessed=preprocessor.transform(real_data)
valid_data_preprocessed=preprocessor.transform(valid_data)
synth_data_preprocessed=preprocessor.transform(synth_data)
# Statistical properties and mutual informationnum_features_stats, cat_features_stats, temporal_feat_stats=compute_statistical_metrics(real_data, synth_data)
corr_real, corr_synth, corr_difference=compute_mutual_info(real_data_preprocessed, synth_data_preprocessed)
# ML utility: TSTR - Train on Synthetic, Test on RealX_train=real_data_preprocessed.drop("label", axis=1) # Assuming the datasets have a “label” column for the machine learning task they are intended fory_train=real_data_preprocessed["label"]
X_synth=synth_data_preprocessed.drop("label", axis=1)
y_synth=synth_data_preprocessed["label"]
X_test=valid_data_preprocessed.drop("label", axis=1).limit(10000) # Test the trained models on a portion of the original real dataset (first 10k rows)y_test=valid_data_preprocessed["label"].limit(10000)
TSTR_metrics=compute_utility_metrics_class(X_train, X_synth, X_test, y_train, y_synth, y_test)
# Distance to closest recorddcr_synth_train=distance_to_closest_record("synth_train", synth_data, real_data)
dcr_synth_valid=distance_to_closest_record("synth_val", synth_data, valid_data)
dcr_stats_synth_train=dcr_stats("synth_train", dcr_synth_train)
dcr_stats_synth_valid=dcr_stats("synth_val", dcr_synth_valid)
dcr_zero_synth_train=number_of_dcr_equal_to_zero("synth_train", dcr_synth_train)
dcr_zero_synth_valid=number_of_dcr_equal_to_zero("synth_val", dcr_synth_valid)
share=validation_dcr_test(dcr_synth_train, dcr_synth_valid)
# Detection Scoredetection_score=detection(real_data, synth_data, preprocessor)
# Query Powerquery_power_score=query_power(real_data, synth_data, preprocessor)
# ML privacy attack sandbox initialization and simulationadversary_df=adversary_dataset(real_data, valid_data)
# The function adversary_dataset adds a column "privacy_test_is_training" to the adversary dataset, indicating whether the record was part of the training set or notadversary_guesses_ground_truth=adversary_df["privacy_test_is_training"] MIA=membership_inference_test(adversary_df, synth_data, adversary_guesses_ground_truth)
# Report generation as HTML pagereport(real_data, synth_data)

Follow the step-by-step guide to test the library.

About

An open-source Python library for the assessment of utility and privacy performance of any tabular synthetic dataset.

Resources

Stars

23 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Documentation StatusPyPIDownloadsGitHub stars

Synthetic Data: Utility, Regulatory compliance, and Ethical privacy

The SURE package is an open-source Python library intended to be used for the assessment of the utility and privacy performance of any tabular synthetic dataset.

The SURE library features multiple Python modules that can be easily imported and seamlessly integrated into any Python script after installing the library.

Warning

This is a beta version of the library and only runs on Linux and MacOS for the moment.

Important

Requires Python >= 3.10

Installation

To install the library run the following command in your terminal:

$ pip install clearbox-sure

Modules overview

The SURE library features the following modules:

  1. Preprocessor
  2. Statistical similarity metrics
  3. Model garden
  4. ML utility metrics
  5. Distance metrics
  6. Privacy attack sandbox
  7. Report generator

Preprocessor

The input datasets undergo manipulation by the preprocessor module, tailored to conform to the standard structure utilized across the subsequent processes. The Polars library used in the preprocessor makes this operation significantly faster compared to the use of other data processing libraries.

Utility

The statistical similarity metrics, the ML utility metrics and the model garden modules constitute the data utility evaluation part.

The statistical similarity module and the distance metrics module take as input the pre-processed datasets and carry out the operation to assess the statistical similarity between the datasets and how different the content of the synthetic dataset is from the one of the original dataset. In particular, The real and synthetic input datasets are used in the statistical similarity metrics module to assess how close the two datasets are in terms of statistical properties, such as mean, correlation, distribution.

The model garden executes a classification or regression task on the given dataset with multiple machine learning models, returning the performance metrics of each of the models tested on the given task and dataset.

The model garden module’s best performing models are employed in the machine learning utility metrics module to compute the usefulness of the synthetic data on a given ML task (classification or regression).

Privacy

The distance metrics and the privacy attack sandbox make up the synthetic data privacy assessment modules.

The distance metrics module computes the Gower distance between the two input datasets and the distance to the closest record for each line of the first dataset.

The ML privacy attack sandbox allows to simulate a Membership Inference Attack for re-identification of vulnerable records identified with the distance metrics module and evaluate how exposed the synthetic dataset is to this kind of assault.

Report

Eventually, the report generator provides a summary of the utility and privacy metrics computed in the previous modules, providing a visual digest with charts and tables of the results.

This following diagram serves as a visual representation of how each module contributes to the utility-privacy assessment process and highlights the seamless interconnection and synergy between individual blocks.

drawing

Usage

The library leverages Polars, which ensures faster computations compared to other data manipulation libraries. It supports both Polars and Pandas dataframes.

The user must provide both the original real training dataset (which was used to train the generative model that produced the synthetic dataset), the real holdout dataset (which was NOT used to train the generative model that produced the synthetic dataset) and the corresponding synthetic dataset to enable the library's modules to perform the necessary computations for evaluation.

Below is a code snippet example for the usage of the library:

# Import the necessary modules from the SURE libraryfromsureimportPreprocessor, reportfromsure.utilityimport (compute_statistical_metrics, compute_mutual_info,
compute_utility_metrics_class,
detection,
query_power)
fromsure.privacyimport (distance_to_closest_record, dcr_stats, number_of_dcr_equal_to_zero, validation_dcr_test, adversary_dataset, membership_inference_test)
# Assuming real_data, valid_data and synth_data are three pandas DataFrames# Preprocessor initialization and query execution on the real, synthetic and validation datasetspreprocessor=Preprocessor(real_data, num_fill_null='forward', scaling='standardize')
real_data_preprocessed=preprocessor.transform(real_data)
valid_data_preprocessed=preprocessor.transform(valid_data)
synth_data_preprocessed=preprocessor.transform(synth_data)
# Statistical properties and mutual informationnum_features_stats, cat_features_stats, temporal_feat_stats=compute_statistical_metrics(real_data, synth_data)
corr_real, corr_synth, corr_difference=compute_mutual_info(real_data_preprocessed, synth_data_preprocessed)
# ML utility: TSTR - Train on Synthetic, Test on RealX_train=real_data_preprocessed.drop("label", axis=1) # Assuming the datasets have a “label” column for the machine learning task they are intended fory_train=real_data_preprocessed["label"]
X_synth=synth_data_preprocessed.drop("label", axis=1)
y_synth=synth_data_preprocessed["label"]
X_test=valid_data_preprocessed.drop("label", axis=1).limit(10000) # Test the trained models on a portion of the original real dataset (first 10k rows)y_test=valid_data_preprocessed["label"].limit(10000)
TSTR_metrics=compute_utility_metrics_class(X_train, X_synth, X_test, y_train, y_synth, y_test)
# Distance to closest recorddcr_synth_train=distance_to_closest_record("synth_train", synth_data, real_data)
dcr_synth_valid=distance_to_closest_record("synth_val", synth_data, valid_data)
dcr_stats_synth_train=dcr_stats("synth_train", dcr_synth_train)
dcr_stats_synth_valid=dcr_stats("synth_val", dcr_synth_valid)
dcr_zero_synth_train=number_of_dcr_equal_to_zero("synth_train", dcr_synth_train)
dcr_zero_synth_valid=number_of_dcr_equal_to_zero("synth_val", dcr_synth_valid)
share=validation_dcr_test(dcr_synth_train, dcr_synth_valid)
# Detection Scoredetection_score=detection(real_data, synth_data, preprocessor)
# Query Powerquery_power_score=query_power(real_data, synth_data, preprocessor)
# ML privacy attack sandbox initialization and simulationadversary_df=adversary_dataset(real_data, valid_data)
# The function adversary_dataset adds a column "privacy_test_is_training" to the adversary dataset, indicating whether the record was part of the training set or notadversary_guesses_ground_truth=adversary_df["privacy_test_is_training"] MIA=membership_inference_test(adversary_df, synth_data, adversary_guesses_ground_truth)
# Report generation as HTML pagereport(real_data, synth_data)

Follow the step-by-step guide to test the library.

About

An open-source Python library for the assessment of utility and privacy performance of any tabular synthetic dataset.

Resources

Stars

23 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages