Repository files navigation

Segregation Analysis, Inference, and Decomposition with PySAL

codecovPyPI - Python VersionPyPIConda (channel only)GitHub commits since latest release (branch)DOIDocumentation

The PySAL segregation package is a tool for analyzing patterns of urban segregation. With only a few lines of code, segregation users can

Calculate over 40 segregation measures from simple to state-of-the art, including:

Test whether segregation estimates are statistically significant:

Decompose segregation comparisons into

  • differences arising from spatial structure
  • differences arising from demographic structure

Installation

Released versions of segregation are available on pip and anaconda

pip:

pip install segregation

anaconda:

conda install -c conda-forge segregation

You can also install the current development version from this repository (requires Python >= 3.12). Clone the repository, cd into the directory, and run an editable install:

git clone https://github.com/pysal/segregation.git
cd segregation
pip install -e .

Optionally, create the bundled conda environment first:

conda env create -f environment.yml
conda activate segregation
pip install -e .

Getting started

For a complete guide to the segregation API, see the online documentation.

For code walkthroughs and sample analyses, see the example notebooks

Calculating Segregation Measures

Each index in the segregation module is implemented as a class, which is built from a pandas.DataFrame or a geopandas.GeoDataFrame. To estimate a segregation statistic, a user needs to call the segregation class she wishes to estimate, and pass three arguments:

  • the DataFrame containing population data
  • the name of the column with population counts for the group of interest
  • the name of the column with the total population for each enumeration unit

Every class in segregation has a statistic and a core_data attributes. The first is a direct access to the point estimation of the specific segregation measure and the second attribute gives access to the main data that the module uses internally to perform the estimates.

Single group measures

If, for example, a user was studying income segregation and wanted to know whether high-income residents tend to be more segregated from others. This user may want would want to fit a dissimilarity index (D) to a DataFrame called df to a specific group with columns like "hi_income", "med_income" and "low_income" that store counts of people in each income bracket, and a total column called "total_population". A typical call would be something like this:

fromsegregation.aspatialimportDissimd_index=Dissim(df, "hi_income", "total_population")

To see the estimated D in the first generic example above, the user would have just to run d_index.statistic to see the fitted value.

If a user would want to fit a spatial dissimilarity index (SD), the call would be nearly identical, save for the fact that the DataFrame now needs to be a GeoDataFrame with an appropriate geometry column

fromsegregation.spatialimportSpatialDissimspatial_index=SpatialDissim(gdf, "hi_income", "total_population")

Some spatial indices can also accept either a PySALW object, or a pandarmNetwork object, which allows the user full control over how to parameterize spatial effects. The network functions can be particularly useful for teasing out differences in segregation measures caused by two cities that have two very different spatial structures, like for example Detroit MI (left) and Monroe LA (right):

For point estimation, all single-group indices available are summarized in the following table:

MeasureClass/FunctionSpatial?Specific Arguments
Dissimilarity (D)DissimNo-
Gini (G)GiniSegNo-
Entropy (H)EntropyNo-
Isolation (xPx)IsolationNo-
Exposure (xPy)ExposureNo-
Atkinson (A)AtkinsonNob
Correlation Ratio (V)CorrelationRNo-
Concentration Profile (R)ConProfNom
Modified Dissimilarity (Dct)ModifiedDissimNoiterations
Modified Gini (Gct)ModifiedGiniSegNoiterations
Bias-Corrected Dissimilarity (Dbc)BiasCorrectedDissimNoB
Density-Corrected Dissimilarity (Ddc)DensityCorrectedDissimNoxtol
Minimun-Maximum Index (MM)MinMaxNo
Spatial Proximity Profile (SPP)SpatialProxProfYesm
Spatial Dissimilarity (SD)SpatialDissimYesw, standardize
Boundary Spatial Dissimilarity (BSD)BoundarySpatialDissimYesstandardize
Perimeter Area Ratio Spatial Dissimilarity (PARD)PerimeterAreaRatioSpatialDissimYesstandardize
Distance Decay Isolation (DDxPx)DistanceDecayIsolationYesalpha, beta, metric
Distance Decay Exposure (DDxPy)DistanceDecayExposureYesalpha, beta, metric
Spatial Proximity (SP)SpatialProximityYesalpha, beta, metric
Absolute Clustering (ACL)AbsoluteClusteringYesalpha, beta, metric
Relative Clustering (RCL)RelativeClusteringYesalpha, beta, metric
Delta (DEL)DeltaYes-
Absolute Concentration (ACO)AbsoluteConcentrationYes-
Relative Concentration (RCO)RelativeConcentrationYes-
Absolute Centralization (ACE)AbsoluteCentralizationYes-
Relative Centralization (RCE)RelativeCentralizationYes-
Relative Centralization (RCE)RelativeCentralizationYes-
Spatial Minimun-Maximum (SMM)SpatialMinMaxYesnetwork, w, decay, distance, precompute

Multigroup measures

segregation also facilitates the estimation of multigroup segregation measures.

In this case, the call is nearly identical to the single-group, only now we pass a list of column names rather than a single string; reprising the income segregation example above, an example call might look like this

fromsegregation.aspatialimportMultiDissimindex=MultiDissim(df, ['hi_income', 'med_income', 'low_income'])
index.statistic

Available multi-group indices are summarized in the table below:

MeasureClass/FunctionSpatial?Specific Arguments
Multigroup DissimilarityMultiDissimNo-
Multigroup GiniMultiGiniSegNo-
Multigroup Normalized ExposureMultiNormalizedExposureNo-
Multigroup Information TheoryMultiInformationTheoryNo-
Multigroup Relative DiversityMultiRelativeDiversityNo-
Multigroup Squared Coefficient of VariationMultiSquaredCoefficientVariationNo-
Multigroup DiversityMultiDiversityNonormalized
Simpson’s ConcentrationSimpsonsConcentrationNo-
Simpson’s InteractionSimpsonsInteractionNo-
Multigroup DivergenceMultiDivergenceNo-

Local measures

Also, it is possible to calculate local measures of segregation. A statistics attribute will contain the values of these indexes. Note: in this case the attribute is in the plural since, many statistics are fitted, one for each enumeration unit Local segregation indices have the same signature as their global cousins and are summarized in the table below:

MeasureClass/FunctionSpatial?Specific Arguments
Location QuotientMultiLocationQuotientNo-
Local DiversityMultiLocalDiversityNo-
Local EntropyMultiLocalEntropyNo-
Local Simpson’s ConcentrationMultiLocalSimpsonConcentrationNo-
Local Simpson’s InteractionMultiLocalSimpsonInteractionNo-
Local CentralizationLocalRelativeCentralizationYes-

Testing for Statistical Significance

Once the segregation indexes are fitted, the user can perform inference to shed light for statistical significance in regional analysis. The summary of the inference framework is presented in the table below:

Inference TypeClass/FunctionFunction main InputsFunction Outputs
Single ValueSingleValueTestseg_class, iterations_under_null, null_approach, two_tailedp_value, est_sim, statistic
Two ValuesTwoValueTestseg_class_1, seg_class_2, iterations_under_null, null_approachp_value, est_sim, est_point_diff

Another useful analysis that can be performed with the segregation module is a decompositional approach where two different indexes can be broken down into their spatial component (c_s) and attribute component (c_a). This framework is summarized in the table below:

FrameworkClass/FunctionFunction main InputsFunction Outputs
DecompositionDecomposeSegregationindex1, index2, counterfactual_approachc_a, c_s

In this case, the difference in measured D statistics between Detroit and Monroe is attributable primarily to their demographic makeup, rather than the spatial structure of the two cities. (Note, this is to be expected since D is not a spatial index)

Contributing

PySAL-segregation is under active development and contributors are welcome.

If you have any suggestion, feature request, or bug report, please open a new issue on GitHub. To submit patches, please follow the PySAL development guidelines and open a pull request. Once your changes get merged, you’ll automatically be added to the Contributors List.

Support

If you are having issues, please talk to us in the gitter room.

License

The project is licensed under the BSD license.

Funding

Award #1831615 RIDIR: Scalable Geospatial Analytics for Social Science Research

Renan Xavier Cortes is grateful for the support of Coordenação de Aperfeiçoamento de Pessoal de Nível Superior - Brazil (CAPES) - Process number 88881.170553/2018-01

Citation

To cite segregation, we recommend the following

@software{renan_xavier_cortes_2020,
author = {Renan Xavier Cortes and
eli knaap and
Sergio Rey and
Wei Kang and
Philip Stephens and
James Gaboardi and
Levi John Wolf and
Antti Härkönen and
Dani Arribas-Bel},
title = {PySAL/segregation: Segregation Analysis, Inference, & Decomposition},
month = feb,
year = 2020,
publisher = {Zenodo},
doi = {10.5281/zenodo.3265359},
url = {https://doi.org/10.5281/zenodo.3265359}
}

About

Segregation Measurement, Inferential Statistics, and Decomposition Analysis

Topics

Resources

Stars

121 stars

Watchers

21 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Segregation Analysis, Inference, and Decomposition with PySAL

codecovPyPI - Python VersionPyPIConda (channel only)GitHub commits since latest release (branch)DOIDocumentation

The PySAL segregation package is a tool for analyzing patterns of urban segregation. With only a few lines of code, segregation users can

Calculate over 40 segregation measures from simple to state-of-the art, including:

Test whether segregation estimates are statistically significant:

Decompose segregation comparisons into

  • differences arising from spatial structure
  • differences arising from demographic structure

Installation

Released versions of segregation are available on pip and anaconda

pip:

pip install segregation

anaconda:

conda install -c conda-forge segregation

You can also install the current development version from this repository (requires Python >= 3.12). Clone the repository, cd into the directory, and run an editable install:

git clone https://github.com/pysal/segregation.git
cd segregation
pip install -e .

Optionally, create the bundled conda environment first:

conda env create -f environment.yml
conda activate segregation
pip install -e .

Getting started

For a complete guide to the segregation API, see the online documentation.

For code walkthroughs and sample analyses, see the example notebooks

Calculating Segregation Measures

Each index in the segregation module is implemented as a class, which is built from a pandas.DataFrame or a geopandas.GeoDataFrame. To estimate a segregation statistic, a user needs to call the segregation class she wishes to estimate, and pass three arguments:

  • the DataFrame containing population data
  • the name of the column with population counts for the group of interest
  • the name of the column with the total population for each enumeration unit

Every class in segregation has a statistic and a core_data attributes. The first is a direct access to the point estimation of the specific segregation measure and the second attribute gives access to the main data that the module uses internally to perform the estimates.

Single group measures

If, for example, a user was studying income segregation and wanted to know whether high-income residents tend to be more segregated from others. This user may want would want to fit a dissimilarity index (D) to a DataFrame called df to a specific group with columns like "hi_income", "med_income" and "low_income" that store counts of people in each income bracket, and a total column called "total_population". A typical call would be something like this:

fromsegregation.aspatialimportDissimd_index=Dissim(df, "hi_income", "total_population")

To see the estimated D in the first generic example above, the user would have just to run d_index.statistic to see the fitted value.

If a user would want to fit a spatial dissimilarity index (SD), the call would be nearly identical, save for the fact that the DataFrame now needs to be a GeoDataFrame with an appropriate geometry column

fromsegregation.spatialimportSpatialDissimspatial_index=SpatialDissim(gdf, "hi_income", "total_population")

Some spatial indices can also accept either a PySALW object, or a pandarmNetwork object, which allows the user full control over how to parameterize spatial effects. The network functions can be particularly useful for teasing out differences in segregation measures caused by two cities that have two very different spatial structures, like for example Detroit MI (left) and Monroe LA (right):

For point estimation, all single-group indices available are summarized in the following table:

MeasureClass/FunctionSpatial?Specific Arguments
Dissimilarity (D)DissimNo-
Gini (G)GiniSegNo-
Entropy (H)EntropyNo-
Isolation (xPx)IsolationNo-
Exposure (xPy)ExposureNo-
Atkinson (A)AtkinsonNob
Correlation Ratio (V)CorrelationRNo-
Concentration Profile (R)ConProfNom
Modified Dissimilarity (Dct)ModifiedDissimNoiterations
Modified Gini (Gct)ModifiedGiniSegNoiterations
Bias-Corrected Dissimilarity (Dbc)BiasCorrectedDissimNoB
Density-Corrected Dissimilarity (Ddc)DensityCorrectedDissimNoxtol
Minimun-Maximum Index (MM)MinMaxNo
Spatial Proximity Profile (SPP)SpatialProxProfYesm
Spatial Dissimilarity (SD)SpatialDissimYesw, standardize
Boundary Spatial Dissimilarity (BSD)BoundarySpatialDissimYesstandardize
Perimeter Area Ratio Spatial Dissimilarity (PARD)PerimeterAreaRatioSpatialDissimYesstandardize
Distance Decay Isolation (DDxPx)DistanceDecayIsolationYesalpha, beta, metric
Distance Decay Exposure (DDxPy)DistanceDecayExposureYesalpha, beta, metric
Spatial Proximity (SP)SpatialProximityYesalpha, beta, metric
Absolute Clustering (ACL)AbsoluteClusteringYesalpha, beta, metric
Relative Clustering (RCL)RelativeClusteringYesalpha, beta, metric
Delta (DEL)DeltaYes-
Absolute Concentration (ACO)AbsoluteConcentrationYes-
Relative Concentration (RCO)RelativeConcentrationYes-
Absolute Centralization (ACE)AbsoluteCentralizationYes-
Relative Centralization (RCE)RelativeCentralizationYes-
Relative Centralization (RCE)RelativeCentralizationYes-
Spatial Minimun-Maximum (SMM)SpatialMinMaxYesnetwork, w, decay, distance, precompute

Multigroup measures

segregation also facilitates the estimation of multigroup segregation measures.

In this case, the call is nearly identical to the single-group, only now we pass a list of column names rather than a single string; reprising the income segregation example above, an example call might look like this

fromsegregation.aspatialimportMultiDissimindex=MultiDissim(df, ['hi_income', 'med_income', 'low_income'])
index.statistic

Available multi-group indices are summarized in the table below:

MeasureClass/FunctionSpatial?Specific Arguments
Multigroup DissimilarityMultiDissimNo-
Multigroup GiniMultiGiniSegNo-
Multigroup Normalized ExposureMultiNormalizedExposureNo-
Multigroup Information TheoryMultiInformationTheoryNo-
Multigroup Relative DiversityMultiRelativeDiversityNo-
Multigroup Squared Coefficient of VariationMultiSquaredCoefficientVariationNo-
Multigroup DiversityMultiDiversityNonormalized
Simpson’s ConcentrationSimpsonsConcentrationNo-
Simpson’s InteractionSimpsonsInteractionNo-
Multigroup DivergenceMultiDivergenceNo-

Local measures

Also, it is possible to calculate local measures of segregation. A statistics attribute will contain the values of these indexes. Note: in this case the attribute is in the plural since, many statistics are fitted, one for each enumeration unit Local segregation indices have the same signature as their global cousins and are summarized in the table below:

MeasureClass/FunctionSpatial?Specific Arguments
Location QuotientMultiLocationQuotientNo-
Local DiversityMultiLocalDiversityNo-
Local EntropyMultiLocalEntropyNo-
Local Simpson’s ConcentrationMultiLocalSimpsonConcentrationNo-
Local Simpson’s InteractionMultiLocalSimpsonInteractionNo-
Local CentralizationLocalRelativeCentralizationYes-

Testing for Statistical Significance

Once the segregation indexes are fitted, the user can perform inference to shed light for statistical significance in regional analysis. The summary of the inference framework is presented in the table below:

Inference TypeClass/FunctionFunction main InputsFunction Outputs
Single ValueSingleValueTestseg_class, iterations_under_null, null_approach, two_tailedp_value, est_sim, statistic
Two ValuesTwoValueTestseg_class_1, seg_class_2, iterations_under_null, null_approachp_value, est_sim, est_point_diff

Another useful analysis that can be performed with the segregation module is a decompositional approach where two different indexes can be broken down into their spatial component (c_s) and attribute component (c_a). This framework is summarized in the table below:

FrameworkClass/FunctionFunction main InputsFunction Outputs
DecompositionDecomposeSegregationindex1, index2, counterfactual_approachc_a, c_s

In this case, the difference in measured D statistics between Detroit and Monroe is attributable primarily to their demographic makeup, rather than the spatial structure of the two cities. (Note, this is to be expected since D is not a spatial index)

Contributing

PySAL-segregation is under active development and contributors are welcome.

If you have any suggestion, feature request, or bug report, please open a new issue on GitHub. To submit patches, please follow the PySAL development guidelines and open a pull request. Once your changes get merged, you’ll automatically be added to the Contributors List.

Support

If you are having issues, please talk to us in the gitter room.

License

The project is licensed under the BSD license.

Funding

Award #1831615 RIDIR: Scalable Geospatial Analytics for Social Science Research

Renan Xavier Cortes is grateful for the support of Coordenação de Aperfeiçoamento de Pessoal de Nível Superior - Brazil (CAPES) - Process number 88881.170553/2018-01

Citation

To cite segregation, we recommend the following

@software{renan_xavier_cortes_2020,
author = {Renan Xavier Cortes and
eli knaap and
Sergio Rey and
Wei Kang and
Philip Stephens and
James Gaboardi and
Levi John Wolf and
Antti Härkönen and
Dani Arribas-Bel},
title = {PySAL/segregation: Segregation Analysis, Inference, & Decomposition},
month = feb,
year = 2020,
publisher = {Zenodo},
doi = {10.5281/zenodo.3265359},
url = {https://doi.org/10.5281/zenodo.3265359}
}

About

Segregation Measurement, Inferential Statistics, and Decomposition Analysis

Topics

Resources

Stars

121 stars

Watchers

21 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Segregation Analysis, Inference, and Decomposition with PySAL

codecovPyPI - Python VersionPyPIConda (channel only)GitHub commits since latest release (branch)DOIDocumentation

The PySAL segregation package is a tool for analyzing patterns of urban segregation. With only a few lines of code, segregation users can

Calculate over 40 segregation measures from simple to state-of-the art, including:

Test whether segregation estimates are statistically significant:

Decompose segregation comparisons into

  • differences arising from spatial structure
  • differences arising from demographic structure

Installation

Released versions of segregation are available on pip and anaconda

pip:

pip install segregation

anaconda:

conda install -c conda-forge segregation

You can also install the current development version from this repository (requires Python >= 3.12). Clone the repository, cd into the directory, and run an editable install:

git clone https://github.com/pysal/segregation.git
cd segregation
pip install -e .

Optionally, create the bundled conda environment first:

conda env create -f environment.yml
conda activate segregation
pip install -e .

Getting started

For a complete guide to the segregation API, see the online documentation.

For code walkthroughs and sample analyses, see the example notebooks

Calculating Segregation Measures

Each index in the segregation module is implemented as a class, which is built from a pandas.DataFrame or a geopandas.GeoDataFrame. To estimate a segregation statistic, a user needs to call the segregation class she wishes to estimate, and pass three arguments:

  • the DataFrame containing population data
  • the name of the column with population counts for the group of interest
  • the name of the column with the total population for each enumeration unit

Every class in segregation has a statistic and a core_data attributes. The first is a direct access to the point estimation of the specific segregation measure and the second attribute gives access to the main data that the module uses internally to perform the estimates.

Single group measures

If, for example, a user was studying income segregation and wanted to know whether high-income residents tend to be more segregated from others. This user may want would want to fit a dissimilarity index (D) to a DataFrame called df to a specific group with columns like "hi_income", "med_income" and "low_income" that store counts of people in each income bracket, and a total column called "total_population". A typical call would be something like this:

fromsegregation.aspatialimportDissimd_index=Dissim(df, "hi_income", "total_population")

To see the estimated D in the first generic example above, the user would have just to run d_index.statistic to see the fitted value.

If a user would want to fit a spatial dissimilarity index (SD), the call would be nearly identical, save for the fact that the DataFrame now needs to be a GeoDataFrame with an appropriate geometry column

fromsegregation.spatialimportSpatialDissimspatial_index=SpatialDissim(gdf, "hi_income", "total_population")

Some spatial indices can also accept either a PySALW object, or a pandarmNetwork object, which allows the user full control over how to parameterize spatial effects. The network functions can be particularly useful for teasing out differences in segregation measures caused by two cities that have two very different spatial structures, like for example Detroit MI (left) and Monroe LA (right):

For point estimation, all single-group indices available are summarized in the following table:

MeasureClass/FunctionSpatial?Specific Arguments
Dissimilarity (D)DissimNo-
Gini (G)GiniSegNo-
Entropy (H)EntropyNo-
Isolation (xPx)IsolationNo-
Exposure (xPy)ExposureNo-
Atkinson (A)AtkinsonNob
Correlation Ratio (V)CorrelationRNo-
Concentration Profile (R)ConProfNom
Modified Dissimilarity (Dct)ModifiedDissimNoiterations
Modified Gini (Gct)ModifiedGiniSegNoiterations
Bias-Corrected Dissimilarity (Dbc)BiasCorrectedDissimNoB
Density-Corrected Dissimilarity (Ddc)DensityCorrectedDissimNoxtol
Minimun-Maximum Index (MM)MinMaxNo
Spatial Proximity Profile (SPP)SpatialProxProfYesm
Spatial Dissimilarity (SD)SpatialDissimYesw, standardize
Boundary Spatial Dissimilarity (BSD)BoundarySpatialDissimYesstandardize
Perimeter Area Ratio Spatial Dissimilarity (PARD)PerimeterAreaRatioSpatialDissimYesstandardize
Distance Decay Isolation (DDxPx)DistanceDecayIsolationYesalpha, beta, metric
Distance Decay Exposure (DDxPy)DistanceDecayExposureYesalpha, beta, metric
Spatial Proximity (SP)SpatialProximityYesalpha, beta, metric
Absolute Clustering (ACL)AbsoluteClusteringYesalpha, beta, metric
Relative Clustering (RCL)RelativeClusteringYesalpha, beta, metric
Delta (DEL)DeltaYes-
Absolute Concentration (ACO)AbsoluteConcentrationYes-
Relative Concentration (RCO)RelativeConcentrationYes-
Absolute Centralization (ACE)AbsoluteCentralizationYes-
Relative Centralization (RCE)RelativeCentralizationYes-
Relative Centralization (RCE)RelativeCentralizationYes-
Spatial Minimun-Maximum (SMM)SpatialMinMaxYesnetwork, w, decay, distance, precompute

Multigroup measures

segregation also facilitates the estimation of multigroup segregation measures.

In this case, the call is nearly identical to the single-group, only now we pass a list of column names rather than a single string; reprising the income segregation example above, an example call might look like this

fromsegregation.aspatialimportMultiDissimindex=MultiDissim(df, ['hi_income', 'med_income', 'low_income'])
index.statistic

Available multi-group indices are summarized in the table below:

MeasureClass/FunctionSpatial?Specific Arguments
Multigroup DissimilarityMultiDissimNo-
Multigroup GiniMultiGiniSegNo-
Multigroup Normalized ExposureMultiNormalizedExposureNo-
Multigroup Information TheoryMultiInformationTheoryNo-
Multigroup Relative DiversityMultiRelativeDiversityNo-
Multigroup Squared Coefficient of VariationMultiSquaredCoefficientVariationNo-
Multigroup DiversityMultiDiversityNonormalized
Simpson’s ConcentrationSimpsonsConcentrationNo-
Simpson’s InteractionSimpsonsInteractionNo-
Multigroup DivergenceMultiDivergenceNo-

Local measures

Also, it is possible to calculate local measures of segregation. A statistics attribute will contain the values of these indexes. Note: in this case the attribute is in the plural since, many statistics are fitted, one for each enumeration unit Local segregation indices have the same signature as their global cousins and are summarized in the table below:

MeasureClass/FunctionSpatial?Specific Arguments
Location QuotientMultiLocationQuotientNo-
Local DiversityMultiLocalDiversityNo-
Local EntropyMultiLocalEntropyNo-
Local Simpson’s ConcentrationMultiLocalSimpsonConcentrationNo-
Local Simpson’s InteractionMultiLocalSimpsonInteractionNo-
Local CentralizationLocalRelativeCentralizationYes-

Testing for Statistical Significance

Once the segregation indexes are fitted, the user can perform inference to shed light for statistical significance in regional analysis. The summary of the inference framework is presented in the table below:

Inference TypeClass/FunctionFunction main InputsFunction Outputs
Single ValueSingleValueTestseg_class, iterations_under_null, null_approach, two_tailedp_value, est_sim, statistic
Two ValuesTwoValueTestseg_class_1, seg_class_2, iterations_under_null, null_approachp_value, est_sim, est_point_diff

Another useful analysis that can be performed with the segregation module is a decompositional approach where two different indexes can be broken down into their spatial component (c_s) and attribute component (c_a). This framework is summarized in the table below:

FrameworkClass/FunctionFunction main InputsFunction Outputs
DecompositionDecomposeSegregationindex1, index2, counterfactual_approachc_a, c_s

In this case, the difference in measured D statistics between Detroit and Monroe is attributable primarily to their demographic makeup, rather than the spatial structure of the two cities. (Note, this is to be expected since D is not a spatial index)

Contributing

PySAL-segregation is under active development and contributors are welcome.

If you have any suggestion, feature request, or bug report, please open a new issue on GitHub. To submit patches, please follow the PySAL development guidelines and open a pull request. Once your changes get merged, you’ll automatically be added to the Contributors List.

Support

If you are having issues, please talk to us in the gitter room.

License

The project is licensed under the BSD license.

Funding

Award #1831615 RIDIR: Scalable Geospatial Analytics for Social Science Research

Renan Xavier Cortes is grateful for the support of Coordenação de Aperfeiçoamento de Pessoal de Nível Superior - Brazil (CAPES) - Process number 88881.170553/2018-01

Citation

To cite segregation, we recommend the following

@software{renan_xavier_cortes_2020,
author = {Renan Xavier Cortes and
eli knaap and
Sergio Rey and
Wei Kang and
Philip Stephens and
James Gaboardi and
Levi John Wolf and
Antti Härkönen and
Dani Arribas-Bel},
title = {PySAL/segregation: Segregation Analysis, Inference, & Decomposition},
month = feb,
year = 2020,
publisher = {Zenodo},
doi = {10.5281/zenodo.3265359},
url = {https://doi.org/10.5281/zenodo.3265359}
}

About

Segregation Measurement, Inferential Statistics, and Decomposition Analysis

Topics

Resources

Stars

121 stars

Watchers

21 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Segregation Analysis, Inference, and Decomposition with PySAL

codecovPyPI - Python VersionPyPIConda (channel only)GitHub commits since latest release (branch)DOIDocumentation

The PySAL segregation package is a tool for analyzing patterns of urban segregation. With only a few lines of code, segregation users can

Calculate over 40 segregation measures from simple to state-of-the art, including:

Test whether segregation estimates are statistically significant:

Decompose segregation comparisons into

  • differences arising from spatial structure
  • differences arising from demographic structure

Installation

Released versions of segregation are available on pip and anaconda

pip:

pip install segregation

anaconda:

conda install -c conda-forge segregation

You can also install the current development version from this repository (requires Python >= 3.12). Clone the repository, cd into the directory, and run an editable install:

git clone https://github.com/pysal/segregation.git
cd segregation
pip install -e .

Optionally, create the bundled conda environment first:

conda env create -f environment.yml
conda activate segregation
pip install -e .

Getting started

For a complete guide to the segregation API, see the online documentation.

For code walkthroughs and sample analyses, see the example notebooks

Calculating Segregation Measures

Each index in the segregation module is implemented as a class, which is built from a pandas.DataFrame or a geopandas.GeoDataFrame. To estimate a segregation statistic, a user needs to call the segregation class she wishes to estimate, and pass three arguments:

  • the DataFrame containing population data
  • the name of the column with population counts for the group of interest
  • the name of the column with the total population for each enumeration unit

Every class in segregation has a statistic and a core_data attributes. The first is a direct access to the point estimation of the specific segregation measure and the second attribute gives access to the main data that the module uses internally to perform the estimates.

Single group measures

If, for example, a user was studying income segregation and wanted to know whether high-income residents tend to be more segregated from others. This user may want would want to fit a dissimilarity index (D) to a DataFrame called df to a specific group with columns like "hi_income", "med_income" and "low_income" that store counts of people in each income bracket, and a total column called "total_population". A typical call would be something like this:

fromsegregation.aspatialimportDissimd_index=Dissim(df, "hi_income", "total_population")

To see the estimated D in the first generic example above, the user would have just to run d_index.statistic to see the fitted value.

If a user would want to fit a spatial dissimilarity index (SD), the call would be nearly identical, save for the fact that the DataFrame now needs to be a GeoDataFrame with an appropriate geometry column

fromsegregation.spatialimportSpatialDissimspatial_index=SpatialDissim(gdf, "hi_income", "total_population")

Some spatial indices can also accept either a PySALW object, or a pandarmNetwork object, which allows the user full control over how to parameterize spatial effects. The network functions can be particularly useful for teasing out differences in segregation measures caused by two cities that have two very different spatial structures, like for example Detroit MI (left) and Monroe LA (right):

For point estimation, all single-group indices available are summarized in the following table:

MeasureClass/FunctionSpatial?Specific Arguments
Dissimilarity (D)DissimNo-
Gini (G)GiniSegNo-
Entropy (H)EntropyNo-
Isolation (xPx)IsolationNo-
Exposure (xPy)ExposureNo-
Atkinson (A)AtkinsonNob
Correlation Ratio (V)CorrelationRNo-
Concentration Profile (R)ConProfNom
Modified Dissimilarity (Dct)ModifiedDissimNoiterations
Modified Gini (Gct)ModifiedGiniSegNoiterations
Bias-Corrected Dissimilarity (Dbc)BiasCorrectedDissimNoB
Density-Corrected Dissimilarity (Ddc)DensityCorrectedDissimNoxtol
Minimun-Maximum Index (MM)MinMaxNo
Spatial Proximity Profile (SPP)SpatialProxProfYesm
Spatial Dissimilarity (SD)SpatialDissimYesw, standardize
Boundary Spatial Dissimilarity (BSD)BoundarySpatialDissimYesstandardize
Perimeter Area Ratio Spatial Dissimilarity (PARD)PerimeterAreaRatioSpatialDissimYesstandardize
Distance Decay Isolation (DDxPx)DistanceDecayIsolationYesalpha, beta, metric
Distance Decay Exposure (DDxPy)DistanceDecayExposureYesalpha, beta, metric
Spatial Proximity (SP)SpatialProximityYesalpha, beta, metric
Absolute Clustering (ACL)AbsoluteClusteringYesalpha, beta, metric
Relative Clustering (RCL)RelativeClusteringYesalpha, beta, metric
Delta (DEL)DeltaYes-
Absolute Concentration (ACO)AbsoluteConcentrationYes-
Relative Concentration (RCO)RelativeConcentrationYes-
Absolute Centralization (ACE)AbsoluteCentralizationYes-
Relative Centralization (RCE)RelativeCentralizationYes-
Relative Centralization (RCE)RelativeCentralizationYes-
Spatial Minimun-Maximum (SMM)SpatialMinMaxYesnetwork, w, decay, distance, precompute

Multigroup measures

segregation also facilitates the estimation of multigroup segregation measures.

In this case, the call is nearly identical to the single-group, only now we pass a list of column names rather than a single string; reprising the income segregation example above, an example call might look like this

fromsegregation.aspatialimportMultiDissimindex=MultiDissim(df, ['hi_income', 'med_income', 'low_income'])
index.statistic

Available multi-group indices are summarized in the table below:

MeasureClass/FunctionSpatial?Specific Arguments
Multigroup DissimilarityMultiDissimNo-
Multigroup GiniMultiGiniSegNo-
Multigroup Normalized ExposureMultiNormalizedExposureNo-
Multigroup Information TheoryMultiInformationTheoryNo-
Multigroup Relative DiversityMultiRelativeDiversityNo-
Multigroup Squared Coefficient of VariationMultiSquaredCoefficientVariationNo-
Multigroup DiversityMultiDiversityNonormalized
Simpson’s ConcentrationSimpsonsConcentrationNo-
Simpson’s InteractionSimpsonsInteractionNo-
Multigroup DivergenceMultiDivergenceNo-

Local measures

Also, it is possible to calculate local measures of segregation. A statistics attribute will contain the values of these indexes. Note: in this case the attribute is in the plural since, many statistics are fitted, one for each enumeration unit Local segregation indices have the same signature as their global cousins and are summarized in the table below:

MeasureClass/FunctionSpatial?Specific Arguments
Location QuotientMultiLocationQuotientNo-
Local DiversityMultiLocalDiversityNo-
Local EntropyMultiLocalEntropyNo-
Local Simpson’s ConcentrationMultiLocalSimpsonConcentrationNo-
Local Simpson’s InteractionMultiLocalSimpsonInteractionNo-
Local CentralizationLocalRelativeCentralizationYes-

Testing for Statistical Significance

Once the segregation indexes are fitted, the user can perform inference to shed light for statistical significance in regional analysis. The summary of the inference framework is presented in the table below:

Inference TypeClass/FunctionFunction main InputsFunction Outputs
Single ValueSingleValueTestseg_class, iterations_under_null, null_approach, two_tailedp_value, est_sim, statistic
Two ValuesTwoValueTestseg_class_1, seg_class_2, iterations_under_null, null_approachp_value, est_sim, est_point_diff

Another useful analysis that can be performed with the segregation module is a decompositional approach where two different indexes can be broken down into their spatial component (c_s) and attribute component (c_a). This framework is summarized in the table below:

FrameworkClass/FunctionFunction main InputsFunction Outputs
DecompositionDecomposeSegregationindex1, index2, counterfactual_approachc_a, c_s

In this case, the difference in measured D statistics between Detroit and Monroe is attributable primarily to their demographic makeup, rather than the spatial structure of the two cities. (Note, this is to be expected since D is not a spatial index)

Contributing

PySAL-segregation is under active development and contributors are welcome.

If you have any suggestion, feature request, or bug report, please open a new issue on GitHub. To submit patches, please follow the PySAL development guidelines and open a pull request. Once your changes get merged, you’ll automatically be added to the Contributors List.

Support

If you are having issues, please talk to us in the gitter room.

License

The project is licensed under the BSD license.

Funding

Award #1831615 RIDIR: Scalable Geospatial Analytics for Social Science Research

Renan Xavier Cortes is grateful for the support of Coordenação de Aperfeiçoamento de Pessoal de Nível Superior - Brazil (CAPES) - Process number 88881.170553/2018-01

Citation

To cite segregation, we recommend the following

@software{renan_xavier_cortes_2020,
author = {Renan Xavier Cortes and
eli knaap and
Sergio Rey and
Wei Kang and
Philip Stephens and
James Gaboardi and
Levi John Wolf and
Antti Härkönen and
Dani Arribas-Bel},
title = {PySAL/segregation: Segregation Analysis, Inference, & Decomposition},
month = feb,
year = 2020,
publisher = {Zenodo},
doi = {10.5281/zenodo.3265359},
url = {https://doi.org/10.5281/zenodo.3265359}
}

About

Segregation Measurement, Inferential Statistics, and Decomposition Analysis

Topics

Resources

Stars

121 stars

Watchers

21 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Segregation Analysis, Inference, and Decomposition with PySAL

codecovPyPI - Python VersionPyPIConda (channel only)GitHub commits since latest release (branch)DOIDocumentation

The PySAL segregation package is a tool for analyzing patterns of urban segregation. With only a few lines of code, segregation users can

Calculate over 40 segregation measures from simple to state-of-the art, including:

Test whether segregation estimates are statistically significant:

Decompose segregation comparisons into

  • differences arising from spatial structure
  • differences arising from demographic structure

Installation

Released versions of segregation are available on pip and anaconda

pip:

pip install segregation

anaconda:

conda install -c conda-forge segregation

You can also install the current development version from this repository (requires Python >= 3.12). Clone the repository, cd into the directory, and run an editable install:

git clone https://github.com/pysal/segregation.git
cd segregation
pip install -e .

Optionally, create the bundled conda environment first:

conda env create -f environment.yml
conda activate segregation
pip install -e .

Getting started

For a complete guide to the segregation API, see the online documentation.

For code walkthroughs and sample analyses, see the example notebooks

Calculating Segregation Measures

Each index in the segregation module is implemented as a class, which is built from a pandas.DataFrame or a geopandas.GeoDataFrame. To estimate a segregation statistic, a user needs to call the segregation class she wishes to estimate, and pass three arguments:

  • the DataFrame containing population data
  • the name of the column with population counts for the group of interest
  • the name of the column with the total population for each enumeration unit

Every class in segregation has a statistic and a core_data attributes. The first is a direct access to the point estimation of the specific segregation measure and the second attribute gives access to the main data that the module uses internally to perform the estimates.

Single group measures

If, for example, a user was studying income segregation and wanted to know whether high-income residents tend to be more segregated from others. This user may want would want to fit a dissimilarity index (D) to a DataFrame called df to a specific group with columns like "hi_income", "med_income" and "low_income" that store counts of people in each income bracket, and a total column called "total_population". A typical call would be something like this:

fromsegregation.aspatialimportDissimd_index=Dissim(df, "hi_income", "total_population")

To see the estimated D in the first generic example above, the user would have just to run d_index.statistic to see the fitted value.

If a user would want to fit a spatial dissimilarity index (SD), the call would be nearly identical, save for the fact that the DataFrame now needs to be a GeoDataFrame with an appropriate geometry column

fromsegregation.spatialimportSpatialDissimspatial_index=SpatialDissim(gdf, "hi_income", "total_population")

Some spatial indices can also accept either a PySALW object, or a pandarmNetwork object, which allows the user full control over how to parameterize spatial effects. The network functions can be particularly useful for teasing out differences in segregation measures caused by two cities that have two very different spatial structures, like for example Detroit MI (left) and Monroe LA (right):

For point estimation, all single-group indices available are summarized in the following table:

MeasureClass/FunctionSpatial?Specific Arguments
Dissimilarity (D)DissimNo-
Gini (G)GiniSegNo-
Entropy (H)EntropyNo-
Isolation (xPx)IsolationNo-
Exposure (xPy)ExposureNo-
Atkinson (A)AtkinsonNob
Correlation Ratio (V)CorrelationRNo-
Concentration Profile (R)ConProfNom
Modified Dissimilarity (Dct)ModifiedDissimNoiterations
Modified Gini (Gct)ModifiedGiniSegNoiterations
Bias-Corrected Dissimilarity (Dbc)BiasCorrectedDissimNoB
Density-Corrected Dissimilarity (Ddc)DensityCorrectedDissimNoxtol
Minimun-Maximum Index (MM)MinMaxNo
Spatial Proximity Profile (SPP)SpatialProxProfYesm
Spatial Dissimilarity (SD)SpatialDissimYesw, standardize
Boundary Spatial Dissimilarity (BSD)BoundarySpatialDissimYesstandardize
Perimeter Area Ratio Spatial Dissimilarity (PARD)PerimeterAreaRatioSpatialDissimYesstandardize
Distance Decay Isolation (DDxPx)DistanceDecayIsolationYesalpha, beta, metric
Distance Decay Exposure (DDxPy)DistanceDecayExposureYesalpha, beta, metric
Spatial Proximity (SP)SpatialProximityYesalpha, beta, metric
Absolute Clustering (ACL)AbsoluteClusteringYesalpha, beta, metric
Relative Clustering (RCL)RelativeClusteringYesalpha, beta, metric
Delta (DEL)DeltaYes-
Absolute Concentration (ACO)AbsoluteConcentrationYes-
Relative Concentration (RCO)RelativeConcentrationYes-
Absolute Centralization (ACE)AbsoluteCentralizationYes-
Relative Centralization (RCE)RelativeCentralizationYes-
Relative Centralization (RCE)RelativeCentralizationYes-
Spatial Minimun-Maximum (SMM)SpatialMinMaxYesnetwork, w, decay, distance, precompute

Multigroup measures

segregation also facilitates the estimation of multigroup segregation measures.

In this case, the call is nearly identical to the single-group, only now we pass a list of column names rather than a single string; reprising the income segregation example above, an example call might look like this

fromsegregation.aspatialimportMultiDissimindex=MultiDissim(df, ['hi_income', 'med_income', 'low_income'])
index.statistic

Available multi-group indices are summarized in the table below:

MeasureClass/FunctionSpatial?Specific Arguments
Multigroup DissimilarityMultiDissimNo-
Multigroup GiniMultiGiniSegNo-
Multigroup Normalized ExposureMultiNormalizedExposureNo-
Multigroup Information TheoryMultiInformationTheoryNo-
Multigroup Relative DiversityMultiRelativeDiversityNo-
Multigroup Squared Coefficient of VariationMultiSquaredCoefficientVariationNo-
Multigroup DiversityMultiDiversityNonormalized
Simpson’s ConcentrationSimpsonsConcentrationNo-
Simpson’s InteractionSimpsonsInteractionNo-
Multigroup DivergenceMultiDivergenceNo-

Local measures

Also, it is possible to calculate local measures of segregation. A statistics attribute will contain the values of these indexes. Note: in this case the attribute is in the plural since, many statistics are fitted, one for each enumeration unit Local segregation indices have the same signature as their global cousins and are summarized in the table below:

MeasureClass/FunctionSpatial?Specific Arguments
Location QuotientMultiLocationQuotientNo-
Local DiversityMultiLocalDiversityNo-
Local EntropyMultiLocalEntropyNo-
Local Simpson’s ConcentrationMultiLocalSimpsonConcentrationNo-
Local Simpson’s InteractionMultiLocalSimpsonInteractionNo-
Local CentralizationLocalRelativeCentralizationYes-

Testing for Statistical Significance

Once the segregation indexes are fitted, the user can perform inference to shed light for statistical significance in regional analysis. The summary of the inference framework is presented in the table below:

Inference TypeClass/FunctionFunction main InputsFunction Outputs
Single ValueSingleValueTestseg_class, iterations_under_null, null_approach, two_tailedp_value, est_sim, statistic
Two ValuesTwoValueTestseg_class_1, seg_class_2, iterations_under_null, null_approachp_value, est_sim, est_point_diff

Another useful analysis that can be performed with the segregation module is a decompositional approach where two different indexes can be broken down into their spatial component (c_s) and attribute component (c_a). This framework is summarized in the table below:

FrameworkClass/FunctionFunction main InputsFunction Outputs
DecompositionDecomposeSegregationindex1, index2, counterfactual_approachc_a, c_s

In this case, the difference in measured D statistics between Detroit and Monroe is attributable primarily to their demographic makeup, rather than the spatial structure of the two cities. (Note, this is to be expected since D is not a spatial index)

Contributing

PySAL-segregation is under active development and contributors are welcome.

If you have any suggestion, feature request, or bug report, please open a new issue on GitHub. To submit patches, please follow the PySAL development guidelines and open a pull request. Once your changes get merged, you’ll automatically be added to the Contributors List.

Support

If you are having issues, please talk to us in the gitter room.

License

The project is licensed under the BSD license.

Funding

Award #1831615 RIDIR: Scalable Geospatial Analytics for Social Science Research

Renan Xavier Cortes is grateful for the support of Coordenação de Aperfeiçoamento de Pessoal de Nível Superior - Brazil (CAPES) - Process number 88881.170553/2018-01

Citation

To cite segregation, we recommend the following

@software{renan_xavier_cortes_2020,
author = {Renan Xavier Cortes and
eli knaap and
Sergio Rey and
Wei Kang and
Philip Stephens and
James Gaboardi and
Levi John Wolf and
Antti Härkönen and
Dani Arribas-Bel},
title = {PySAL/segregation: Segregation Analysis, Inference, & Decomposition},
month = feb,
year = 2020,
publisher = {Zenodo},
doi = {10.5281/zenodo.3265359},
url = {https://doi.org/10.5281/zenodo.3265359}
}

About

Segregation Measurement, Inferential Statistics, and Decomposition Analysis

Topics

Resources

Stars

121 stars

Watchers

21 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Segregation Analysis, Inference, and Decomposition with PySAL

codecovPyPI - Python VersionPyPIConda (channel only)GitHub commits since latest release (branch)DOIDocumentation

The PySAL segregation package is a tool for analyzing patterns of urban segregation. With only a few lines of code, segregation users can

Calculate over 40 segregation measures from simple to state-of-the art, including:

Test whether segregation estimates are statistically significant:

Decompose segregation comparisons into

  • differences arising from spatial structure
  • differences arising from demographic structure

Installation

Released versions of segregation are available on pip and anaconda

pip:

pip install segregation

anaconda:

conda install -c conda-forge segregation

You can also install the current development version from this repository (requires Python >= 3.12). Clone the repository, cd into the directory, and run an editable install:

git clone https://github.com/pysal/segregation.git
cd segregation
pip install -e .

Optionally, create the bundled conda environment first:

conda env create -f environment.yml
conda activate segregation
pip install -e .

Getting started

For a complete guide to the segregation API, see the online documentation.

For code walkthroughs and sample analyses, see the example notebooks

Calculating Segregation Measures

Each index in the segregation module is implemented as a class, which is built from a pandas.DataFrame or a geopandas.GeoDataFrame. To estimate a segregation statistic, a user needs to call the segregation class she wishes to estimate, and pass three arguments:

  • the DataFrame containing population data
  • the name of the column with population counts for the group of interest
  • the name of the column with the total population for each enumeration unit

Every class in segregation has a statistic and a core_data attributes. The first is a direct access to the point estimation of the specific segregation measure and the second attribute gives access to the main data that the module uses internally to perform the estimates.

Single group measures

If, for example, a user was studying income segregation and wanted to know whether high-income residents tend to be more segregated from others. This user may want would want to fit a dissimilarity index (D) to a DataFrame called df to a specific group with columns like "hi_income", "med_income" and "low_income" that store counts of people in each income bracket, and a total column called "total_population". A typical call would be something like this:

fromsegregation.aspatialimportDissimd_index=Dissim(df, "hi_income", "total_population")

To see the estimated D in the first generic example above, the user would have just to run d_index.statistic to see the fitted value.

If a user would want to fit a spatial dissimilarity index (SD), the call would be nearly identical, save for the fact that the DataFrame now needs to be a GeoDataFrame with an appropriate geometry column

fromsegregation.spatialimportSpatialDissimspatial_index=SpatialDissim(gdf, "hi_income", "total_population")

Some spatial indices can also accept either a PySALW object, or a pandarmNetwork object, which allows the user full control over how to parameterize spatial effects. The network functions can be particularly useful for teasing out differences in segregation measures caused by two cities that have two very different spatial structures, like for example Detroit MI (left) and Monroe LA (right):

For point estimation, all single-group indices available are summarized in the following table:

MeasureClass/FunctionSpatial?Specific Arguments
Dissimilarity (D)DissimNo-
Gini (G)GiniSegNo-
Entropy (H)EntropyNo-
Isolation (xPx)IsolationNo-
Exposure (xPy)ExposureNo-
Atkinson (A)AtkinsonNob
Correlation Ratio (V)CorrelationRNo-
Concentration Profile (R)ConProfNom
Modified Dissimilarity (Dct)ModifiedDissimNoiterations
Modified Gini (Gct)ModifiedGiniSegNoiterations
Bias-Corrected Dissimilarity (Dbc)BiasCorrectedDissimNoB
Density-Corrected Dissimilarity (Ddc)DensityCorrectedDissimNoxtol
Minimun-Maximum Index (MM)MinMaxNo
Spatial Proximity Profile (SPP)SpatialProxProfYesm
Spatial Dissimilarity (SD)SpatialDissimYesw, standardize
Boundary Spatial Dissimilarity (BSD)BoundarySpatialDissimYesstandardize
Perimeter Area Ratio Spatial Dissimilarity (PARD)PerimeterAreaRatioSpatialDissimYesstandardize
Distance Decay Isolation (DDxPx)DistanceDecayIsolationYesalpha, beta, metric
Distance Decay Exposure (DDxPy)DistanceDecayExposureYesalpha, beta, metric
Spatial Proximity (SP)SpatialProximityYesalpha, beta, metric
Absolute Clustering (ACL)AbsoluteClusteringYesalpha, beta, metric
Relative Clustering (RCL)RelativeClusteringYesalpha, beta, metric
Delta (DEL)DeltaYes-
Absolute Concentration (ACO)AbsoluteConcentrationYes-
Relative Concentration (RCO)RelativeConcentrationYes-
Absolute Centralization (ACE)AbsoluteCentralizationYes-
Relative Centralization (RCE)RelativeCentralizationYes-
Relative Centralization (RCE)RelativeCentralizationYes-
Spatial Minimun-Maximum (SMM)SpatialMinMaxYesnetwork, w, decay, distance, precompute

Multigroup measures

segregation also facilitates the estimation of multigroup segregation measures.

In this case, the call is nearly identical to the single-group, only now we pass a list of column names rather than a single string; reprising the income segregation example above, an example call might look like this

fromsegregation.aspatialimportMultiDissimindex=MultiDissim(df, ['hi_income', 'med_income', 'low_income'])
index.statistic

Available multi-group indices are summarized in the table below:

MeasureClass/FunctionSpatial?Specific Arguments
Multigroup DissimilarityMultiDissimNo-
Multigroup GiniMultiGiniSegNo-
Multigroup Normalized ExposureMultiNormalizedExposureNo-
Multigroup Information TheoryMultiInformationTheoryNo-
Multigroup Relative DiversityMultiRelativeDiversityNo-
Multigroup Squared Coefficient of VariationMultiSquaredCoefficientVariationNo-
Multigroup DiversityMultiDiversityNonormalized
Simpson’s ConcentrationSimpsonsConcentrationNo-
Simpson’s InteractionSimpsonsInteractionNo-
Multigroup DivergenceMultiDivergenceNo-

Local measures

Also, it is possible to calculate local measures of segregation. A statistics attribute will contain the values of these indexes. Note: in this case the attribute is in the plural since, many statistics are fitted, one for each enumeration unit Local segregation indices have the same signature as their global cousins and are summarized in the table below:

MeasureClass/FunctionSpatial?Specific Arguments
Location QuotientMultiLocationQuotientNo-
Local DiversityMultiLocalDiversityNo-
Local EntropyMultiLocalEntropyNo-
Local Simpson’s ConcentrationMultiLocalSimpsonConcentrationNo-
Local Simpson’s InteractionMultiLocalSimpsonInteractionNo-
Local CentralizationLocalRelativeCentralizationYes-

Testing for Statistical Significance

Once the segregation indexes are fitted, the user can perform inference to shed light for statistical significance in regional analysis. The summary of the inference framework is presented in the table below:

Inference TypeClass/FunctionFunction main InputsFunction Outputs
Single ValueSingleValueTestseg_class, iterations_under_null, null_approach, two_tailedp_value, est_sim, statistic
Two ValuesTwoValueTestseg_class_1, seg_class_2, iterations_under_null, null_approachp_value, est_sim, est_point_diff

Another useful analysis that can be performed with the segregation module is a decompositional approach where two different indexes can be broken down into their spatial component (c_s) and attribute component (c_a). This framework is summarized in the table below:

FrameworkClass/FunctionFunction main InputsFunction Outputs
DecompositionDecomposeSegregationindex1, index2, counterfactual_approachc_a, c_s

In this case, the difference in measured D statistics between Detroit and Monroe is attributable primarily to their demographic makeup, rather than the spatial structure of the two cities. (Note, this is to be expected since D is not a spatial index)

Contributing

PySAL-segregation is under active development and contributors are welcome.

If you have any suggestion, feature request, or bug report, please open a new issue on GitHub. To submit patches, please follow the PySAL development guidelines and open a pull request. Once your changes get merged, you’ll automatically be added to the Contributors List.

Support

If you are having issues, please talk to us in the gitter room.

License

The project is licensed under the BSD license.

Funding

Award #1831615 RIDIR: Scalable Geospatial Analytics for Social Science Research

Renan Xavier Cortes is grateful for the support of Coordenação de Aperfeiçoamento de Pessoal de Nível Superior - Brazil (CAPES) - Process number 88881.170553/2018-01

Citation

To cite segregation, we recommend the following

@software{renan_xavier_cortes_2020,
author = {Renan Xavier Cortes and
eli knaap and
Sergio Rey and
Wei Kang and
Philip Stephens and
James Gaboardi and
Levi John Wolf and
Antti Härkönen and
Dani Arribas-Bel},
title = {PySAL/segregation: Segregation Analysis, Inference, & Decomposition},
month = feb,
year = 2020,
publisher = {Zenodo},
doi = {10.5281/zenodo.3265359},
url = {https://doi.org/10.5281/zenodo.3265359}
}

About

Segregation Measurement, Inferential Statistics, and Decomposition Analysis

Topics

Resources

Stars

121 stars

Watchers

21 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Segregation Analysis, Inference, and Decomposition with PySAL

codecovPyPI - Python VersionPyPIConda (channel only)GitHub commits since latest release (branch)DOIDocumentation

The PySAL segregation package is a tool for analyzing patterns of urban segregation. With only a few lines of code, segregation users can

Calculate over 40 segregation measures from simple to state-of-the art, including:

Test whether segregation estimates are statistically significant:

Decompose segregation comparisons into

  • differences arising from spatial structure
  • differences arising from demographic structure

Installation

Released versions of segregation are available on pip and anaconda

pip:

pip install segregation

anaconda:

conda install -c conda-forge segregation

You can also install the current development version from this repository (requires Python >= 3.12). Clone the repository, cd into the directory, and run an editable install:

git clone https://github.com/pysal/segregation.git
cd segregation
pip install -e .

Optionally, create the bundled conda environment first:

conda env create -f environment.yml
conda activate segregation
pip install -e .

Getting started

For a complete guide to the segregation API, see the online documentation.

For code walkthroughs and sample analyses, see the example notebooks

Calculating Segregation Measures

Each index in the segregation module is implemented as a class, which is built from a pandas.DataFrame or a geopandas.GeoDataFrame. To estimate a segregation statistic, a user needs to call the segregation class she wishes to estimate, and pass three arguments:

  • the DataFrame containing population data
  • the name of the column with population counts for the group of interest
  • the name of the column with the total population for each enumeration unit

Every class in segregation has a statistic and a core_data attributes. The first is a direct access to the point estimation of the specific segregation measure and the second attribute gives access to the main data that the module uses internally to perform the estimates.

Single group measures

If, for example, a user was studying income segregation and wanted to know whether high-income residents tend to be more segregated from others. This user may want would want to fit a dissimilarity index (D) to a DataFrame called df to a specific group with columns like "hi_income", "med_income" and "low_income" that store counts of people in each income bracket, and a total column called "total_population". A typical call would be something like this:

fromsegregation.aspatialimportDissimd_index=Dissim(df, "hi_income", "total_population")

To see the estimated D in the first generic example above, the user would have just to run d_index.statistic to see the fitted value.

If a user would want to fit a spatial dissimilarity index (SD), the call would be nearly identical, save for the fact that the DataFrame now needs to be a GeoDataFrame with an appropriate geometry column

fromsegregation.spatialimportSpatialDissimspatial_index=SpatialDissim(gdf, "hi_income", "total_population")

Some spatial indices can also accept either a PySALW object, or a pandarmNetwork object, which allows the user full control over how to parameterize spatial effects. The network functions can be particularly useful for teasing out differences in segregation measures caused by two cities that have two very different spatial structures, like for example Detroit MI (left) and Monroe LA (right):

For point estimation, all single-group indices available are summarized in the following table:

MeasureClass/FunctionSpatial?Specific Arguments
Dissimilarity (D)DissimNo-
Gini (G)GiniSegNo-
Entropy (H)EntropyNo-
Isolation (xPx)IsolationNo-
Exposure (xPy)ExposureNo-
Atkinson (A)AtkinsonNob
Correlation Ratio (V)CorrelationRNo-
Concentration Profile (R)ConProfNom
Modified Dissimilarity (Dct)ModifiedDissimNoiterations
Modified Gini (Gct)ModifiedGiniSegNoiterations
Bias-Corrected Dissimilarity (Dbc)BiasCorrectedDissimNoB
Density-Corrected Dissimilarity (Ddc)DensityCorrectedDissimNoxtol
Minimun-Maximum Index (MM)MinMaxNo
Spatial Proximity Profile (SPP)SpatialProxProfYesm
Spatial Dissimilarity (SD)SpatialDissimYesw, standardize
Boundary Spatial Dissimilarity (BSD)BoundarySpatialDissimYesstandardize
Perimeter Area Ratio Spatial Dissimilarity (PARD)PerimeterAreaRatioSpatialDissimYesstandardize
Distance Decay Isolation (DDxPx)DistanceDecayIsolationYesalpha, beta, metric
Distance Decay Exposure (DDxPy)DistanceDecayExposureYesalpha, beta, metric
Spatial Proximity (SP)SpatialProximityYesalpha, beta, metric
Absolute Clustering (ACL)AbsoluteClusteringYesalpha, beta, metric
Relative Clustering (RCL)RelativeClusteringYesalpha, beta, metric
Delta (DEL)DeltaYes-
Absolute Concentration (ACO)AbsoluteConcentrationYes-
Relative Concentration (RCO)RelativeConcentrationYes-
Absolute Centralization (ACE)AbsoluteCentralizationYes-
Relative Centralization (RCE)RelativeCentralizationYes-
Relative Centralization (RCE)RelativeCentralizationYes-
Spatial Minimun-Maximum (SMM)SpatialMinMaxYesnetwork, w, decay, distance, precompute

Multigroup measures

segregation also facilitates the estimation of multigroup segregation measures.

In this case, the call is nearly identical to the single-group, only now we pass a list of column names rather than a single string; reprising the income segregation example above, an example call might look like this

fromsegregation.aspatialimportMultiDissimindex=MultiDissim(df, ['hi_income', 'med_income', 'low_income'])
index.statistic

Available multi-group indices are summarized in the table below:

MeasureClass/FunctionSpatial?Specific Arguments
Multigroup DissimilarityMultiDissimNo-
Multigroup GiniMultiGiniSegNo-
Multigroup Normalized ExposureMultiNormalizedExposureNo-
Multigroup Information TheoryMultiInformationTheoryNo-
Multigroup Relative DiversityMultiRelativeDiversityNo-
Multigroup Squared Coefficient of VariationMultiSquaredCoefficientVariationNo-
Multigroup DiversityMultiDiversityNonormalized
Simpson’s ConcentrationSimpsonsConcentrationNo-
Simpson’s InteractionSimpsonsInteractionNo-
Multigroup DivergenceMultiDivergenceNo-

Local measures

Also, it is possible to calculate local measures of segregation. A statistics attribute will contain the values of these indexes. Note: in this case the attribute is in the plural since, many statistics are fitted, one for each enumeration unit Local segregation indices have the same signature as their global cousins and are summarized in the table below:

MeasureClass/FunctionSpatial?Specific Arguments
Location QuotientMultiLocationQuotientNo-
Local DiversityMultiLocalDiversityNo-
Local EntropyMultiLocalEntropyNo-
Local Simpson’s ConcentrationMultiLocalSimpsonConcentrationNo-
Local Simpson’s InteractionMultiLocalSimpsonInteractionNo-
Local CentralizationLocalRelativeCentralizationYes-

Testing for Statistical Significance

Once the segregation indexes are fitted, the user can perform inference to shed light for statistical significance in regional analysis. The summary of the inference framework is presented in the table below:

Inference TypeClass/FunctionFunction main InputsFunction Outputs
Single ValueSingleValueTestseg_class, iterations_under_null, null_approach, two_tailedp_value, est_sim, statistic
Two ValuesTwoValueTestseg_class_1, seg_class_2, iterations_under_null, null_approachp_value, est_sim, est_point_diff

Another useful analysis that can be performed with the segregation module is a decompositional approach where two different indexes can be broken down into their spatial component (c_s) and attribute component (c_a). This framework is summarized in the table below:

FrameworkClass/FunctionFunction main InputsFunction Outputs
DecompositionDecomposeSegregationindex1, index2, counterfactual_approachc_a, c_s

In this case, the difference in measured D statistics between Detroit and Monroe is attributable primarily to their demographic makeup, rather than the spatial structure of the two cities. (Note, this is to be expected since D is not a spatial index)

Contributing

PySAL-segregation is under active development and contributors are welcome.

If you have any suggestion, feature request, or bug report, please open a new issue on GitHub. To submit patches, please follow the PySAL development guidelines and open a pull request. Once your changes get merged, you’ll automatically be added to the Contributors List.

Support

If you are having issues, please talk to us in the gitter room.

License

The project is licensed under the BSD license.

Funding

Award #1831615 RIDIR: Scalable Geospatial Analytics for Social Science Research

Renan Xavier Cortes is grateful for the support of Coordenação de Aperfeiçoamento de Pessoal de Nível Superior - Brazil (CAPES) - Process number 88881.170553/2018-01

Citation

To cite segregation, we recommend the following

@software{renan_xavier_cortes_2020,
author = {Renan Xavier Cortes and
eli knaap and
Sergio Rey and
Wei Kang and
Philip Stephens and
James Gaboardi and
Levi John Wolf and
Antti Härkönen and
Dani Arribas-Bel},
title = {PySAL/segregation: Segregation Analysis, Inference, & Decomposition},
month = feb,
year = 2020,
publisher = {Zenodo},
doi = {10.5281/zenodo.3265359},
url = {https://doi.org/10.5281/zenodo.3265359}
}

About

Segregation Measurement, Inferential Statistics, and Decomposition Analysis

Topics

Resources

Stars

121 stars

Watchers

21 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Segregation Analysis, Inference, and Decomposition with PySAL

codecovPyPI - Python VersionPyPIConda (channel only)GitHub commits since latest release (branch)DOIDocumentation

The PySAL segregation package is a tool for analyzing patterns of urban segregation. With only a few lines of code, segregation users can

Calculate over 40 segregation measures from simple to state-of-the art, including:

Test whether segregation estimates are statistically significant:

Decompose segregation comparisons into

  • differences arising from spatial structure
  • differences arising from demographic structure

Installation

Released versions of segregation are available on pip and anaconda

pip:

pip install segregation

anaconda:

conda install -c conda-forge segregation

You can also install the current development version from this repository (requires Python >= 3.12). Clone the repository, cd into the directory, and run an editable install:

git clone https://github.com/pysal/segregation.git
cd segregation
pip install -e .

Optionally, create the bundled conda environment first:

conda env create -f environment.yml
conda activate segregation
pip install -e .

Getting started

For a complete guide to the segregation API, see the online documentation.

For code walkthroughs and sample analyses, see the example notebooks

Calculating Segregation Measures

Each index in the segregation module is implemented as a class, which is built from a pandas.DataFrame or a geopandas.GeoDataFrame. To estimate a segregation statistic, a user needs to call the segregation class she wishes to estimate, and pass three arguments:

  • the DataFrame containing population data
  • the name of the column with population counts for the group of interest
  • the name of the column with the total population for each enumeration unit

Every class in segregation has a statistic and a core_data attributes. The first is a direct access to the point estimation of the specific segregation measure and the second attribute gives access to the main data that the module uses internally to perform the estimates.

Single group measures

If, for example, a user was studying income segregation and wanted to know whether high-income residents tend to be more segregated from others. This user may want would want to fit a dissimilarity index (D) to a DataFrame called df to a specific group with columns like "hi_income", "med_income" and "low_income" that store counts of people in each income bracket, and a total column called "total_population". A typical call would be something like this:

fromsegregation.aspatialimportDissimd_index=Dissim(df, "hi_income", "total_population")

To see the estimated D in the first generic example above, the user would have just to run d_index.statistic to see the fitted value.

If a user would want to fit a spatial dissimilarity index (SD), the call would be nearly identical, save for the fact that the DataFrame now needs to be a GeoDataFrame with an appropriate geometry column

fromsegregation.spatialimportSpatialDissimspatial_index=SpatialDissim(gdf, "hi_income", "total_population")

Some spatial indices can also accept either a PySALW object, or a pandarmNetwork object, which allows the user full control over how to parameterize spatial effects. The network functions can be particularly useful for teasing out differences in segregation measures caused by two cities that have two very different spatial structures, like for example Detroit MI (left) and Monroe LA (right):

For point estimation, all single-group indices available are summarized in the following table:

MeasureClass/FunctionSpatial?Specific Arguments
Dissimilarity (D)DissimNo-
Gini (G)GiniSegNo-
Entropy (H)EntropyNo-
Isolation (xPx)IsolationNo-
Exposure (xPy)ExposureNo-
Atkinson (A)AtkinsonNob
Correlation Ratio (V)CorrelationRNo-
Concentration Profile (R)ConProfNom
Modified Dissimilarity (Dct)ModifiedDissimNoiterations
Modified Gini (Gct)ModifiedGiniSegNoiterations
Bias-Corrected Dissimilarity (Dbc)BiasCorrectedDissimNoB
Density-Corrected Dissimilarity (Ddc)DensityCorrectedDissimNoxtol
Minimun-Maximum Index (MM)MinMaxNo
Spatial Proximity Profile (SPP)SpatialProxProfYesm
Spatial Dissimilarity (SD)SpatialDissimYesw, standardize
Boundary Spatial Dissimilarity (BSD)BoundarySpatialDissimYesstandardize
Perimeter Area Ratio Spatial Dissimilarity (PARD)PerimeterAreaRatioSpatialDissimYesstandardize
Distance Decay Isolation (DDxPx)DistanceDecayIsolationYesalpha, beta, metric
Distance Decay Exposure (DDxPy)DistanceDecayExposureYesalpha, beta, metric
Spatial Proximity (SP)SpatialProximityYesalpha, beta, metric
Absolute Clustering (ACL)AbsoluteClusteringYesalpha, beta, metric
Relative Clustering (RCL)RelativeClusteringYesalpha, beta, metric
Delta (DEL)DeltaYes-
Absolute Concentration (ACO)AbsoluteConcentrationYes-
Relative Concentration (RCO)RelativeConcentrationYes-
Absolute Centralization (ACE)AbsoluteCentralizationYes-
Relative Centralization (RCE)RelativeCentralizationYes-
Relative Centralization (RCE)RelativeCentralizationYes-
Spatial Minimun-Maximum (SMM)SpatialMinMaxYesnetwork, w, decay, distance, precompute

Multigroup measures

segregation also facilitates the estimation of multigroup segregation measures.

In this case, the call is nearly identical to the single-group, only now we pass a list of column names rather than a single string; reprising the income segregation example above, an example call might look like this

fromsegregation.aspatialimportMultiDissimindex=MultiDissim(df, ['hi_income', 'med_income', 'low_income'])
index.statistic

Available multi-group indices are summarized in the table below:

MeasureClass/FunctionSpatial?Specific Arguments
Multigroup DissimilarityMultiDissimNo-
Multigroup GiniMultiGiniSegNo-
Multigroup Normalized ExposureMultiNormalizedExposureNo-
Multigroup Information TheoryMultiInformationTheoryNo-
Multigroup Relative DiversityMultiRelativeDiversityNo-
Multigroup Squared Coefficient of VariationMultiSquaredCoefficientVariationNo-
Multigroup DiversityMultiDiversityNonormalized
Simpson’s ConcentrationSimpsonsConcentrationNo-
Simpson’s InteractionSimpsonsInteractionNo-
Multigroup DivergenceMultiDivergenceNo-

Local measures

Also, it is possible to calculate local measures of segregation. A statistics attribute will contain the values of these indexes. Note: in this case the attribute is in the plural since, many statistics are fitted, one for each enumeration unit Local segregation indices have the same signature as their global cousins and are summarized in the table below:

MeasureClass/FunctionSpatial?Specific Arguments
Location QuotientMultiLocationQuotientNo-
Local DiversityMultiLocalDiversityNo-
Local EntropyMultiLocalEntropyNo-
Local Simpson’s ConcentrationMultiLocalSimpsonConcentrationNo-
Local Simpson’s InteractionMultiLocalSimpsonInteractionNo-
Local CentralizationLocalRelativeCentralizationYes-

Testing for Statistical Significance

Once the segregation indexes are fitted, the user can perform inference to shed light for statistical significance in regional analysis. The summary of the inference framework is presented in the table below:

Inference TypeClass/FunctionFunction main InputsFunction Outputs
Single ValueSingleValueTestseg_class, iterations_under_null, null_approach, two_tailedp_value, est_sim, statistic
Two ValuesTwoValueTestseg_class_1, seg_class_2, iterations_under_null, null_approachp_value, est_sim, est_point_diff

Another useful analysis that can be performed with the segregation module is a decompositional approach where two different indexes can be broken down into their spatial component (c_s) and attribute component (c_a). This framework is summarized in the table below:

FrameworkClass/FunctionFunction main InputsFunction Outputs
DecompositionDecomposeSegregationindex1, index2, counterfactual_approachc_a, c_s

In this case, the difference in measured D statistics between Detroit and Monroe is attributable primarily to their demographic makeup, rather than the spatial structure of the two cities. (Note, this is to be expected since D is not a spatial index)

Contributing

PySAL-segregation is under active development and contributors are welcome.

If you have any suggestion, feature request, or bug report, please open a new issue on GitHub. To submit patches, please follow the PySAL development guidelines and open a pull request. Once your changes get merged, you’ll automatically be added to the Contributors List.

Support

If you are having issues, please talk to us in the gitter room.

License

The project is licensed under the BSD license.

Funding

Award #1831615 RIDIR: Scalable Geospatial Analytics for Social Science Research

Renan Xavier Cortes is grateful for the support of Coordenação de Aperfeiçoamento de Pessoal de Nível Superior - Brazil (CAPES) - Process number 88881.170553/2018-01

Citation

To cite segregation, we recommend the following

@software{renan_xavier_cortes_2020,
author = {Renan Xavier Cortes and
eli knaap and
Sergio Rey and
Wei Kang and
Philip Stephens and
James Gaboardi and
Levi John Wolf and
Antti Härkönen and
Dani Arribas-Bel},
title = {PySAL/segregation: Segregation Analysis, Inference, & Decomposition},
month = feb,
year = 2020,
publisher = {Zenodo},
doi = {10.5281/zenodo.3265359},
url = {https://doi.org/10.5281/zenodo.3265359}
}

About

Segregation Measurement, Inferential Statistics, and Decomposition Analysis

Topics

Resources

Stars

121 stars

Watchers

21 watching

Forks

Releases

Packages

Used by

Contributors

Languages