Skip to content

Repository files navigation

GeoLinear

Boosted piecewise-linear models on cooperative geometry partitions.

PyPI versionLicense: AGPL-3.0Python 3.10+


What is GeoLinear?

GeoLinear discovers cooperative geometry regimes in your data — groups of observations where features interact in similar ways — and fits an interpretable linear model inside each regime. Predictions are accumulated across boosting rounds.

The key insight is that many real-world relationships are piecewise-linear in cooperative geometry. A global linear model averages over regimes and loses the regime-specific signal. GeoLinear finds the regimes automatically (via HVRT) and lets the coefficients vary across them.

Round 1: HVRT partitions X → {cooperative, competitive, mixed} regimes
Ridge fits within each partition on residuals
Round 2: New partitioning on updated residuals → new local Ridge models
...
Final prediction = Σ (learning_rate × stage_predictions) + intercept

Each partition's Ridge coefficients are directly interpretable as local relativities — exactly the quantities actuaries file with regulators.


Why it matters for actuaries

Insurance pricing models must be both accurate and interpretable. Regulators require filed relativities; black-box models are inadmissible. The usual compromise — vanilla GLM — leaves accuracy on the table when the true relationship is regime-switching.

GeoLinear bridges this gap:

  1. Fit GeoLinear on the training portfolio. Each partition's Ridge coefficients are the relativities for that cooperative geometry segment.
  2. File a meta-GLM: fit OLS(X → ŷ_GL) to compress the ensemble into a single interpretable GLM. The compression loss is typically small (R² drop < 2%).
  3. Audit trail: every prediction can be traced to a specific partition and its local coefficient vector.

See examples/insurance_pricing_demo.py for a worked actuarial example including relativities tables and the GLM bridge.


Installation

pip install geolinear

Requires a C++17 compiler and CMake (automatically satisfied on most systems; the cmake PyPI package is a reliable fallback).

For the examples:

pip install geolinear[examples] # adds xgboost, optuna, matplotlib

Quick start

Regression

fromgeolinearimportGeoLinearfromsklearn.datasetsimportfetch_california_housingfromsklearn.model_selectionimporttrain_test_splitfromsklearn.metricsimportr2_scoreX, y=fetch_california_housing(return_X_y=True)
X_tr, X_te, y_tr, y_te=train_test_split(X, y, test_size=0.2, random_state=0)
model=GeoLinear(n_rounds=20, learning_rate=0.1, alpha=1.0)
model.fit(X_tr, y_tr)
print(r2_score(y_te, model.predict(X_te))) # ~0.72 default params

Classification

fromgeolinearimportGeoLinearClassifierfromsklearn.datasetsimportload_breast_cancerfromsklearn.model_selectionimporttrain_test_splitfromsklearn.metricsimportroc_auc_scoreX, y=load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te=train_test_split(X, y, test_size=0.2, random_state=0)
clf=GeoLinearClassifier(n_rounds=20, alpha=1.0)
clf.fit(X_tr, y_tr)
print(roc_auc_score(y_te, clf.predict_proba(X_te)[:, 1])) # ~0.993

Pipeline + GridSearchCV

fromsklearn.pipelineimportPipelinefromsklearn.preprocessingimportStandardScalerfromsklearn.model_selectionimportGridSearchCVfromgeolinearimportGeoLinearpipe=Pipeline([("scaler", StandardScaler()), ("gl", GeoLinear())])
grid=GridSearchCV(pipe, {"gl__alpha": [0.1, 1.0, 10.0], "gl__n_rounds": [10, 20]}, cv=5)
grid.fit(X_tr, y_tr)

Optuna HPO

importoptunafromgeolinearimportGeoLinearfromsklearn.model_selectionimportcross_val_scoredefobjective(trial):
model=GeoLinear(
n_rounds=trial.suggest_int("n_rounds", 10, 100),
learning_rate=trial.suggest_float("learning_rate", 0.01, 0.3, log=True),
alpha=trial.suggest_float("alpha", 0.01, 20.0, log=True),
y_weight=trial.suggest_float("y_weight", 0.1, 1.0),
base_learner=trial.suggest_categorical("base_learner", ["ridge", "lasso"]),
hvrt_n_partitions=trial.suggest_int("hvrt_n_partitions", 3, 14),
min_samples_partition=trial.suggest_int("min_samples_partition", 3, 25),
hvrt_model=trial.suggest_categorical("hvrt_model", ["hvrt", "pyramid_hart", "fast_hvrt"]),
use_t_feature=trial.suggest_categorical("use_t_feature", [False, True]),
refit_interval=trial.suggest_categorical("refit_interval", [0, 1, 2]),
)
returncross_val_score(model, X_tr, y_tr, cv=5, scoring="r2").mean()
study=optuna.create_study(direction="maximize")
study.optimize(objective, n_trials=60)

Ecosystem: which tool to use?

GeoLinear is part of a family of cooperative-geometry libraries. The right choice depends on what "interpretable" means in your domain and the structure of your data.

DomainPrimary needRecommended
Insurance pricingFiled relativities; regulator-auditable linear tariffGeoLinear
Credit risk / scoringScorecard linearity; SR 11-7 model risk complianceGeoLinear
Utility rate-settingFiled tariff schedules with linear rate factorsGeoLinear
Healthcare / clinicalXGBoost-level performance with fully deterministic, auditable predictionsGeoXGB
Ecology / environmentalNonlinear regime detection; richer explanation than SHAP via 100% determinismGeoXGB
Public policyAlgorithmic accountability without linearity constraintsGeoXGB
Personalised interventionsPer-entity longitudinal data; individual treatment trajectoriesAutoITE

Rule of thumb: if your regulator or ethics board requires a linear equation you can file or defend, use GeoLinear. If you need XGBoost-class accuracy with fully deterministic predictions that provide richer interpretability than SHAP, use GeoXGB. If you have repeated observations per individual and want to model how treatment effects evolve over time for each entity, use AutoITE.


Benchmark results

80-trial Optuna HPO, GeoLinear (v0.3.0) vs XGBoost. use_t_feature included in GeoLinear's HPO search space so the optimizer can exploit cooperative geometry when it exists.

Regression R² — v0.3.0 (80-trial HPO)

DGPGL-HPOuse_tXGB-HPOGap
T-regime0.5690.296+0.273 GL wins
3-regime0.3580.329+0.029 GL wins
Friedman10.9840.988−0.003 (tie)
Linear0.9650.957+0.008 GL wins
CalHousing0.5840.856−0.271 XGB wins

Score: 3 wins · 1 tie · 1 loss.

use_t_feature is selected by HPO whenever the data has cooperative structure (T-regime, 3-regime). On smooth or axis-aligned DGPs (Friedman1, Linear) HPO correctly opts out. CalHousing remains XGBoost territory: spatial autocorrelation and threshold effects are better captured by axis-aligned splits than by pairwise cooperative geometry.

Classification AUC

DatasetGL-Ridge (default)GL-HPOXGB-HPO
Synthetic0.9400.9570.981
BreastCancer0.9930.9720.993

Default-parameter GL-Ridge matches XGBoost-HPO on BreastCancer (AUC 0.993 each).


API reference

GeoLinear (regressor)

ParameterTypeDefaultDescription
n_roundsint20Boosting rounds
learning_ratefloat0.1Shrinkage per round
y_weightfloat0.5HVRT outcome-blend (0 = geometry-only, 1 = y-driven)
base_learnerstr"ridge""ridge", "ols", or "lasso"
alphafloat1.0L2 regularisation within each partition
min_samples_partitionint5Minimum samples to fit a partition model
hvrt_n_partitionsint|NoneNoneTarget partitions (None = HVRT auto-tune)
hvrt_min_samples_leafint|NoneNoneHVRT min leaf size
hvrt_inner_roundsint1HVRT T-residual inner rounds per stage
partition_inner_roundsint1Base-learner rounds within each partition
refit_intervalint00 = fresh HVRT each round; k>0 = refit every k rounds (faster)
use_t_featureboolFalseAppend per-sample T-statistic (S²−Q) as an extra linear feature
use_coop_weightsboolFalseScale features by |corr(z_k, S−z_k)|² before linear fit
random_stateint42Seed (incremented per round for diverse partitionings)

GeoLinearClassifier accepts the same parameters.

Key methods

model.fit(X, y) # fitmodel.predict(X) # regression predictionsclf.predict_proba(X) # class probabilities, shape (n, 2)model.feature_importances(feature_names) # weighted mean |coef| across partitions/stagesmodel.stages_# list of (None, dict[partition_id, RidgeModel])

License

GNU Affero General Public License v3.0 or later. See LICENSE.

About

HVRT-informed linear modelling, comparable with XGBoost's accuracy while retaining interpretability.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - jpeaceau/GeoLinear: HVRT-informed linear modelling, comparable with XGBoost's accuracy while retaining interpretability. · GitHub
Skip to content

Repository files navigation

GeoLinear

Boosted piecewise-linear models on cooperative geometry partitions.

PyPI versionLicense: AGPL-3.0Python 3.10+


What is GeoLinear?

GeoLinear discovers cooperative geometry regimes in your data — groups of observations where features interact in similar ways — and fits an interpretable linear model inside each regime. Predictions are accumulated across boosting rounds.

The key insight is that many real-world relationships are piecewise-linear in cooperative geometry. A global linear model averages over regimes and loses the regime-specific signal. GeoLinear finds the regimes automatically (via HVRT) and lets the coefficients vary across them.

Round 1: HVRT partitions X → {cooperative, competitive, mixed} regimes
Ridge fits within each partition on residuals
Round 2: New partitioning on updated residuals → new local Ridge models
...
Final prediction = Σ (learning_rate × stage_predictions) + intercept

Each partition's Ridge coefficients are directly interpretable as local relativities — exactly the quantities actuaries file with regulators.


Why it matters for actuaries

Insurance pricing models must be both accurate and interpretable. Regulators require filed relativities; black-box models are inadmissible. The usual compromise — vanilla GLM — leaves accuracy on the table when the true relationship is regime-switching.

GeoLinear bridges this gap:

  1. Fit GeoLinear on the training portfolio. Each partition's Ridge coefficients are the relativities for that cooperative geometry segment.
  2. File a meta-GLM: fit OLS(X → ŷ_GL) to compress the ensemble into a single interpretable GLM. The compression loss is typically small (R² drop < 2%).
  3. Audit trail: every prediction can be traced to a specific partition and its local coefficient vector.

See examples/insurance_pricing_demo.py for a worked actuarial example including relativities tables and the GLM bridge.


Installation

pip install geolinear

Requires a C++17 compiler and CMake (automatically satisfied on most systems; the cmake PyPI package is a reliable fallback).

For the examples:

pip install geolinear[examples] # adds xgboost, optuna, matplotlib

Quick start

Regression

fromgeolinearimportGeoLinearfromsklearn.datasetsimportfetch_california_housingfromsklearn.model_selectionimporttrain_test_splitfromsklearn.metricsimportr2_scoreX, y=fetch_california_housing(return_X_y=True)
X_tr, X_te, y_tr, y_te=train_test_split(X, y, test_size=0.2, random_state=0)
model=GeoLinear(n_rounds=20, learning_rate=0.1, alpha=1.0)
model.fit(X_tr, y_tr)
print(r2_score(y_te, model.predict(X_te))) # ~0.72 default params

Classification

fromgeolinearimportGeoLinearClassifierfromsklearn.datasetsimportload_breast_cancerfromsklearn.model_selectionimporttrain_test_splitfromsklearn.metricsimportroc_auc_scoreX, y=load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te=train_test_split(X, y, test_size=0.2, random_state=0)
clf=GeoLinearClassifier(n_rounds=20, alpha=1.0)
clf.fit(X_tr, y_tr)
print(roc_auc_score(y_te, clf.predict_proba(X_te)[:, 1])) # ~0.993

Pipeline + GridSearchCV

fromsklearn.pipelineimportPipelinefromsklearn.preprocessingimportStandardScalerfromsklearn.model_selectionimportGridSearchCVfromgeolinearimportGeoLinearpipe=Pipeline([("scaler", StandardScaler()), ("gl", GeoLinear())])
grid=GridSearchCV(pipe, {"gl__alpha": [0.1, 1.0, 10.0], "gl__n_rounds": [10, 20]}, cv=5)
grid.fit(X_tr, y_tr)

Optuna HPO

importoptunafromgeolinearimportGeoLinearfromsklearn.model_selectionimportcross_val_scoredefobjective(trial):
model=GeoLinear(
n_rounds=trial.suggest_int("n_rounds", 10, 100),
learning_rate=trial.suggest_float("learning_rate", 0.01, 0.3, log=True),
alpha=trial.suggest_float("alpha", 0.01, 20.0, log=True),
y_weight=trial.suggest_float("y_weight", 0.1, 1.0),
base_learner=trial.suggest_categorical("base_learner", ["ridge", "lasso"]),
hvrt_n_partitions=trial.suggest_int("hvrt_n_partitions", 3, 14),
min_samples_partition=trial.suggest_int("min_samples_partition", 3, 25),
hvrt_model=trial.suggest_categorical("hvrt_model", ["hvrt", "pyramid_hart", "fast_hvrt"]),
use_t_feature=trial.suggest_categorical("use_t_feature", [False, True]),
refit_interval=trial.suggest_categorical("refit_interval", [0, 1, 2]),
)
returncross_val_score(model, X_tr, y_tr, cv=5, scoring="r2").mean()
study=optuna.create_study(direction="maximize")
study.optimize(objective, n_trials=60)

Ecosystem: which tool to use?

GeoLinear is part of a family of cooperative-geometry libraries. The right choice depends on what "interpretable" means in your domain and the structure of your data.

DomainPrimary needRecommended
Insurance pricingFiled relativities; regulator-auditable linear tariffGeoLinear
Credit risk / scoringScorecard linearity; SR 11-7 model risk complianceGeoLinear
Utility rate-settingFiled tariff schedules with linear rate factorsGeoLinear
Healthcare / clinicalXGBoost-level performance with fully deterministic, auditable predictionsGeoXGB
Ecology / environmentalNonlinear regime detection; richer explanation than SHAP via 100% determinismGeoXGB
Public policyAlgorithmic accountability without linearity constraintsGeoXGB
Personalised interventionsPer-entity longitudinal data; individual treatment trajectoriesAutoITE

Rule of thumb: if your regulator or ethics board requires a linear equation you can file or defend, use GeoLinear. If you need XGBoost-class accuracy with fully deterministic predictions that provide richer interpretability than SHAP, use GeoXGB. If you have repeated observations per individual and want to model how treatment effects evolve over time for each entity, use AutoITE.


Benchmark results

80-trial Optuna HPO, GeoLinear (v0.3.0) vs XGBoost. use_t_feature included in GeoLinear's HPO search space so the optimizer can exploit cooperative geometry when it exists.

Regression R² — v0.3.0 (80-trial HPO)

DGPGL-HPOuse_tXGB-HPOGap
T-regime0.5690.296+0.273 GL wins
3-regime0.3580.329+0.029 GL wins
Friedman10.9840.988−0.003 (tie)
Linear0.9650.957+0.008 GL wins
CalHousing0.5840.856−0.271 XGB wins

Score: 3 wins · 1 tie · 1 loss.

use_t_feature is selected by HPO whenever the data has cooperative structure (T-regime, 3-regime). On smooth or axis-aligned DGPs (Friedman1, Linear) HPO correctly opts out. CalHousing remains XGBoost territory: spatial autocorrelation and threshold effects are better captured by axis-aligned splits than by pairwise cooperative geometry.

Classification AUC

DatasetGL-Ridge (default)GL-HPOXGB-HPO
Synthetic0.9400.9570.981
BreastCancer0.9930.9720.993

Default-parameter GL-Ridge matches XGBoost-HPO on BreastCancer (AUC 0.993 each).


API reference

GeoLinear (regressor)

ParameterTypeDefaultDescription
n_roundsint20Boosting rounds
learning_ratefloat0.1Shrinkage per round
y_weightfloat0.5HVRT outcome-blend (0 = geometry-only, 1 = y-driven)
base_learnerstr"ridge""ridge", "ols", or "lasso"
alphafloat1.0L2 regularisation within each partition
min_samples_partitionint5Minimum samples to fit a partition model
hvrt_n_partitionsint|NoneNoneTarget partitions (None = HVRT auto-tune)
hvrt_min_samples_leafint|NoneNoneHVRT min leaf size
hvrt_inner_roundsint1HVRT T-residual inner rounds per stage
partition_inner_roundsint1Base-learner rounds within each partition
refit_intervalint00 = fresh HVRT each round; k>0 = refit every k rounds (faster)
use_t_featureboolFalseAppend per-sample T-statistic (S²−Q) as an extra linear feature
use_coop_weightsboolFalseScale features by |corr(z_k, S−z_k)|² before linear fit
random_stateint42Seed (incremented per round for diverse partitionings)

GeoLinearClassifier accepts the same parameters.

Key methods

model.fit(X, y) # fitmodel.predict(X) # regression predictionsclf.predict_proba(X) # class probabilities, shape (n, 2)model.feature_importances(feature_names) # weighted mean |coef| across partitions/stagesmodel.stages_# list of (None, dict[partition_id, RidgeModel])

License

GNU Affero General Public License v3.0 or later. See LICENSE.

About

HVRT-informed linear modelling, comparable with XGBoost's accuracy while retaining interpretability.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - jpeaceau/GeoLinear: HVRT-informed linear modelling, comparable with XGBoost's accuracy while retaining interpretability. · GitHub
Skip to content

Repository files navigation

GeoLinear

Boosted piecewise-linear models on cooperative geometry partitions.

PyPI versionLicense: AGPL-3.0Python 3.10+


What is GeoLinear?

GeoLinear discovers cooperative geometry regimes in your data — groups of observations where features interact in similar ways — and fits an interpretable linear model inside each regime. Predictions are accumulated across boosting rounds.

The key insight is that many real-world relationships are piecewise-linear in cooperative geometry. A global linear model averages over regimes and loses the regime-specific signal. GeoLinear finds the regimes automatically (via HVRT) and lets the coefficients vary across them.

Round 1: HVRT partitions X → {cooperative, competitive, mixed} regimes
Ridge fits within each partition on residuals
Round 2: New partitioning on updated residuals → new local Ridge models
...
Final prediction = Σ (learning_rate × stage_predictions) + intercept

Each partition's Ridge coefficients are directly interpretable as local relativities — exactly the quantities actuaries file with regulators.


Why it matters for actuaries

Insurance pricing models must be both accurate and interpretable. Regulators require filed relativities; black-box models are inadmissible. The usual compromise — vanilla GLM — leaves accuracy on the table when the true relationship is regime-switching.

GeoLinear bridges this gap:

  1. Fit GeoLinear on the training portfolio. Each partition's Ridge coefficients are the relativities for that cooperative geometry segment.
  2. File a meta-GLM: fit OLS(X → ŷ_GL) to compress the ensemble into a single interpretable GLM. The compression loss is typically small (R² drop < 2%).
  3. Audit trail: every prediction can be traced to a specific partition and its local coefficient vector.

See examples/insurance_pricing_demo.py for a worked actuarial example including relativities tables and the GLM bridge.


Installation

pip install geolinear

Requires a C++17 compiler and CMake (automatically satisfied on most systems; the cmake PyPI package is a reliable fallback).

For the examples:

pip install geolinear[examples] # adds xgboost, optuna, matplotlib

Quick start

Regression

fromgeolinearimportGeoLinearfromsklearn.datasetsimportfetch_california_housingfromsklearn.model_selectionimporttrain_test_splitfromsklearn.metricsimportr2_scoreX, y=fetch_california_housing(return_X_y=True)
X_tr, X_te, y_tr, y_te=train_test_split(X, y, test_size=0.2, random_state=0)
model=GeoLinear(n_rounds=20, learning_rate=0.1, alpha=1.0)
model.fit(X_tr, y_tr)
print(r2_score(y_te, model.predict(X_te))) # ~0.72 default params

Classification

fromgeolinearimportGeoLinearClassifierfromsklearn.datasetsimportload_breast_cancerfromsklearn.model_selectionimporttrain_test_splitfromsklearn.metricsimportroc_auc_scoreX, y=load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te=train_test_split(X, y, test_size=0.2, random_state=0)
clf=GeoLinearClassifier(n_rounds=20, alpha=1.0)
clf.fit(X_tr, y_tr)
print(roc_auc_score(y_te, clf.predict_proba(X_te)[:, 1])) # ~0.993

Pipeline + GridSearchCV

fromsklearn.pipelineimportPipelinefromsklearn.preprocessingimportStandardScalerfromsklearn.model_selectionimportGridSearchCVfromgeolinearimportGeoLinearpipe=Pipeline([("scaler", StandardScaler()), ("gl", GeoLinear())])
grid=GridSearchCV(pipe, {"gl__alpha": [0.1, 1.0, 10.0], "gl__n_rounds": [10, 20]}, cv=5)
grid.fit(X_tr, y_tr)

Optuna HPO

importoptunafromgeolinearimportGeoLinearfromsklearn.model_selectionimportcross_val_scoredefobjective(trial):
model=GeoLinear(
n_rounds=trial.suggest_int("n_rounds", 10, 100),
learning_rate=trial.suggest_float("learning_rate", 0.01, 0.3, log=True),
alpha=trial.suggest_float("alpha", 0.01, 20.0, log=True),
y_weight=trial.suggest_float("y_weight", 0.1, 1.0),
base_learner=trial.suggest_categorical("base_learner", ["ridge", "lasso"]),
hvrt_n_partitions=trial.suggest_int("hvrt_n_partitions", 3, 14),
min_samples_partition=trial.suggest_int("min_samples_partition", 3, 25),
hvrt_model=trial.suggest_categorical("hvrt_model", ["hvrt", "pyramid_hart", "fast_hvrt"]),
use_t_feature=trial.suggest_categorical("use_t_feature", [False, True]),
refit_interval=trial.suggest_categorical("refit_interval", [0, 1, 2]),
)
returncross_val_score(model, X_tr, y_tr, cv=5, scoring="r2").mean()
study=optuna.create_study(direction="maximize")
study.optimize(objective, n_trials=60)

Ecosystem: which tool to use?

GeoLinear is part of a family of cooperative-geometry libraries. The right choice depends on what "interpretable" means in your domain and the structure of your data.

DomainPrimary needRecommended
Insurance pricingFiled relativities; regulator-auditable linear tariffGeoLinear
Credit risk / scoringScorecard linearity; SR 11-7 model risk complianceGeoLinear
Utility rate-settingFiled tariff schedules with linear rate factorsGeoLinear
Healthcare / clinicalXGBoost-level performance with fully deterministic, auditable predictionsGeoXGB
Ecology / environmentalNonlinear regime detection; richer explanation than SHAP via 100% determinismGeoXGB
Public policyAlgorithmic accountability without linearity constraintsGeoXGB
Personalised interventionsPer-entity longitudinal data; individual treatment trajectoriesAutoITE

Rule of thumb: if your regulator or ethics board requires a linear equation you can file or defend, use GeoLinear. If you need XGBoost-class accuracy with fully deterministic predictions that provide richer interpretability than SHAP, use GeoXGB. If you have repeated observations per individual and want to model how treatment effects evolve over time for each entity, use AutoITE.


Benchmark results

80-trial Optuna HPO, GeoLinear (v0.3.0) vs XGBoost. use_t_feature included in GeoLinear's HPO search space so the optimizer can exploit cooperative geometry when it exists.

Regression R² — v0.3.0 (80-trial HPO)

DGPGL-HPOuse_tXGB-HPOGap
T-regime0.5690.296+0.273 GL wins
3-regime0.3580.329+0.029 GL wins
Friedman10.9840.988−0.003 (tie)
Linear0.9650.957+0.008 GL wins
CalHousing0.5840.856−0.271 XGB wins

Score: 3 wins · 1 tie · 1 loss.

use_t_feature is selected by HPO whenever the data has cooperative structure (T-regime, 3-regime). On smooth or axis-aligned DGPs (Friedman1, Linear) HPO correctly opts out. CalHousing remains XGBoost territory: spatial autocorrelation and threshold effects are better captured by axis-aligned splits than by pairwise cooperative geometry.

Classification AUC

DatasetGL-Ridge (default)GL-HPOXGB-HPO
Synthetic0.9400.9570.981
BreastCancer0.9930.9720.993

Default-parameter GL-Ridge matches XGBoost-HPO on BreastCancer (AUC 0.993 each).


API reference

GeoLinear (regressor)

ParameterTypeDefaultDescription
n_roundsint20Boosting rounds
learning_ratefloat0.1Shrinkage per round
y_weightfloat0.5HVRT outcome-blend (0 = geometry-only, 1 = y-driven)
base_learnerstr"ridge""ridge", "ols", or "lasso"
alphafloat1.0L2 regularisation within each partition
min_samples_partitionint5Minimum samples to fit a partition model
hvrt_n_partitionsint|NoneNoneTarget partitions (None = HVRT auto-tune)
hvrt_min_samples_leafint|NoneNoneHVRT min leaf size
hvrt_inner_roundsint1HVRT T-residual inner rounds per stage
partition_inner_roundsint1Base-learner rounds within each partition
refit_intervalint00 = fresh HVRT each round; k>0 = refit every k rounds (faster)
use_t_featureboolFalseAppend per-sample T-statistic (S²−Q) as an extra linear feature
use_coop_weightsboolFalseScale features by |corr(z_k, S−z_k)|² before linear fit
random_stateint42Seed (incremented per round for diverse partitionings)

GeoLinearClassifier accepts the same parameters.

Key methods

model.fit(X, y) # fitmodel.predict(X) # regression predictionsclf.predict_proba(X) # class probabilities, shape (n, 2)model.feature_importances(feature_names) # weighted mean |coef| across partitions/stagesmodel.stages_# list of (None, dict[partition_id, RidgeModel])

License

GNU Affero General Public License v3.0 or later. See LICENSE.

About

HVRT-informed linear modelling, comparable with XGBoost's accuracy while retaining interpretability.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - jpeaceau/GeoLinear: HVRT-informed linear modelling, comparable with XGBoost's accuracy while retaining interpretability. · GitHub
Skip to content

Repository files navigation

GeoLinear

Boosted piecewise-linear models on cooperative geometry partitions.

PyPI versionLicense: AGPL-3.0Python 3.10+


What is GeoLinear?

GeoLinear discovers cooperative geometry regimes in your data — groups of observations where features interact in similar ways — and fits an interpretable linear model inside each regime. Predictions are accumulated across boosting rounds.

The key insight is that many real-world relationships are piecewise-linear in cooperative geometry. A global linear model averages over regimes and loses the regime-specific signal. GeoLinear finds the regimes automatically (via HVRT) and lets the coefficients vary across them.

Round 1: HVRT partitions X → {cooperative, competitive, mixed} regimes
Ridge fits within each partition on residuals
Round 2: New partitioning on updated residuals → new local Ridge models
...
Final prediction = Σ (learning_rate × stage_predictions) + intercept

Each partition's Ridge coefficients are directly interpretable as local relativities — exactly the quantities actuaries file with regulators.


Why it matters for actuaries

Insurance pricing models must be both accurate and interpretable. Regulators require filed relativities; black-box models are inadmissible. The usual compromise — vanilla GLM — leaves accuracy on the table when the true relationship is regime-switching.

GeoLinear bridges this gap:

  1. Fit GeoLinear on the training portfolio. Each partition's Ridge coefficients are the relativities for that cooperative geometry segment.
  2. File a meta-GLM: fit OLS(X → ŷ_GL) to compress the ensemble into a single interpretable GLM. The compression loss is typically small (R² drop < 2%).
  3. Audit trail: every prediction can be traced to a specific partition and its local coefficient vector.

See examples/insurance_pricing_demo.py for a worked actuarial example including relativities tables and the GLM bridge.


Installation

pip install geolinear

Requires a C++17 compiler and CMake (automatically satisfied on most systems; the cmake PyPI package is a reliable fallback).

For the examples:

pip install geolinear[examples] # adds xgboost, optuna, matplotlib

Quick start

Regression

fromgeolinearimportGeoLinearfromsklearn.datasetsimportfetch_california_housingfromsklearn.model_selectionimporttrain_test_splitfromsklearn.metricsimportr2_scoreX, y=fetch_california_housing(return_X_y=True)
X_tr, X_te, y_tr, y_te=train_test_split(X, y, test_size=0.2, random_state=0)
model=GeoLinear(n_rounds=20, learning_rate=0.1, alpha=1.0)
model.fit(X_tr, y_tr)
print(r2_score(y_te, model.predict(X_te))) # ~0.72 default params

Classification

fromgeolinearimportGeoLinearClassifierfromsklearn.datasetsimportload_breast_cancerfromsklearn.model_selectionimporttrain_test_splitfromsklearn.metricsimportroc_auc_scoreX, y=load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te=train_test_split(X, y, test_size=0.2, random_state=0)
clf=GeoLinearClassifier(n_rounds=20, alpha=1.0)
clf.fit(X_tr, y_tr)
print(roc_auc_score(y_te, clf.predict_proba(X_te)[:, 1])) # ~0.993

Pipeline + GridSearchCV

fromsklearn.pipelineimportPipelinefromsklearn.preprocessingimportStandardScalerfromsklearn.model_selectionimportGridSearchCVfromgeolinearimportGeoLinearpipe=Pipeline([("scaler", StandardScaler()), ("gl", GeoLinear())])
grid=GridSearchCV(pipe, {"gl__alpha": [0.1, 1.0, 10.0], "gl__n_rounds": [10, 20]}, cv=5)
grid.fit(X_tr, y_tr)

Optuna HPO

importoptunafromgeolinearimportGeoLinearfromsklearn.model_selectionimportcross_val_scoredefobjective(trial):
model=GeoLinear(
n_rounds=trial.suggest_int("n_rounds", 10, 100),
learning_rate=trial.suggest_float("learning_rate", 0.01, 0.3, log=True),
alpha=trial.suggest_float("alpha", 0.01, 20.0, log=True),
y_weight=trial.suggest_float("y_weight", 0.1, 1.0),
base_learner=trial.suggest_categorical("base_learner", ["ridge", "lasso"]),
hvrt_n_partitions=trial.suggest_int("hvrt_n_partitions", 3, 14),
min_samples_partition=trial.suggest_int("min_samples_partition", 3, 25),
hvrt_model=trial.suggest_categorical("hvrt_model", ["hvrt", "pyramid_hart", "fast_hvrt"]),
use_t_feature=trial.suggest_categorical("use_t_feature", [False, True]),
refit_interval=trial.suggest_categorical("refit_interval", [0, 1, 2]),
)
returncross_val_score(model, X_tr, y_tr, cv=5, scoring="r2").mean()
study=optuna.create_study(direction="maximize")
study.optimize(objective, n_trials=60)

Ecosystem: which tool to use?

GeoLinear is part of a family of cooperative-geometry libraries. The right choice depends on what "interpretable" means in your domain and the structure of your data.

DomainPrimary needRecommended
Insurance pricingFiled relativities; regulator-auditable linear tariffGeoLinear
Credit risk / scoringScorecard linearity; SR 11-7 model risk complianceGeoLinear
Utility rate-settingFiled tariff schedules with linear rate factorsGeoLinear
Healthcare / clinicalXGBoost-level performance with fully deterministic, auditable predictionsGeoXGB
Ecology / environmentalNonlinear regime detection; richer explanation than SHAP via 100% determinismGeoXGB
Public policyAlgorithmic accountability without linearity constraintsGeoXGB
Personalised interventionsPer-entity longitudinal data; individual treatment trajectoriesAutoITE

Rule of thumb: if your regulator or ethics board requires a linear equation you can file or defend, use GeoLinear. If you need XGBoost-class accuracy with fully deterministic predictions that provide richer interpretability than SHAP, use GeoXGB. If you have repeated observations per individual and want to model how treatment effects evolve over time for each entity, use AutoITE.


Benchmark results

80-trial Optuna HPO, GeoLinear (v0.3.0) vs XGBoost. use_t_feature included in GeoLinear's HPO search space so the optimizer can exploit cooperative geometry when it exists.

Regression R² — v0.3.0 (80-trial HPO)

DGPGL-HPOuse_tXGB-HPOGap
T-regime0.5690.296+0.273 GL wins
3-regime0.3580.329+0.029 GL wins
Friedman10.9840.988−0.003 (tie)
Linear0.9650.957+0.008 GL wins
CalHousing0.5840.856−0.271 XGB wins

Score: 3 wins · 1 tie · 1 loss.

use_t_feature is selected by HPO whenever the data has cooperative structure (T-regime, 3-regime). On smooth or axis-aligned DGPs (Friedman1, Linear) HPO correctly opts out. CalHousing remains XGBoost territory: spatial autocorrelation and threshold effects are better captured by axis-aligned splits than by pairwise cooperative geometry.

Classification AUC

DatasetGL-Ridge (default)GL-HPOXGB-HPO
Synthetic0.9400.9570.981
BreastCancer0.9930.9720.993

Default-parameter GL-Ridge matches XGBoost-HPO on BreastCancer (AUC 0.993 each).


API reference

GeoLinear (regressor)

ParameterTypeDefaultDescription
n_roundsint20Boosting rounds
learning_ratefloat0.1Shrinkage per round
y_weightfloat0.5HVRT outcome-blend (0 = geometry-only, 1 = y-driven)
base_learnerstr"ridge""ridge", "ols", or "lasso"
alphafloat1.0L2 regularisation within each partition
min_samples_partitionint5Minimum samples to fit a partition model
hvrt_n_partitionsint|NoneNoneTarget partitions (None = HVRT auto-tune)
hvrt_min_samples_leafint|NoneNoneHVRT min leaf size
hvrt_inner_roundsint1HVRT T-residual inner rounds per stage
partition_inner_roundsint1Base-learner rounds within each partition
refit_intervalint00 = fresh HVRT each round; k>0 = refit every k rounds (faster)
use_t_featureboolFalseAppend per-sample T-statistic (S²−Q) as an extra linear feature
use_coop_weightsboolFalseScale features by |corr(z_k, S−z_k)|² before linear fit
random_stateint42Seed (incremented per round for diverse partitionings)

GeoLinearClassifier accepts the same parameters.

Key methods

model.fit(X, y) # fitmodel.predict(X) # regression predictionsclf.predict_proba(X) # class probabilities, shape (n, 2)model.feature_importances(feature_names) # weighted mean |coef| across partitions/stagesmodel.stages_# list of (None, dict[partition_id, RidgeModel])

License

GNU Affero General Public License v3.0 or later. See LICENSE.

About

HVRT-informed linear modelling, comparable with XGBoost's accuracy while retaining interpretability.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - jpeaceau/GeoLinear: HVRT-informed linear modelling, comparable with XGBoost's accuracy while retaining interpretability. · GitHub
Skip to content

Repository files navigation

GeoLinear

Boosted piecewise-linear models on cooperative geometry partitions.

PyPI versionLicense: AGPL-3.0Python 3.10+


What is GeoLinear?

GeoLinear discovers cooperative geometry regimes in your data — groups of observations where features interact in similar ways — and fits an interpretable linear model inside each regime. Predictions are accumulated across boosting rounds.

The key insight is that many real-world relationships are piecewise-linear in cooperative geometry. A global linear model averages over regimes and loses the regime-specific signal. GeoLinear finds the regimes automatically (via HVRT) and lets the coefficients vary across them.

Round 1: HVRT partitions X → {cooperative, competitive, mixed} regimes
Ridge fits within each partition on residuals
Round 2: New partitioning on updated residuals → new local Ridge models
...
Final prediction = Σ (learning_rate × stage_predictions) + intercept

Each partition's Ridge coefficients are directly interpretable as local relativities — exactly the quantities actuaries file with regulators.


Why it matters for actuaries

Insurance pricing models must be both accurate and interpretable. Regulators require filed relativities; black-box models are inadmissible. The usual compromise — vanilla GLM — leaves accuracy on the table when the true relationship is regime-switching.

GeoLinear bridges this gap:

  1. Fit GeoLinear on the training portfolio. Each partition's Ridge coefficients are the relativities for that cooperative geometry segment.
  2. File a meta-GLM: fit OLS(X → ŷ_GL) to compress the ensemble into a single interpretable GLM. The compression loss is typically small (R² drop < 2%).
  3. Audit trail: every prediction can be traced to a specific partition and its local coefficient vector.

See examples/insurance_pricing_demo.py for a worked actuarial example including relativities tables and the GLM bridge.


Installation

pip install geolinear

Requires a C++17 compiler and CMake (automatically satisfied on most systems; the cmake PyPI package is a reliable fallback).

For the examples:

pip install geolinear[examples] # adds xgboost, optuna, matplotlib

Quick start

Regression

fromgeolinearimportGeoLinearfromsklearn.datasetsimportfetch_california_housingfromsklearn.model_selectionimporttrain_test_splitfromsklearn.metricsimportr2_scoreX, y=fetch_california_housing(return_X_y=True)
X_tr, X_te, y_tr, y_te=train_test_split(X, y, test_size=0.2, random_state=0)
model=GeoLinear(n_rounds=20, learning_rate=0.1, alpha=1.0)
model.fit(X_tr, y_tr)
print(r2_score(y_te, model.predict(X_te))) # ~0.72 default params

Classification

fromgeolinearimportGeoLinearClassifierfromsklearn.datasetsimportload_breast_cancerfromsklearn.model_selectionimporttrain_test_splitfromsklearn.metricsimportroc_auc_scoreX, y=load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te=train_test_split(X, y, test_size=0.2, random_state=0)
clf=GeoLinearClassifier(n_rounds=20, alpha=1.0)
clf.fit(X_tr, y_tr)
print(roc_auc_score(y_te, clf.predict_proba(X_te)[:, 1])) # ~0.993

Pipeline + GridSearchCV

fromsklearn.pipelineimportPipelinefromsklearn.preprocessingimportStandardScalerfromsklearn.model_selectionimportGridSearchCVfromgeolinearimportGeoLinearpipe=Pipeline([("scaler", StandardScaler()), ("gl", GeoLinear())])
grid=GridSearchCV(pipe, {"gl__alpha": [0.1, 1.0, 10.0], "gl__n_rounds": [10, 20]}, cv=5)
grid.fit(X_tr, y_tr)

Optuna HPO

importoptunafromgeolinearimportGeoLinearfromsklearn.model_selectionimportcross_val_scoredefobjective(trial):
model=GeoLinear(
n_rounds=trial.suggest_int("n_rounds", 10, 100),
learning_rate=trial.suggest_float("learning_rate", 0.01, 0.3, log=True),
alpha=trial.suggest_float("alpha", 0.01, 20.0, log=True),
y_weight=trial.suggest_float("y_weight", 0.1, 1.0),
base_learner=trial.suggest_categorical("base_learner", ["ridge", "lasso"]),
hvrt_n_partitions=trial.suggest_int("hvrt_n_partitions", 3, 14),
min_samples_partition=trial.suggest_int("min_samples_partition", 3, 25),
hvrt_model=trial.suggest_categorical("hvrt_model", ["hvrt", "pyramid_hart", "fast_hvrt"]),
use_t_feature=trial.suggest_categorical("use_t_feature", [False, True]),
refit_interval=trial.suggest_categorical("refit_interval", [0, 1, 2]),
)
returncross_val_score(model, X_tr, y_tr, cv=5, scoring="r2").mean()
study=optuna.create_study(direction="maximize")
study.optimize(objective, n_trials=60)

Ecosystem: which tool to use?

GeoLinear is part of a family of cooperative-geometry libraries. The right choice depends on what "interpretable" means in your domain and the structure of your data.

DomainPrimary needRecommended
Insurance pricingFiled relativities; regulator-auditable linear tariffGeoLinear
Credit risk / scoringScorecard linearity; SR 11-7 model risk complianceGeoLinear
Utility rate-settingFiled tariff schedules with linear rate factorsGeoLinear
Healthcare / clinicalXGBoost-level performance with fully deterministic, auditable predictionsGeoXGB
Ecology / environmentalNonlinear regime detection; richer explanation than SHAP via 100% determinismGeoXGB
Public policyAlgorithmic accountability without linearity constraintsGeoXGB
Personalised interventionsPer-entity longitudinal data; individual treatment trajectoriesAutoITE

Rule of thumb: if your regulator or ethics board requires a linear equation you can file or defend, use GeoLinear. If you need XGBoost-class accuracy with fully deterministic predictions that provide richer interpretability than SHAP, use GeoXGB. If you have repeated observations per individual and want to model how treatment effects evolve over time for each entity, use AutoITE.


Benchmark results

80-trial Optuna HPO, GeoLinear (v0.3.0) vs XGBoost. use_t_feature included in GeoLinear's HPO search space so the optimizer can exploit cooperative geometry when it exists.

Regression R² — v0.3.0 (80-trial HPO)

DGPGL-HPOuse_tXGB-HPOGap
T-regime0.5690.296+0.273 GL wins
3-regime0.3580.329+0.029 GL wins
Friedman10.9840.988−0.003 (tie)
Linear0.9650.957+0.008 GL wins
CalHousing0.5840.856−0.271 XGB wins

Score: 3 wins · 1 tie · 1 loss.

use_t_feature is selected by HPO whenever the data has cooperative structure (T-regime, 3-regime). On smooth or axis-aligned DGPs (Friedman1, Linear) HPO correctly opts out. CalHousing remains XGBoost territory: spatial autocorrelation and threshold effects are better captured by axis-aligned splits than by pairwise cooperative geometry.

Classification AUC

DatasetGL-Ridge (default)GL-HPOXGB-HPO
Synthetic0.9400.9570.981
BreastCancer0.9930.9720.993

Default-parameter GL-Ridge matches XGBoost-HPO on BreastCancer (AUC 0.993 each).


API reference

GeoLinear (regressor)

ParameterTypeDefaultDescription
n_roundsint20Boosting rounds
learning_ratefloat0.1Shrinkage per round
y_weightfloat0.5HVRT outcome-blend (0 = geometry-only, 1 = y-driven)
base_learnerstr"ridge""ridge", "ols", or "lasso"
alphafloat1.0L2 regularisation within each partition
min_samples_partitionint5Minimum samples to fit a partition model
hvrt_n_partitionsint|NoneNoneTarget partitions (None = HVRT auto-tune)
hvrt_min_samples_leafint|NoneNoneHVRT min leaf size
hvrt_inner_roundsint1HVRT T-residual inner rounds per stage
partition_inner_roundsint1Base-learner rounds within each partition
refit_intervalint00 = fresh HVRT each round; k>0 = refit every k rounds (faster)
use_t_featureboolFalseAppend per-sample T-statistic (S²−Q) as an extra linear feature
use_coop_weightsboolFalseScale features by |corr(z_k, S−z_k)|² before linear fit
random_stateint42Seed (incremented per round for diverse partitionings)

GeoLinearClassifier accepts the same parameters.

Key methods

model.fit(X, y) # fitmodel.predict(X) # regression predictionsclf.predict_proba(X) # class probabilities, shape (n, 2)model.feature_importances(feature_names) # weighted mean |coef| across partitions/stagesmodel.stages_# list of (None, dict[partition_id, RidgeModel])

License

GNU Affero General Public License v3.0 or later. See LICENSE.

About

HVRT-informed linear modelling, comparable with XGBoost's accuracy while retaining interpretability.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - jpeaceau/GeoLinear: HVRT-informed linear modelling, comparable with XGBoost's accuracy while retaining interpretability. · GitHub
Skip to content

Repository files navigation

GeoLinear

Boosted piecewise-linear models on cooperative geometry partitions.

PyPI versionLicense: AGPL-3.0Python 3.10+


What is GeoLinear?

GeoLinear discovers cooperative geometry regimes in your data — groups of observations where features interact in similar ways — and fits an interpretable linear model inside each regime. Predictions are accumulated across boosting rounds.

The key insight is that many real-world relationships are piecewise-linear in cooperative geometry. A global linear model averages over regimes and loses the regime-specific signal. GeoLinear finds the regimes automatically (via HVRT) and lets the coefficients vary across them.

Round 1: HVRT partitions X → {cooperative, competitive, mixed} regimes
Ridge fits within each partition on residuals
Round 2: New partitioning on updated residuals → new local Ridge models
...
Final prediction = Σ (learning_rate × stage_predictions) + intercept

Each partition's Ridge coefficients are directly interpretable as local relativities — exactly the quantities actuaries file with regulators.


Why it matters for actuaries

Insurance pricing models must be both accurate and interpretable. Regulators require filed relativities; black-box models are inadmissible. The usual compromise — vanilla GLM — leaves accuracy on the table when the true relationship is regime-switching.

GeoLinear bridges this gap:

  1. Fit GeoLinear on the training portfolio. Each partition's Ridge coefficients are the relativities for that cooperative geometry segment.
  2. File a meta-GLM: fit OLS(X → ŷ_GL) to compress the ensemble into a single interpretable GLM. The compression loss is typically small (R² drop < 2%).
  3. Audit trail: every prediction can be traced to a specific partition and its local coefficient vector.

See examples/insurance_pricing_demo.py for a worked actuarial example including relativities tables and the GLM bridge.


Installation

pip install geolinear

Requires a C++17 compiler and CMake (automatically satisfied on most systems; the cmake PyPI package is a reliable fallback).

For the examples:

pip install geolinear[examples] # adds xgboost, optuna, matplotlib

Quick start

Regression

fromgeolinearimportGeoLinearfromsklearn.datasetsimportfetch_california_housingfromsklearn.model_selectionimporttrain_test_splitfromsklearn.metricsimportr2_scoreX, y=fetch_california_housing(return_X_y=True)
X_tr, X_te, y_tr, y_te=train_test_split(X, y, test_size=0.2, random_state=0)
model=GeoLinear(n_rounds=20, learning_rate=0.1, alpha=1.0)
model.fit(X_tr, y_tr)
print(r2_score(y_te, model.predict(X_te))) # ~0.72 default params

Classification

fromgeolinearimportGeoLinearClassifierfromsklearn.datasetsimportload_breast_cancerfromsklearn.model_selectionimporttrain_test_splitfromsklearn.metricsimportroc_auc_scoreX, y=load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te=train_test_split(X, y, test_size=0.2, random_state=0)
clf=GeoLinearClassifier(n_rounds=20, alpha=1.0)
clf.fit(X_tr, y_tr)
print(roc_auc_score(y_te, clf.predict_proba(X_te)[:, 1])) # ~0.993

Pipeline + GridSearchCV

fromsklearn.pipelineimportPipelinefromsklearn.preprocessingimportStandardScalerfromsklearn.model_selectionimportGridSearchCVfromgeolinearimportGeoLinearpipe=Pipeline([("scaler", StandardScaler()), ("gl", GeoLinear())])
grid=GridSearchCV(pipe, {"gl__alpha": [0.1, 1.0, 10.0], "gl__n_rounds": [10, 20]}, cv=5)
grid.fit(X_tr, y_tr)

Optuna HPO

importoptunafromgeolinearimportGeoLinearfromsklearn.model_selectionimportcross_val_scoredefobjective(trial):
model=GeoLinear(
n_rounds=trial.suggest_int("n_rounds", 10, 100),
learning_rate=trial.suggest_float("learning_rate", 0.01, 0.3, log=True),
alpha=trial.suggest_float("alpha", 0.01, 20.0, log=True),
y_weight=trial.suggest_float("y_weight", 0.1, 1.0),
base_learner=trial.suggest_categorical("base_learner", ["ridge", "lasso"]),
hvrt_n_partitions=trial.suggest_int("hvrt_n_partitions", 3, 14),
min_samples_partition=trial.suggest_int("min_samples_partition", 3, 25),
hvrt_model=trial.suggest_categorical("hvrt_model", ["hvrt", "pyramid_hart", "fast_hvrt"]),
use_t_feature=trial.suggest_categorical("use_t_feature", [False, True]),
refit_interval=trial.suggest_categorical("refit_interval", [0, 1, 2]),
)
returncross_val_score(model, X_tr, y_tr, cv=5, scoring="r2").mean()
study=optuna.create_study(direction="maximize")
study.optimize(objective, n_trials=60)

Ecosystem: which tool to use?

GeoLinear is part of a family of cooperative-geometry libraries. The right choice depends on what "interpretable" means in your domain and the structure of your data.

DomainPrimary needRecommended
Insurance pricingFiled relativities; regulator-auditable linear tariffGeoLinear
Credit risk / scoringScorecard linearity; SR 11-7 model risk complianceGeoLinear
Utility rate-settingFiled tariff schedules with linear rate factorsGeoLinear
Healthcare / clinicalXGBoost-level performance with fully deterministic, auditable predictionsGeoXGB
Ecology / environmentalNonlinear regime detection; richer explanation than SHAP via 100% determinismGeoXGB
Public policyAlgorithmic accountability without linearity constraintsGeoXGB
Personalised interventionsPer-entity longitudinal data; individual treatment trajectoriesAutoITE

Rule of thumb: if your regulator or ethics board requires a linear equation you can file or defend, use GeoLinear. If you need XGBoost-class accuracy with fully deterministic predictions that provide richer interpretability than SHAP, use GeoXGB. If you have repeated observations per individual and want to model how treatment effects evolve over time for each entity, use AutoITE.


Benchmark results

80-trial Optuna HPO, GeoLinear (v0.3.0) vs XGBoost. use_t_feature included in GeoLinear's HPO search space so the optimizer can exploit cooperative geometry when it exists.

Regression R² — v0.3.0 (80-trial HPO)

DGPGL-HPOuse_tXGB-HPOGap
T-regime0.5690.296+0.273 GL wins
3-regime0.3580.329+0.029 GL wins
Friedman10.9840.988−0.003 (tie)
Linear0.9650.957+0.008 GL wins
CalHousing0.5840.856−0.271 XGB wins

Score: 3 wins · 1 tie · 1 loss.

use_t_feature is selected by HPO whenever the data has cooperative structure (T-regime, 3-regime). On smooth or axis-aligned DGPs (Friedman1, Linear) HPO correctly opts out. CalHousing remains XGBoost territory: spatial autocorrelation and threshold effects are better captured by axis-aligned splits than by pairwise cooperative geometry.

Classification AUC

DatasetGL-Ridge (default)GL-HPOXGB-HPO
Synthetic0.9400.9570.981
BreastCancer0.9930.9720.993

Default-parameter GL-Ridge matches XGBoost-HPO on BreastCancer (AUC 0.993 each).


API reference

GeoLinear (regressor)

ParameterTypeDefaultDescription
n_roundsint20Boosting rounds
learning_ratefloat0.1Shrinkage per round
y_weightfloat0.5HVRT outcome-blend (0 = geometry-only, 1 = y-driven)
base_learnerstr"ridge""ridge", "ols", or "lasso"
alphafloat1.0L2 regularisation within each partition
min_samples_partitionint5Minimum samples to fit a partition model
hvrt_n_partitionsint|NoneNoneTarget partitions (None = HVRT auto-tune)
hvrt_min_samples_leafint|NoneNoneHVRT min leaf size
hvrt_inner_roundsint1HVRT T-residual inner rounds per stage
partition_inner_roundsint1Base-learner rounds within each partition
refit_intervalint00 = fresh HVRT each round; k>0 = refit every k rounds (faster)
use_t_featureboolFalseAppend per-sample T-statistic (S²−Q) as an extra linear feature
use_coop_weightsboolFalseScale features by |corr(z_k, S−z_k)|² before linear fit
random_stateint42Seed (incremented per round for diverse partitionings)

GeoLinearClassifier accepts the same parameters.

Key methods

model.fit(X, y) # fitmodel.predict(X) # regression predictionsclf.predict_proba(X) # class probabilities, shape (n, 2)model.feature_importances(feature_names) # weighted mean |coef| across partitions/stagesmodel.stages_# list of (None, dict[partition_id, RidgeModel])

License

GNU Affero General Public License v3.0 or later. See LICENSE.

About

HVRT-informed linear modelling, comparable with XGBoost's accuracy while retaining interpretability.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - jpeaceau/GeoLinear: HVRT-informed linear modelling, comparable with XGBoost's accuracy while retaining interpretability. · GitHub
Skip to content

Repository files navigation

GeoLinear

Boosted piecewise-linear models on cooperative geometry partitions.

PyPI versionLicense: AGPL-3.0Python 3.10+


What is GeoLinear?

GeoLinear discovers cooperative geometry regimes in your data — groups of observations where features interact in similar ways — and fits an interpretable linear model inside each regime. Predictions are accumulated across boosting rounds.

The key insight is that many real-world relationships are piecewise-linear in cooperative geometry. A global linear model averages over regimes and loses the regime-specific signal. GeoLinear finds the regimes automatically (via HVRT) and lets the coefficients vary across them.

Round 1: HVRT partitions X → {cooperative, competitive, mixed} regimes
Ridge fits within each partition on residuals
Round 2: New partitioning on updated residuals → new local Ridge models
...
Final prediction = Σ (learning_rate × stage_predictions) + intercept

Each partition's Ridge coefficients are directly interpretable as local relativities — exactly the quantities actuaries file with regulators.


Why it matters for actuaries

Insurance pricing models must be both accurate and interpretable. Regulators require filed relativities; black-box models are inadmissible. The usual compromise — vanilla GLM — leaves accuracy on the table when the true relationship is regime-switching.

GeoLinear bridges this gap:

  1. Fit GeoLinear on the training portfolio. Each partition's Ridge coefficients are the relativities for that cooperative geometry segment.
  2. File a meta-GLM: fit OLS(X → ŷ_GL) to compress the ensemble into a single interpretable GLM. The compression loss is typically small (R² drop < 2%).
  3. Audit trail: every prediction can be traced to a specific partition and its local coefficient vector.

See examples/insurance_pricing_demo.py for a worked actuarial example including relativities tables and the GLM bridge.


Installation

pip install geolinear

Requires a C++17 compiler and CMake (automatically satisfied on most systems; the cmake PyPI package is a reliable fallback).

For the examples:

pip install geolinear[examples] # adds xgboost, optuna, matplotlib

Quick start

Regression

fromgeolinearimportGeoLinearfromsklearn.datasetsimportfetch_california_housingfromsklearn.model_selectionimporttrain_test_splitfromsklearn.metricsimportr2_scoreX, y=fetch_california_housing(return_X_y=True)
X_tr, X_te, y_tr, y_te=train_test_split(X, y, test_size=0.2, random_state=0)
model=GeoLinear(n_rounds=20, learning_rate=0.1, alpha=1.0)
model.fit(X_tr, y_tr)
print(r2_score(y_te, model.predict(X_te))) # ~0.72 default params

Classification

fromgeolinearimportGeoLinearClassifierfromsklearn.datasetsimportload_breast_cancerfromsklearn.model_selectionimporttrain_test_splitfromsklearn.metricsimportroc_auc_scoreX, y=load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te=train_test_split(X, y, test_size=0.2, random_state=0)
clf=GeoLinearClassifier(n_rounds=20, alpha=1.0)
clf.fit(X_tr, y_tr)
print(roc_auc_score(y_te, clf.predict_proba(X_te)[:, 1])) # ~0.993

Pipeline + GridSearchCV

fromsklearn.pipelineimportPipelinefromsklearn.preprocessingimportStandardScalerfromsklearn.model_selectionimportGridSearchCVfromgeolinearimportGeoLinearpipe=Pipeline([("scaler", StandardScaler()), ("gl", GeoLinear())])
grid=GridSearchCV(pipe, {"gl__alpha": [0.1, 1.0, 10.0], "gl__n_rounds": [10, 20]}, cv=5)
grid.fit(X_tr, y_tr)

Optuna HPO

importoptunafromgeolinearimportGeoLinearfromsklearn.model_selectionimportcross_val_scoredefobjective(trial):
model=GeoLinear(
n_rounds=trial.suggest_int("n_rounds", 10, 100),
learning_rate=trial.suggest_float("learning_rate", 0.01, 0.3, log=True),
alpha=trial.suggest_float("alpha", 0.01, 20.0, log=True),
y_weight=trial.suggest_float("y_weight", 0.1, 1.0),
base_learner=trial.suggest_categorical("base_learner", ["ridge", "lasso"]),
hvrt_n_partitions=trial.suggest_int("hvrt_n_partitions", 3, 14),
min_samples_partition=trial.suggest_int("min_samples_partition", 3, 25),
hvrt_model=trial.suggest_categorical("hvrt_model", ["hvrt", "pyramid_hart", "fast_hvrt"]),
use_t_feature=trial.suggest_categorical("use_t_feature", [False, True]),
refit_interval=trial.suggest_categorical("refit_interval", [0, 1, 2]),
)
returncross_val_score(model, X_tr, y_tr, cv=5, scoring="r2").mean()
study=optuna.create_study(direction="maximize")
study.optimize(objective, n_trials=60)

Ecosystem: which tool to use?

GeoLinear is part of a family of cooperative-geometry libraries. The right choice depends on what "interpretable" means in your domain and the structure of your data.

DomainPrimary needRecommended
Insurance pricingFiled relativities; regulator-auditable linear tariffGeoLinear
Credit risk / scoringScorecard linearity; SR 11-7 model risk complianceGeoLinear
Utility rate-settingFiled tariff schedules with linear rate factorsGeoLinear
Healthcare / clinicalXGBoost-level performance with fully deterministic, auditable predictionsGeoXGB
Ecology / environmentalNonlinear regime detection; richer explanation than SHAP via 100% determinismGeoXGB
Public policyAlgorithmic accountability without linearity constraintsGeoXGB
Personalised interventionsPer-entity longitudinal data; individual treatment trajectoriesAutoITE

Rule of thumb: if your regulator or ethics board requires a linear equation you can file or defend, use GeoLinear. If you need XGBoost-class accuracy with fully deterministic predictions that provide richer interpretability than SHAP, use GeoXGB. If you have repeated observations per individual and want to model how treatment effects evolve over time for each entity, use AutoITE.


Benchmark results

80-trial Optuna HPO, GeoLinear (v0.3.0) vs XGBoost. use_t_feature included in GeoLinear's HPO search space so the optimizer can exploit cooperative geometry when it exists.

Regression R² — v0.3.0 (80-trial HPO)

DGPGL-HPOuse_tXGB-HPOGap
T-regime0.5690.296+0.273 GL wins
3-regime0.3580.329+0.029 GL wins
Friedman10.9840.988−0.003 (tie)
Linear0.9650.957+0.008 GL wins
CalHousing0.5840.856−0.271 XGB wins

Score: 3 wins · 1 tie · 1 loss.

use_t_feature is selected by HPO whenever the data has cooperative structure (T-regime, 3-regime). On smooth or axis-aligned DGPs (Friedman1, Linear) HPO correctly opts out. CalHousing remains XGBoost territory: spatial autocorrelation and threshold effects are better captured by axis-aligned splits than by pairwise cooperative geometry.

Classification AUC

DatasetGL-Ridge (default)GL-HPOXGB-HPO
Synthetic0.9400.9570.981
BreastCancer0.9930.9720.993

Default-parameter GL-Ridge matches XGBoost-HPO on BreastCancer (AUC 0.993 each).


API reference

GeoLinear (regressor)

ParameterTypeDefaultDescription
n_roundsint20Boosting rounds
learning_ratefloat0.1Shrinkage per round
y_weightfloat0.5HVRT outcome-blend (0 = geometry-only, 1 = y-driven)
base_learnerstr"ridge""ridge", "ols", or "lasso"
alphafloat1.0L2 regularisation within each partition
min_samples_partitionint5Minimum samples to fit a partition model
hvrt_n_partitionsint|NoneNoneTarget partitions (None = HVRT auto-tune)
hvrt_min_samples_leafint|NoneNoneHVRT min leaf size
hvrt_inner_roundsint1HVRT T-residual inner rounds per stage
partition_inner_roundsint1Base-learner rounds within each partition
refit_intervalint00 = fresh HVRT each round; k>0 = refit every k rounds (faster)
use_t_featureboolFalseAppend per-sample T-statistic (S²−Q) as an extra linear feature
use_coop_weightsboolFalseScale features by |corr(z_k, S−z_k)|² before linear fit
random_stateint42Seed (incremented per round for diverse partitionings)

GeoLinearClassifier accepts the same parameters.

Key methods

model.fit(X, y) # fitmodel.predict(X) # regression predictionsclf.predict_proba(X) # class probabilities, shape (n, 2)model.feature_importances(feature_names) # weighted mean |coef| across partitions/stagesmodel.stages_# list of (None, dict[partition_id, RidgeModel])

License

GNU Affero General Public License v3.0 or later. See LICENSE.

About

HVRT-informed linear modelling, comparable with XGBoost's accuracy while retaining interpretability.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - jpeaceau/GeoLinear: HVRT-informed linear modelling, comparable with XGBoost's accuracy while retaining interpretability. · GitHub
Skip to content

Repository files navigation

GeoLinear

Boosted piecewise-linear models on cooperative geometry partitions.

PyPI versionLicense: AGPL-3.0Python 3.10+


What is GeoLinear?

GeoLinear discovers cooperative geometry regimes in your data — groups of observations where features interact in similar ways — and fits an interpretable linear model inside each regime. Predictions are accumulated across boosting rounds.

The key insight is that many real-world relationships are piecewise-linear in cooperative geometry. A global linear model averages over regimes and loses the regime-specific signal. GeoLinear finds the regimes automatically (via HVRT) and lets the coefficients vary across them.

Round 1: HVRT partitions X → {cooperative, competitive, mixed} regimes
Ridge fits within each partition on residuals
Round 2: New partitioning on updated residuals → new local Ridge models
...
Final prediction = Σ (learning_rate × stage_predictions) + intercept

Each partition's Ridge coefficients are directly interpretable as local relativities — exactly the quantities actuaries file with regulators.


Why it matters for actuaries

Insurance pricing models must be both accurate and interpretable. Regulators require filed relativities; black-box models are inadmissible. The usual compromise — vanilla GLM — leaves accuracy on the table when the true relationship is regime-switching.

GeoLinear bridges this gap:

  1. Fit GeoLinear on the training portfolio. Each partition's Ridge coefficients are the relativities for that cooperative geometry segment.
  2. File a meta-GLM: fit OLS(X → ŷ_GL) to compress the ensemble into a single interpretable GLM. The compression loss is typically small (R² drop < 2%).
  3. Audit trail: every prediction can be traced to a specific partition and its local coefficient vector.

See examples/insurance_pricing_demo.py for a worked actuarial example including relativities tables and the GLM bridge.


Installation

pip install geolinear

Requires a C++17 compiler and CMake (automatically satisfied on most systems; the cmake PyPI package is a reliable fallback).

For the examples:

pip install geolinear[examples] # adds xgboost, optuna, matplotlib

Quick start

Regression

fromgeolinearimportGeoLinearfromsklearn.datasetsimportfetch_california_housingfromsklearn.model_selectionimporttrain_test_splitfromsklearn.metricsimportr2_scoreX, y=fetch_california_housing(return_X_y=True)
X_tr, X_te, y_tr, y_te=train_test_split(X, y, test_size=0.2, random_state=0)
model=GeoLinear(n_rounds=20, learning_rate=0.1, alpha=1.0)
model.fit(X_tr, y_tr)
print(r2_score(y_te, model.predict(X_te))) # ~0.72 default params

Classification

fromgeolinearimportGeoLinearClassifierfromsklearn.datasetsimportload_breast_cancerfromsklearn.model_selectionimporttrain_test_splitfromsklearn.metricsimportroc_auc_scoreX, y=load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te=train_test_split(X, y, test_size=0.2, random_state=0)
clf=GeoLinearClassifier(n_rounds=20, alpha=1.0)
clf.fit(X_tr, y_tr)
print(roc_auc_score(y_te, clf.predict_proba(X_te)[:, 1])) # ~0.993

Pipeline + GridSearchCV

fromsklearn.pipelineimportPipelinefromsklearn.preprocessingimportStandardScalerfromsklearn.model_selectionimportGridSearchCVfromgeolinearimportGeoLinearpipe=Pipeline([("scaler", StandardScaler()), ("gl", GeoLinear())])
grid=GridSearchCV(pipe, {"gl__alpha": [0.1, 1.0, 10.0], "gl__n_rounds": [10, 20]}, cv=5)
grid.fit(X_tr, y_tr)

Optuna HPO

importoptunafromgeolinearimportGeoLinearfromsklearn.model_selectionimportcross_val_scoredefobjective(trial):
model=GeoLinear(
n_rounds=trial.suggest_int("n_rounds", 10, 100),
learning_rate=trial.suggest_float("learning_rate", 0.01, 0.3, log=True),
alpha=trial.suggest_float("alpha", 0.01, 20.0, log=True),
y_weight=trial.suggest_float("y_weight", 0.1, 1.0),
base_learner=trial.suggest_categorical("base_learner", ["ridge", "lasso"]),
hvrt_n_partitions=trial.suggest_int("hvrt_n_partitions", 3, 14),
min_samples_partition=trial.suggest_int("min_samples_partition", 3, 25),
hvrt_model=trial.suggest_categorical("hvrt_model", ["hvrt", "pyramid_hart", "fast_hvrt"]),
use_t_feature=trial.suggest_categorical("use_t_feature", [False, True]),
refit_interval=trial.suggest_categorical("refit_interval", [0, 1, 2]),
)
returncross_val_score(model, X_tr, y_tr, cv=5, scoring="r2").mean()
study=optuna.create_study(direction="maximize")
study.optimize(objective, n_trials=60)

Ecosystem: which tool to use?

GeoLinear is part of a family of cooperative-geometry libraries. The right choice depends on what "interpretable" means in your domain and the structure of your data.

DomainPrimary needRecommended
Insurance pricingFiled relativities; regulator-auditable linear tariffGeoLinear
Credit risk / scoringScorecard linearity; SR 11-7 model risk complianceGeoLinear
Utility rate-settingFiled tariff schedules with linear rate factorsGeoLinear
Healthcare / clinicalXGBoost-level performance with fully deterministic, auditable predictionsGeoXGB
Ecology / environmentalNonlinear regime detection; richer explanation than SHAP via 100% determinismGeoXGB
Public policyAlgorithmic accountability without linearity constraintsGeoXGB
Personalised interventionsPer-entity longitudinal data; individual treatment trajectoriesAutoITE

Rule of thumb: if your regulator or ethics board requires a linear equation you can file or defend, use GeoLinear. If you need XGBoost-class accuracy with fully deterministic predictions that provide richer interpretability than SHAP, use GeoXGB. If you have repeated observations per individual and want to model how treatment effects evolve over time for each entity, use AutoITE.


Benchmark results

80-trial Optuna HPO, GeoLinear (v0.3.0) vs XGBoost. use_t_feature included in GeoLinear's HPO search space so the optimizer can exploit cooperative geometry when it exists.

Regression R² — v0.3.0 (80-trial HPO)

DGPGL-HPOuse_tXGB-HPOGap
T-regime0.5690.296+0.273 GL wins
3-regime0.3580.329+0.029 GL wins
Friedman10.9840.988−0.003 (tie)
Linear0.9650.957+0.008 GL wins
CalHousing0.5840.856−0.271 XGB wins

Score: 3 wins · 1 tie · 1 loss.

use_t_feature is selected by HPO whenever the data has cooperative structure (T-regime, 3-regime). On smooth or axis-aligned DGPs (Friedman1, Linear) HPO correctly opts out. CalHousing remains XGBoost territory: spatial autocorrelation and threshold effects are better captured by axis-aligned splits than by pairwise cooperative geometry.

Classification AUC

DatasetGL-Ridge (default)GL-HPOXGB-HPO
Synthetic0.9400.9570.981
BreastCancer0.9930.9720.993

Default-parameter GL-Ridge matches XGBoost-HPO on BreastCancer (AUC 0.993 each).


API reference

GeoLinear (regressor)

ParameterTypeDefaultDescription
n_roundsint20Boosting rounds
learning_ratefloat0.1Shrinkage per round
y_weightfloat0.5HVRT outcome-blend (0 = geometry-only, 1 = y-driven)
base_learnerstr"ridge""ridge", "ols", or "lasso"
alphafloat1.0L2 regularisation within each partition
min_samples_partitionint5Minimum samples to fit a partition model
hvrt_n_partitionsint|NoneNoneTarget partitions (None = HVRT auto-tune)
hvrt_min_samples_leafint|NoneNoneHVRT min leaf size
hvrt_inner_roundsint1HVRT T-residual inner rounds per stage
partition_inner_roundsint1Base-learner rounds within each partition
refit_intervalint00 = fresh HVRT each round; k>0 = refit every k rounds (faster)
use_t_featureboolFalseAppend per-sample T-statistic (S²−Q) as an extra linear feature
use_coop_weightsboolFalseScale features by |corr(z_k, S−z_k)|² before linear fit
random_stateint42Seed (incremented per round for diverse partitionings)

GeoLinearClassifier accepts the same parameters.

Key methods

model.fit(X, y) # fitmodel.predict(X) # regression predictionsclf.predict_proba(X) # class probabilities, shape (n, 2)model.feature_importances(feature_names) # weighted mean |coef| across partitions/stagesmodel.stages_# list of (None, dict[partition_id, RidgeModel])

License

GNU Affero General Public License v3.0 or later. See LICENSE.

About

HVRT-informed linear modelling, comparable with XGBoost's accuracy while retaining interpretability.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages