Skip to content

Repository files navigation

Distribution-Inference

Reproducible distribution fitting, model discrimination, and likelihood-based uncertainty quantification across Python, R, MATLAB, and Fortran.

Distribution-Inference is research software for inferring a probability model from one or more observed samples. It fits domain-compatible candidate distributions by maximum likelihood, calibrates goodness-of-fit through a refitted parametric Bootstrap, ranks statistically admissible models, and quantifies parameter uncertainty with both profile-likelihood and Wald confidence intervals. Four independent language implementations execute one canonical statistical specification and emit a shared result schema.

Inference workflow

Figure 1. Statistical workflow implemented independently by all four computational backends.

Methodological scope

For observations $x_1,\ldots,x_n$ and a candidate density or probability mass function $f(x\mid\theta)$, parameters are estimated by

$$ \hat\theta=\arg\max_\theta \ell(\theta), \qquad \ell(\theta)=\sum_{i=1}^{n}\log f(x_i\mid\theta). $$

Each successful fit reports K-S, AIC, small-sample corrected AIC, and BIC:

$$ \mathrm{AIC}=2k-2\ell(\hat\theta),\qquad \mathrm{BIC}=k\log n-2\ell(\hat\theta). $$

Because the model parameters are estimated from the tested sample, the K-S p-value is calibrated by parametric Bootstrap with parameter re-estimation in every replicate. Continuous and discrete samples are analyzed separately; the discrete statistic is evaluated on both sides of each probability-mass jump. The recommended model is the minimum-BIC candidate among those not rejected at the configured K-S level. If every model is rejected, the lowest-BIC result is reported only as a diagnostic fallback.

Parameter uncertainty is evaluated at 90%, 95%, and 99% by default:

  • Profile-likelihood LR: nuisance parameters are re-optimized for every fixed value of the parameter of interest.
  • Wald: the observed information matrix is inverted on an unconstrained parameter scale and transformed back to the canonical scale.

Supported probability models

Data classDistributionCanonical parameters
Continuous, realNormalmu, sigma
Continuous, positiveLognormalmu_log, sigma_log
Continuous, nonnegativeExponentialrate
Continuous, positiveGammashape, scale
Continuous, nonnegativeWeibullshape, scale
Continuous, boundedBetaalpha, beta, known bounds
Continuous, realGumbel maximumloc, scale
Continuous, constrained supportGEVloc, scale, xi
Discrete, nonnegative integerPoissonlambda

The GEV convention is

$$ F(x)=\exp\left[-\left(1+\xi\frac{x-\mathrm{loc}}{\mathrm{scale}}\right)^{-1/\xi}\right], $$

with the Gumbel limit at $\xi\rightarrow0$. The canonical convention takes precedence over language-library aliases or sign conventions.

Cross-language benchmark

Cross-language benchmark

Figure 2. Model evidence, cross-language numerical spread, confidence intervals, and the selected-model fit for the bundled synthetic benchmark. The figure is generated from machine-readable outputs by scripts/generate_readme_figures.py.

On the bundled data, every implementation selects the same model:

Data setSelected modelScientific role
normal_sampleNormalreal-valued symmetric observations
positive_sampleLognormalpositive right-skewed observations
beta_sampleBetabounded observations on (0,1)
count_samplePoissonnonnegative event counts

Closed-form model likelihoods agree to numerical precision across the four backends. Iterative Beta, Gumbel, Gamma, Weibull, and GEV fits are compared with model-specific tolerances documented in the statistical specification.

Reproducible installation

Create the project environment from the repository root:

conda env create -f environment.yml
conda activate distribution-inference

The environment contains both Python and R dependencies. MATLAB and gfortran are detected but are not installed automatically. If Python and R already reside in separate environments, pass -PythonEnvironment and -REnvironment on PowerShell, or --python-env and --r-env on Shell.

Running an analysis

Windows PowerShell:

.\run_all.ps1 -Input examples\samples.csv `-DatasetConfig examples\datasets.csv `-Output results\analysis

Linux or a research server:

./run_all.sh \
--input examples/samples.csv \
--dataset-config examples/datasets.csv \
--output results/analysis

All repository paths remain relative to the project root. Production runs use 1000 Bootstrap replicates; use --bootstrap 99 or -Bootstrap 99 for development.

Input and output contract

The sample CSV contains one independent data set per numeric column. Blank cells are omitted column-wise. The optional data-set configuration defines:

dataset,data_type,beta_lower,beta_upper,distributions
beta_sample,continuous,0,1,beta;normal;gumbel

data_type accepts auto, continuous, or discrete. Automatic classification treats an entirely nonnegative integer sample as discrete; override it for rounded continuous measurements.

Each backend writes the same files below its result directory:

  • summary.json: settings, selection status, and data-set metadata.
  • candidates.csv: convergence, likelihood, K-S, AIC/AICc/BIC, and selection status.
  • parameters.csv: canonical MLEs and selected-model standard errors.
  • confidence_intervals.csv: LR and Wald limits with explicit status fields.

The shared reporting layer adds comparison.csv, comparison_differences.csv, report.md, fitted-density plots, and Q-Q plots.

Repository organization

src/python/ Python reference implementation and reporting
src/r/ Independent R implementation
src/matlab/ Independent MATLAB implementation
src/fortran/ Self-contained Fortran 2008 numerical implementation
spec/ Canonical parameterization and inferential rules
examples/ Synthetic benchmark samples and data-set configuration
tests/ Deterministic unit and command-line tests
scripts/ Setup, compilation, validation, and figure generation

Verification and development

pytest
./scripts/build_fortran.sh debug
./scripts/smoke_test.sh
python scripts/generate_readme_figures.py

PowerShell equivalents are provided for setup, Fortran compilation, and smoke testing. Generated results and compiler products are ignored by version control.

Statistical limitations

  • Bootstrap precision is finite and depends on the requested number of replicates; interpret p-values near the selection threshold cautiously.
  • Beta support bounds are treated as known, not estimated from sample minima and maxima.
  • Profile intervals can be unbounded or numerically unavailable near support boundaries; such cases are reported rather than silently replaced.
  • Automatic model selection is conditional on the supplied candidate set and does not establish that the data-generating mechanism is uniquely identified.
  • The software supports reproducible research workflows, but domain-specific assumptions and study design remain the analyst's responsibility.

About

Reproducible distribution fitting, model selection, and likelihood-based uncertainty quantification across Python, R, MATLAB, and Fortran.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages