Latest commit

History

83 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SSGS — Spectral-State Guided Synthesis

© 2025 Damien Davison & Michael Maillet & Sacha Davison

Recursive AI Devs

Licensed under the Apache License 2.0 psychcoherence@gmail.com OR therealmichaelmaillet@gmail.com

SSGS (Spectral-State Guided Synthesis) is a hybrid generative audio model that combines linear predictive coding (LPC), hidden Markov models (HMMs), spectral clustering, global graph search, and a physically-inspired Karplus–Strong excitation engine into a single synthesis framework.

Unlike classical systems that commit to a single paradigm (neural, statistical, or physical), SSGS deliberately smashes algorithms together and extracts the useful computational primitives from each.
The result is a model that:

  • learns spectral structure from real signals,
  • organizes that structure into Markov states,
  • decodes the optimal state trajectory over time via A* search,
  • generates excitation signals using string-model dynamics, and
  • reconstructs audio by filtering excitation through LPC envelopes.

SSGS is not meant to imitate any existing model family.
It is a new hybrid class: detailed like DSP, global like HMMs, and expressive like physical models.


Key Idea: “Algorithm Deconstruction”

SSGS is built using the same philosophy behind the Symbo project:

Break apart many algorithms, strip them to their mathematical essence, recombine only the primitives that are actually useful.

Instead of inheriting algorithms whole (e.g., a standard HMM, standard LPC vocoder, standard Karplus–Strong), SSGS takes:

  • LPC → spectral envelope representation
  • HMM → temporal clustering + statistical state transitions
  • EM → unsupervised parameter refinement
  • A* search → globally optimal state sequence
  • Karplus–Strong → natural resonance and excitation noise
  • Heuristics → spectral smoothness constraints
  • Graph theory → prune invalid or degenerate state structures

SSGS then recomposes these pieces into a single generative pipeline that didn’t exist before.

This approach is extremely flexible: swap the envelope model, swap the excitation, modify the search heuristic — the system keeps working.


Pipeline Overview

1. LPC Analysis

The training signal is segmented and analyzed with LPC to extract:

  • LPC coefficients
  • excitation residual
  • power envelope

These become the feature vectors for clustering.


2. Spectral Clustering via HMM Initialization

Frames are embedded into a feature space (typically derived from LPC spectra). States are initialized using k-means, then upgraded into a full HMM with:

  • initial state distribution
  • transition matrix
  • mean vectors and covariance matrices for each state

3. Expectation-Maximization (EM) Refinement

A custom EM implementation updates:

  • state responsibilities (γ)
  • pairwise transitions (ξ)
  • transition probabilities
  • Gaussian parameters (means, covariances)

Result: the HMM becomes a structured map of repeating spectral “modes.”


4. Graph Constraint Pruning

To prevent degenerate solutions, SSGS analyzes the transition graph:

  • identifies strongly connected components (SCCs)
  • removes invalid or isolated states
  • ensures state sequences remain musically plausible

5. Global State Decoding (A Search)*

Instead of using Viterbi (which is greedy and purely local), SSGS uses A*:

  • cost function = negative log-likelihood + spectral smoothness heuristic
  • ensures global consistency in the decoded state path
  • supports long-range structure better than classical decoding

6. Physically Inspired Excitation

The decoded state sequence modulates a Karplus–Strong string model, producing a dynamic excitation that is:

  • rich in overtones
  • noisy where appropriate
  • resonant and evolving

7. LPC Synthesis

Finally, excitation is passed through the LPC filters of each decoded state:

  • reconstructs spectral envelopes
  • restores formants and resonances
  • yields new but structurally coherent audio

This is how SSGS generates novel signals even from short or simple training data.


Model Export, Compression, and Checkpointing

After training you can persist the learned parameters in a compact archive:

ssgs.export_model("models/ssgs_baseline", use_compression=True, pack_covariances=True)
  • .npz output uses ZIP compression by default when use_compression=True.
  • Covariance matrices are serialized as lower-triangular slices to halve their stored size.
  • Set include_training_artifacts=True if you also need cached LPC and residual buffers for later analysis.
  • Reload the model with SpectralStateGuidedSynthesis.load_model(path).

For long-running training, SSGS also supports .safetensors checkpoints:

ssgs.save_checkpoint("models/checkpoint_epoch_5.safetensors")
ssgs=SpectralStateGuidedSynthesis.load_checkpoint("models/checkpoint_epoch_5.safetensors")

Checkpoints can optionally persist adaptive statistics so continuous learning can resume without reprocessing previous audio.


Generative Capability Levels

SSGS now offers enhanced generative capabilities through configurable state counts:

  • Standard (16 states): Original baseline configuration
  • Enhanced (34 states): 112% increase - recommended for most applications
  • Maximum (42 states): 163% increase - highest generative capacity
# Enhanced generative capabilities (default)ssgs=SpectralStateGuidedSynthesis(n_states=34)
# Maximum generative capabilitiesssgs=SpectralStateGuidedSynthesis(n_states=42)
# Original baselinessgs=SpectralStateGuidedSynthesis(n_states=16)

More states = more spectral patterns the model can learn and generate = richer, more diverse output.


Example Output

SSGS produces time-domain and spectral visualizations like this:

SSGS demo output

For a deeper comparison diagnostic, see ssgs_analysis.png.

The generated signal is not a copy —
it is a new trajectory through learned spectral states.


Quick Start

Install dependencies:

pip install -r requirements.txt

Run the demo training + generation script:

python test_ssgs.py

Run the enhanced capabilities demo (defaults to 34 states):

python demo_enhanced_capabilities.py

Train on folders of audio with checkpointing:

python train_on_folder.py --folders training_001 training_002

Adaptive Learning

SSGS can adapt to new audio using persistent statistics:

ssgs.adapt_to_audio(new_audio, sample_rate=16000, adaptation_rate=0.5, memory_blend=0.3)

Adaptation can be exported through .npz or .safetensors artifacts and reloaded later.

Feature Indexing & Retrieval

Use built-in nearest-neighbor search over frame features for exploration and analysis:

index=ssgs.build_feature_index(feature_space="lpc", normalize=True)
result=ssgs.query_similar_frames(frame_idx=42, k=8, feature_space="lpc")
vector_result=ssgs.search_by_feature_vector(
vector=index.tree.data[0],
k=5,
feature_space="lpc",
)

The returned FrameQueryResult includes indices, optional distances, and frame metadata.

Performance: Native Compiled Backend

The numeric hot loops are also implemented as a small, dependency-light C library (ssgs_native.c) and exposed through ssgs_native.py. ssgs.py automatically uses them when libssgs_native.so is loadable, and falls back to the existing SciPy/NumPy paths otherwise — so the package still works on a machine with no compiler installed.

Kernels covered: LPC (Levinson–Durbin), Gaussian emission (hand Cholesky + forward solve), EM forward–backward, and the smoothness-first Viterbi decode. The decode is the dominant synthesis cost (a pure-Python triple loop in the SciPy build); the compiled kernel is ~170× faster there, with identical output:

kernelSciPy/NumPyCspeedup
LPC (Levinson–Durbin)0.172 s0.172 s1.0×
Gaussian emission0.0014 s0.0012 s1.2×
decode (2000 frames)0.4998 s0.0029 s170×

The C kernels are numerically identical to ssgs.py (enforced by test_regression_guard.py). Build the library with:

make # produces libssgs_native.so (requires gcc)

A pure-NumPy reference implementation lives in ssgs_core.py and is used by the regression guard to prove equivalence even when the compiled library is absent.


Testing

Run individual checks as standalone scripts:

python test_ssgs.py
python test_checkpoints.py
python test_adaptive_persistence.py
python test_data_indexing.py
python test_gain_control.py
python test_generative_capabilities.py
python test_utils.py

Numeric-equivalence and performance gates:

python ssgs_core.py # pure-NumPy core self-test (equivalence oracle)
python ssgs_native.py # compiled-kernel verify + benchmark
python test_regression_guard.py # native load + core/pipeline equivalence

Requirements

numpy>=1.21.0
scipy>=1.7.0
matplotlib>=3.4.0
soundfile>=0.10.0
PyWavelets>=1.1.0
safetensors>=0.4.0

Optional: gcc and make to build the native backend (libssgs_native.so). If unavailable, SSGS falls back to the SciPy/NumPy paths automatically.

Project Structure

ssgs.py # Full model implementation
ssgs_core.py # Pure-NumPy reference kernels (equivalence oracle)
ssgs_native.c # Compiled C kernels (Levinson-Durbin, emission, EM, decode)
ssgs_native.py # ctypes wrapper + benchmark for the C kernels
train_on_folder.py # Folder-based training with checkpointing
demo_enhanced_capabilities.py # State-count showcase
example_checkpoints.py # Checkpoint usage example
example_fidelity.py # Fidelity-focused generation example
test_ssgs.py # Training + generation demo
test_checkpoints.py # Checkpoint save/load tests
test_adaptive_persistence.py # Adaptive learning tests
test_data_indexing.py # Feature index tests
test_gain_control.py # Gain control tests
test_generative_capabilities.py # State-count tests
test_regression_guard.py # Native load + core/pipeline equivalence guard
Makefile # Builds libssgs_native.so
requirements.txt # Dependencies
ssgs_demo.png # Example output
ssgs_analysis.png # Diagnostic output
README.md # This document \

Authors

Damien Davison & Michael Maillet & Sacha Davison Recursive AI Devs

We build hybrid-symbolic neural, statistical, and physical AI systems by algorithm decomposition — extracting the useful primitives and recombining them into new model classes.

If you use SSGS in research or production, please cite the authors.

License — Apache 2.0

This project is licensed under the Apache License, Version 2.0.

Copyright 2025 Damien Davison & Michael Maillet Recursive AI Devs

Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy at:

http://www.apache.org/licenses/LICENSE-2.0

NOTICE

This product includes original work by:
Damien Davison & Michael Maillet & Sacha Davison (Recursive AI Devs)
Additional details can be found in the project's LICENSE and source headers.

About

Spectral-State Guided Synthesis

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

83 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SSGS — Spectral-State Guided Synthesis

© 2025 Damien Davison & Michael Maillet & Sacha Davison

Recursive AI Devs

Licensed under the Apache License 2.0 psychcoherence@gmail.com OR therealmichaelmaillet@gmail.com

SSGS (Spectral-State Guided Synthesis) is a hybrid generative audio model that combines linear predictive coding (LPC), hidden Markov models (HMMs), spectral clustering, global graph search, and a physically-inspired Karplus–Strong excitation engine into a single synthesis framework.

Unlike classical systems that commit to a single paradigm (neural, statistical, or physical), SSGS deliberately smashes algorithms together and extracts the useful computational primitives from each.
The result is a model that:

  • learns spectral structure from real signals,
  • organizes that structure into Markov states,
  • decodes the optimal state trajectory over time via A* search,
  • generates excitation signals using string-model dynamics, and
  • reconstructs audio by filtering excitation through LPC envelopes.

SSGS is not meant to imitate any existing model family.
It is a new hybrid class: detailed like DSP, global like HMMs, and expressive like physical models.


Key Idea: “Algorithm Deconstruction”

SSGS is built using the same philosophy behind the Symbo project:

Break apart many algorithms, strip them to their mathematical essence, recombine only the primitives that are actually useful.

Instead of inheriting algorithms whole (e.g., a standard HMM, standard LPC vocoder, standard Karplus–Strong), SSGS takes:

  • LPC → spectral envelope representation
  • HMM → temporal clustering + statistical state transitions
  • EM → unsupervised parameter refinement
  • A* search → globally optimal state sequence
  • Karplus–Strong → natural resonance and excitation noise
  • Heuristics → spectral smoothness constraints
  • Graph theory → prune invalid or degenerate state structures

SSGS then recomposes these pieces into a single generative pipeline that didn’t exist before.

This approach is extremely flexible: swap the envelope model, swap the excitation, modify the search heuristic — the system keeps working.


Pipeline Overview

1. LPC Analysis

The training signal is segmented and analyzed with LPC to extract:

  • LPC coefficients
  • excitation residual
  • power envelope

These become the feature vectors for clustering.


2. Spectral Clustering via HMM Initialization

Frames are embedded into a feature space (typically derived from LPC spectra). States are initialized using k-means, then upgraded into a full HMM with:

  • initial state distribution
  • transition matrix
  • mean vectors and covariance matrices for each state

3. Expectation-Maximization (EM) Refinement

A custom EM implementation updates:

  • state responsibilities (γ)
  • pairwise transitions (ξ)
  • transition probabilities
  • Gaussian parameters (means, covariances)

Result: the HMM becomes a structured map of repeating spectral “modes.”


4. Graph Constraint Pruning

To prevent degenerate solutions, SSGS analyzes the transition graph:

  • identifies strongly connected components (SCCs)
  • removes invalid or isolated states
  • ensures state sequences remain musically plausible

5. Global State Decoding (A Search)*

Instead of using Viterbi (which is greedy and purely local), SSGS uses A*:

  • cost function = negative log-likelihood + spectral smoothness heuristic
  • ensures global consistency in the decoded state path
  • supports long-range structure better than classical decoding

6. Physically Inspired Excitation

The decoded state sequence modulates a Karplus–Strong string model, producing a dynamic excitation that is:

  • rich in overtones
  • noisy where appropriate
  • resonant and evolving

7. LPC Synthesis

Finally, excitation is passed through the LPC filters of each decoded state:

  • reconstructs spectral envelopes
  • restores formants and resonances
  • yields new but structurally coherent audio

This is how SSGS generates novel signals even from short or simple training data.


Model Export, Compression, and Checkpointing

After training you can persist the learned parameters in a compact archive:

ssgs.export_model("models/ssgs_baseline", use_compression=True, pack_covariances=True)
  • .npz output uses ZIP compression by default when use_compression=True.
  • Covariance matrices are serialized as lower-triangular slices to halve their stored size.
  • Set include_training_artifacts=True if you also need cached LPC and residual buffers for later analysis.
  • Reload the model with SpectralStateGuidedSynthesis.load_model(path).

For long-running training, SSGS also supports .safetensors checkpoints:

ssgs.save_checkpoint("models/checkpoint_epoch_5.safetensors")
ssgs=SpectralStateGuidedSynthesis.load_checkpoint("models/checkpoint_epoch_5.safetensors")

Checkpoints can optionally persist adaptive statistics so continuous learning can resume without reprocessing previous audio.


Generative Capability Levels

SSGS now offers enhanced generative capabilities through configurable state counts:

  • Standard (16 states): Original baseline configuration
  • Enhanced (34 states): 112% increase - recommended for most applications
  • Maximum (42 states): 163% increase - highest generative capacity
# Enhanced generative capabilities (default)ssgs=SpectralStateGuidedSynthesis(n_states=34)
# Maximum generative capabilitiesssgs=SpectralStateGuidedSynthesis(n_states=42)
# Original baselinessgs=SpectralStateGuidedSynthesis(n_states=16)

More states = more spectral patterns the model can learn and generate = richer, more diverse output.


Example Output

SSGS produces time-domain and spectral visualizations like this:

SSGS demo output

For a deeper comparison diagnostic, see ssgs_analysis.png.

The generated signal is not a copy —
it is a new trajectory through learned spectral states.


Quick Start

Install dependencies:

pip install -r requirements.txt

Run the demo training + generation script:

python test_ssgs.py

Run the enhanced capabilities demo (defaults to 34 states):

python demo_enhanced_capabilities.py

Train on folders of audio with checkpointing:

python train_on_folder.py --folders training_001 training_002

Adaptive Learning

SSGS can adapt to new audio using persistent statistics:

ssgs.adapt_to_audio(new_audio, sample_rate=16000, adaptation_rate=0.5, memory_blend=0.3)

Adaptation can be exported through .npz or .safetensors artifacts and reloaded later.

Feature Indexing & Retrieval

Use built-in nearest-neighbor search over frame features for exploration and analysis:

index=ssgs.build_feature_index(feature_space="lpc", normalize=True)
result=ssgs.query_similar_frames(frame_idx=42, k=8, feature_space="lpc")
vector_result=ssgs.search_by_feature_vector(
vector=index.tree.data[0],
k=5,
feature_space="lpc",
)

The returned FrameQueryResult includes indices, optional distances, and frame metadata.

Performance: Native Compiled Backend

The numeric hot loops are also implemented as a small, dependency-light C library (ssgs_native.c) and exposed through ssgs_native.py. ssgs.py automatically uses them when libssgs_native.so is loadable, and falls back to the existing SciPy/NumPy paths otherwise — so the package still works on a machine with no compiler installed.

Kernels covered: LPC (Levinson–Durbin), Gaussian emission (hand Cholesky + forward solve), EM forward–backward, and the smoothness-first Viterbi decode. The decode is the dominant synthesis cost (a pure-Python triple loop in the SciPy build); the compiled kernel is ~170× faster there, with identical output:

kernelSciPy/NumPyCspeedup
LPC (Levinson–Durbin)0.172 s0.172 s1.0×
Gaussian emission0.0014 s0.0012 s1.2×
decode (2000 frames)0.4998 s0.0029 s170×

The C kernels are numerically identical to ssgs.py (enforced by test_regression_guard.py). Build the library with:

make # produces libssgs_native.so (requires gcc)

A pure-NumPy reference implementation lives in ssgs_core.py and is used by the regression guard to prove equivalence even when the compiled library is absent.


Testing

Run individual checks as standalone scripts:

python test_ssgs.py
python test_checkpoints.py
python test_adaptive_persistence.py
python test_data_indexing.py
python test_gain_control.py
python test_generative_capabilities.py
python test_utils.py

Numeric-equivalence and performance gates:

python ssgs_core.py # pure-NumPy core self-test (equivalence oracle)
python ssgs_native.py # compiled-kernel verify + benchmark
python test_regression_guard.py # native load + core/pipeline equivalence

Requirements

numpy>=1.21.0
scipy>=1.7.0
matplotlib>=3.4.0
soundfile>=0.10.0
PyWavelets>=1.1.0
safetensors>=0.4.0

Optional: gcc and make to build the native backend (libssgs_native.so). If unavailable, SSGS falls back to the SciPy/NumPy paths automatically.

Project Structure

ssgs.py # Full model implementation
ssgs_core.py # Pure-NumPy reference kernels (equivalence oracle)
ssgs_native.c # Compiled C kernels (Levinson-Durbin, emission, EM, decode)
ssgs_native.py # ctypes wrapper + benchmark for the C kernels
train_on_folder.py # Folder-based training with checkpointing
demo_enhanced_capabilities.py # State-count showcase
example_checkpoints.py # Checkpoint usage example
example_fidelity.py # Fidelity-focused generation example
test_ssgs.py # Training + generation demo
test_checkpoints.py # Checkpoint save/load tests
test_adaptive_persistence.py # Adaptive learning tests
test_data_indexing.py # Feature index tests
test_gain_control.py # Gain control tests
test_generative_capabilities.py # State-count tests
test_regression_guard.py # Native load + core/pipeline equivalence guard
Makefile # Builds libssgs_native.so
requirements.txt # Dependencies
ssgs_demo.png # Example output
ssgs_analysis.png # Diagnostic output
README.md # This document \

Authors

Damien Davison & Michael Maillet & Sacha Davison Recursive AI Devs

We build hybrid-symbolic neural, statistical, and physical AI systems by algorithm decomposition — extracting the useful primitives and recombining them into new model classes.

If you use SSGS in research or production, please cite the authors.

License — Apache 2.0

This project is licensed under the Apache License, Version 2.0.

Copyright 2025 Damien Davison & Michael Maillet Recursive AI Devs

Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy at:

http://www.apache.org/licenses/LICENSE-2.0

NOTICE

This product includes original work by:
Damien Davison & Michael Maillet & Sacha Davison (Recursive AI Devs)
Additional details can be found in the project's LICENSE and source headers.

About

Spectral-State Guided Synthesis

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

83 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SSGS — Spectral-State Guided Synthesis

© 2025 Damien Davison & Michael Maillet & Sacha Davison

Recursive AI Devs

Licensed under the Apache License 2.0 psychcoherence@gmail.com OR therealmichaelmaillet@gmail.com

SSGS (Spectral-State Guided Synthesis) is a hybrid generative audio model that combines linear predictive coding (LPC), hidden Markov models (HMMs), spectral clustering, global graph search, and a physically-inspired Karplus–Strong excitation engine into a single synthesis framework.

Unlike classical systems that commit to a single paradigm (neural, statistical, or physical), SSGS deliberately smashes algorithms together and extracts the useful computational primitives from each.
The result is a model that:

  • learns spectral structure from real signals,
  • organizes that structure into Markov states,
  • decodes the optimal state trajectory over time via A* search,
  • generates excitation signals using string-model dynamics, and
  • reconstructs audio by filtering excitation through LPC envelopes.

SSGS is not meant to imitate any existing model family.
It is a new hybrid class: detailed like DSP, global like HMMs, and expressive like physical models.


Key Idea: “Algorithm Deconstruction”

SSGS is built using the same philosophy behind the Symbo project:

Break apart many algorithms, strip them to their mathematical essence, recombine only the primitives that are actually useful.

Instead of inheriting algorithms whole (e.g., a standard HMM, standard LPC vocoder, standard Karplus–Strong), SSGS takes:

  • LPC → spectral envelope representation
  • HMM → temporal clustering + statistical state transitions
  • EM → unsupervised parameter refinement
  • A* search → globally optimal state sequence
  • Karplus–Strong → natural resonance and excitation noise
  • Heuristics → spectral smoothness constraints
  • Graph theory → prune invalid or degenerate state structures

SSGS then recomposes these pieces into a single generative pipeline that didn’t exist before.

This approach is extremely flexible: swap the envelope model, swap the excitation, modify the search heuristic — the system keeps working.


Pipeline Overview

1. LPC Analysis

The training signal is segmented and analyzed with LPC to extract:

  • LPC coefficients
  • excitation residual
  • power envelope

These become the feature vectors for clustering.


2. Spectral Clustering via HMM Initialization

Frames are embedded into a feature space (typically derived from LPC spectra). States are initialized using k-means, then upgraded into a full HMM with:

  • initial state distribution
  • transition matrix
  • mean vectors and covariance matrices for each state

3. Expectation-Maximization (EM) Refinement

A custom EM implementation updates:

  • state responsibilities (γ)
  • pairwise transitions (ξ)
  • transition probabilities
  • Gaussian parameters (means, covariances)

Result: the HMM becomes a structured map of repeating spectral “modes.”


4. Graph Constraint Pruning

To prevent degenerate solutions, SSGS analyzes the transition graph:

  • identifies strongly connected components (SCCs)
  • removes invalid or isolated states
  • ensures state sequences remain musically plausible

5. Global State Decoding (A Search)*

Instead of using Viterbi (which is greedy and purely local), SSGS uses A*:

  • cost function = negative log-likelihood + spectral smoothness heuristic
  • ensures global consistency in the decoded state path
  • supports long-range structure better than classical decoding

6. Physically Inspired Excitation

The decoded state sequence modulates a Karplus–Strong string model, producing a dynamic excitation that is:

  • rich in overtones
  • noisy where appropriate
  • resonant and evolving

7. LPC Synthesis

Finally, excitation is passed through the LPC filters of each decoded state:

  • reconstructs spectral envelopes
  • restores formants and resonances
  • yields new but structurally coherent audio

This is how SSGS generates novel signals even from short or simple training data.


Model Export, Compression, and Checkpointing

After training you can persist the learned parameters in a compact archive:

ssgs.export_model("models/ssgs_baseline", use_compression=True, pack_covariances=True)
  • .npz output uses ZIP compression by default when use_compression=True.
  • Covariance matrices are serialized as lower-triangular slices to halve their stored size.
  • Set include_training_artifacts=True if you also need cached LPC and residual buffers for later analysis.
  • Reload the model with SpectralStateGuidedSynthesis.load_model(path).

For long-running training, SSGS also supports .safetensors checkpoints:

ssgs.save_checkpoint("models/checkpoint_epoch_5.safetensors")
ssgs=SpectralStateGuidedSynthesis.load_checkpoint("models/checkpoint_epoch_5.safetensors")

Checkpoints can optionally persist adaptive statistics so continuous learning can resume without reprocessing previous audio.


Generative Capability Levels

SSGS now offers enhanced generative capabilities through configurable state counts:

  • Standard (16 states): Original baseline configuration
  • Enhanced (34 states): 112% increase - recommended for most applications
  • Maximum (42 states): 163% increase - highest generative capacity
# Enhanced generative capabilities (default)ssgs=SpectralStateGuidedSynthesis(n_states=34)
# Maximum generative capabilitiesssgs=SpectralStateGuidedSynthesis(n_states=42)
# Original baselinessgs=SpectralStateGuidedSynthesis(n_states=16)

More states = more spectral patterns the model can learn and generate = richer, more diverse output.


Example Output

SSGS produces time-domain and spectral visualizations like this:

SSGS demo output

For a deeper comparison diagnostic, see ssgs_analysis.png.

The generated signal is not a copy —
it is a new trajectory through learned spectral states.


Quick Start

Install dependencies:

pip install -r requirements.txt

Run the demo training + generation script:

python test_ssgs.py

Run the enhanced capabilities demo (defaults to 34 states):

python demo_enhanced_capabilities.py

Train on folders of audio with checkpointing:

python train_on_folder.py --folders training_001 training_002

Adaptive Learning

SSGS can adapt to new audio using persistent statistics:

ssgs.adapt_to_audio(new_audio, sample_rate=16000, adaptation_rate=0.5, memory_blend=0.3)

Adaptation can be exported through .npz or .safetensors artifacts and reloaded later.

Feature Indexing & Retrieval

Use built-in nearest-neighbor search over frame features for exploration and analysis:

index=ssgs.build_feature_index(feature_space="lpc", normalize=True)
result=ssgs.query_similar_frames(frame_idx=42, k=8, feature_space="lpc")
vector_result=ssgs.search_by_feature_vector(
vector=index.tree.data[0],
k=5,
feature_space="lpc",
)

The returned FrameQueryResult includes indices, optional distances, and frame metadata.

Performance: Native Compiled Backend

The numeric hot loops are also implemented as a small, dependency-light C library (ssgs_native.c) and exposed through ssgs_native.py. ssgs.py automatically uses them when libssgs_native.so is loadable, and falls back to the existing SciPy/NumPy paths otherwise — so the package still works on a machine with no compiler installed.

Kernels covered: LPC (Levinson–Durbin), Gaussian emission (hand Cholesky + forward solve), EM forward–backward, and the smoothness-first Viterbi decode. The decode is the dominant synthesis cost (a pure-Python triple loop in the SciPy build); the compiled kernel is ~170× faster there, with identical output:

kernelSciPy/NumPyCspeedup
LPC (Levinson–Durbin)0.172 s0.172 s1.0×
Gaussian emission0.0014 s0.0012 s1.2×
decode (2000 frames)0.4998 s0.0029 s170×

The C kernels are numerically identical to ssgs.py (enforced by test_regression_guard.py). Build the library with:

make # produces libssgs_native.so (requires gcc)

A pure-NumPy reference implementation lives in ssgs_core.py and is used by the regression guard to prove equivalence even when the compiled library is absent.


Testing

Run individual checks as standalone scripts:

python test_ssgs.py
python test_checkpoints.py
python test_adaptive_persistence.py
python test_data_indexing.py
python test_gain_control.py
python test_generative_capabilities.py
python test_utils.py

Numeric-equivalence and performance gates:

python ssgs_core.py # pure-NumPy core self-test (equivalence oracle)
python ssgs_native.py # compiled-kernel verify + benchmark
python test_regression_guard.py # native load + core/pipeline equivalence

Requirements

numpy>=1.21.0
scipy>=1.7.0
matplotlib>=3.4.0
soundfile>=0.10.0
PyWavelets>=1.1.0
safetensors>=0.4.0

Optional: gcc and make to build the native backend (libssgs_native.so). If unavailable, SSGS falls back to the SciPy/NumPy paths automatically.

Project Structure

ssgs.py # Full model implementation
ssgs_core.py # Pure-NumPy reference kernels (equivalence oracle)
ssgs_native.c # Compiled C kernels (Levinson-Durbin, emission, EM, decode)
ssgs_native.py # ctypes wrapper + benchmark for the C kernels
train_on_folder.py # Folder-based training with checkpointing
demo_enhanced_capabilities.py # State-count showcase
example_checkpoints.py # Checkpoint usage example
example_fidelity.py # Fidelity-focused generation example
test_ssgs.py # Training + generation demo
test_checkpoints.py # Checkpoint save/load tests
test_adaptive_persistence.py # Adaptive learning tests
test_data_indexing.py # Feature index tests
test_gain_control.py # Gain control tests
test_generative_capabilities.py # State-count tests
test_regression_guard.py # Native load + core/pipeline equivalence guard
Makefile # Builds libssgs_native.so
requirements.txt # Dependencies
ssgs_demo.png # Example output
ssgs_analysis.png # Diagnostic output
README.md # This document \

Authors

Damien Davison & Michael Maillet & Sacha Davison Recursive AI Devs

We build hybrid-symbolic neural, statistical, and physical AI systems by algorithm decomposition — extracting the useful primitives and recombining them into new model classes.

If you use SSGS in research or production, please cite the authors.

License — Apache 2.0

This project is licensed under the Apache License, Version 2.0.

Copyright 2025 Damien Davison & Michael Maillet Recursive AI Devs

Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy at:

http://www.apache.org/licenses/LICENSE-2.0

NOTICE

This product includes original work by:
Damien Davison & Michael Maillet & Sacha Davison (Recursive AI Devs)
Additional details can be found in the project's LICENSE and source headers.

About

Spectral-State Guided Synthesis

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

83 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SSGS — Spectral-State Guided Synthesis

© 2025 Damien Davison & Michael Maillet & Sacha Davison

Recursive AI Devs

Licensed under the Apache License 2.0 psychcoherence@gmail.com OR therealmichaelmaillet@gmail.com

SSGS (Spectral-State Guided Synthesis) is a hybrid generative audio model that combines linear predictive coding (LPC), hidden Markov models (HMMs), spectral clustering, global graph search, and a physically-inspired Karplus–Strong excitation engine into a single synthesis framework.

Unlike classical systems that commit to a single paradigm (neural, statistical, or physical), SSGS deliberately smashes algorithms together and extracts the useful computational primitives from each.
The result is a model that:

  • learns spectral structure from real signals,
  • organizes that structure into Markov states,
  • decodes the optimal state trajectory over time via A* search,
  • generates excitation signals using string-model dynamics, and
  • reconstructs audio by filtering excitation through LPC envelopes.

SSGS is not meant to imitate any existing model family.
It is a new hybrid class: detailed like DSP, global like HMMs, and expressive like physical models.


Key Idea: “Algorithm Deconstruction”

SSGS is built using the same philosophy behind the Symbo project:

Break apart many algorithms, strip them to their mathematical essence, recombine only the primitives that are actually useful.

Instead of inheriting algorithms whole (e.g., a standard HMM, standard LPC vocoder, standard Karplus–Strong), SSGS takes:

  • LPC → spectral envelope representation
  • HMM → temporal clustering + statistical state transitions
  • EM → unsupervised parameter refinement
  • A* search → globally optimal state sequence
  • Karplus–Strong → natural resonance and excitation noise
  • Heuristics → spectral smoothness constraints
  • Graph theory → prune invalid or degenerate state structures

SSGS then recomposes these pieces into a single generative pipeline that didn’t exist before.

This approach is extremely flexible: swap the envelope model, swap the excitation, modify the search heuristic — the system keeps working.


Pipeline Overview

1. LPC Analysis

The training signal is segmented and analyzed with LPC to extract:

  • LPC coefficients
  • excitation residual
  • power envelope

These become the feature vectors for clustering.


2. Spectral Clustering via HMM Initialization

Frames are embedded into a feature space (typically derived from LPC spectra). States are initialized using k-means, then upgraded into a full HMM with:

  • initial state distribution
  • transition matrix
  • mean vectors and covariance matrices for each state

3. Expectation-Maximization (EM) Refinement

A custom EM implementation updates:

  • state responsibilities (γ)
  • pairwise transitions (ξ)
  • transition probabilities
  • Gaussian parameters (means, covariances)

Result: the HMM becomes a structured map of repeating spectral “modes.”


4. Graph Constraint Pruning

To prevent degenerate solutions, SSGS analyzes the transition graph:

  • identifies strongly connected components (SCCs)
  • removes invalid or isolated states
  • ensures state sequences remain musically plausible

5. Global State Decoding (A Search)*

Instead of using Viterbi (which is greedy and purely local), SSGS uses A*:

  • cost function = negative log-likelihood + spectral smoothness heuristic
  • ensures global consistency in the decoded state path
  • supports long-range structure better than classical decoding

6. Physically Inspired Excitation

The decoded state sequence modulates a Karplus–Strong string model, producing a dynamic excitation that is:

  • rich in overtones
  • noisy where appropriate
  • resonant and evolving

7. LPC Synthesis

Finally, excitation is passed through the LPC filters of each decoded state:

  • reconstructs spectral envelopes
  • restores formants and resonances
  • yields new but structurally coherent audio

This is how SSGS generates novel signals even from short or simple training data.


Model Export, Compression, and Checkpointing

After training you can persist the learned parameters in a compact archive:

ssgs.export_model("models/ssgs_baseline", use_compression=True, pack_covariances=True)
  • .npz output uses ZIP compression by default when use_compression=True.
  • Covariance matrices are serialized as lower-triangular slices to halve their stored size.
  • Set include_training_artifacts=True if you also need cached LPC and residual buffers for later analysis.
  • Reload the model with SpectralStateGuidedSynthesis.load_model(path).

For long-running training, SSGS also supports .safetensors checkpoints:

ssgs.save_checkpoint("models/checkpoint_epoch_5.safetensors")
ssgs=SpectralStateGuidedSynthesis.load_checkpoint("models/checkpoint_epoch_5.safetensors")

Checkpoints can optionally persist adaptive statistics so continuous learning can resume without reprocessing previous audio.


Generative Capability Levels

SSGS now offers enhanced generative capabilities through configurable state counts:

  • Standard (16 states): Original baseline configuration
  • Enhanced (34 states): 112% increase - recommended for most applications
  • Maximum (42 states): 163% increase - highest generative capacity
# Enhanced generative capabilities (default)ssgs=SpectralStateGuidedSynthesis(n_states=34)
# Maximum generative capabilitiesssgs=SpectralStateGuidedSynthesis(n_states=42)
# Original baselinessgs=SpectralStateGuidedSynthesis(n_states=16)

More states = more spectral patterns the model can learn and generate = richer, more diverse output.


Example Output

SSGS produces time-domain and spectral visualizations like this:

SSGS demo output

For a deeper comparison diagnostic, see ssgs_analysis.png.

The generated signal is not a copy —
it is a new trajectory through learned spectral states.


Quick Start

Install dependencies:

pip install -r requirements.txt

Run the demo training + generation script:

python test_ssgs.py

Run the enhanced capabilities demo (defaults to 34 states):

python demo_enhanced_capabilities.py

Train on folders of audio with checkpointing:

python train_on_folder.py --folders training_001 training_002

Adaptive Learning

SSGS can adapt to new audio using persistent statistics:

ssgs.adapt_to_audio(new_audio, sample_rate=16000, adaptation_rate=0.5, memory_blend=0.3)

Adaptation can be exported through .npz or .safetensors artifacts and reloaded later.

Feature Indexing & Retrieval

Use built-in nearest-neighbor search over frame features for exploration and analysis:

index=ssgs.build_feature_index(feature_space="lpc", normalize=True)
result=ssgs.query_similar_frames(frame_idx=42, k=8, feature_space="lpc")
vector_result=ssgs.search_by_feature_vector(
vector=index.tree.data[0],
k=5,
feature_space="lpc",
)

The returned FrameQueryResult includes indices, optional distances, and frame metadata.

Performance: Native Compiled Backend

The numeric hot loops are also implemented as a small, dependency-light C library (ssgs_native.c) and exposed through ssgs_native.py. ssgs.py automatically uses them when libssgs_native.so is loadable, and falls back to the existing SciPy/NumPy paths otherwise — so the package still works on a machine with no compiler installed.

Kernels covered: LPC (Levinson–Durbin), Gaussian emission (hand Cholesky + forward solve), EM forward–backward, and the smoothness-first Viterbi decode. The decode is the dominant synthesis cost (a pure-Python triple loop in the SciPy build); the compiled kernel is ~170× faster there, with identical output:

kernelSciPy/NumPyCspeedup
LPC (Levinson–Durbin)0.172 s0.172 s1.0×
Gaussian emission0.0014 s0.0012 s1.2×
decode (2000 frames)0.4998 s0.0029 s170×

The C kernels are numerically identical to ssgs.py (enforced by test_regression_guard.py). Build the library with:

make # produces libssgs_native.so (requires gcc)

A pure-NumPy reference implementation lives in ssgs_core.py and is used by the regression guard to prove equivalence even when the compiled library is absent.


Testing

Run individual checks as standalone scripts:

python test_ssgs.py
python test_checkpoints.py
python test_adaptive_persistence.py
python test_data_indexing.py
python test_gain_control.py
python test_generative_capabilities.py
python test_utils.py

Numeric-equivalence and performance gates:

python ssgs_core.py # pure-NumPy core self-test (equivalence oracle)
python ssgs_native.py # compiled-kernel verify + benchmark
python test_regression_guard.py # native load + core/pipeline equivalence

Requirements

numpy>=1.21.0
scipy>=1.7.0
matplotlib>=3.4.0
soundfile>=0.10.0
PyWavelets>=1.1.0
safetensors>=0.4.0

Optional: gcc and make to build the native backend (libssgs_native.so). If unavailable, SSGS falls back to the SciPy/NumPy paths automatically.

Project Structure

ssgs.py # Full model implementation
ssgs_core.py # Pure-NumPy reference kernels (equivalence oracle)
ssgs_native.c # Compiled C kernels (Levinson-Durbin, emission, EM, decode)
ssgs_native.py # ctypes wrapper + benchmark for the C kernels
train_on_folder.py # Folder-based training with checkpointing
demo_enhanced_capabilities.py # State-count showcase
example_checkpoints.py # Checkpoint usage example
example_fidelity.py # Fidelity-focused generation example
test_ssgs.py # Training + generation demo
test_checkpoints.py # Checkpoint save/load tests
test_adaptive_persistence.py # Adaptive learning tests
test_data_indexing.py # Feature index tests
test_gain_control.py # Gain control tests
test_generative_capabilities.py # State-count tests
test_regression_guard.py # Native load + core/pipeline equivalence guard
Makefile # Builds libssgs_native.so
requirements.txt # Dependencies
ssgs_demo.png # Example output
ssgs_analysis.png # Diagnostic output
README.md # This document \

Authors

Damien Davison & Michael Maillet & Sacha Davison Recursive AI Devs

We build hybrid-symbolic neural, statistical, and physical AI systems by algorithm decomposition — extracting the useful primitives and recombining them into new model classes.

If you use SSGS in research or production, please cite the authors.

License — Apache 2.0

This project is licensed under the Apache License, Version 2.0.

Copyright 2025 Damien Davison & Michael Maillet Recursive AI Devs

Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy at:

http://www.apache.org/licenses/LICENSE-2.0

NOTICE

This product includes original work by:
Damien Davison & Michael Maillet & Sacha Davison (Recursive AI Devs)
Additional details can be found in the project's LICENSE and source headers.

About

Spectral-State Guided Synthesis

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

83 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SSGS — Spectral-State Guided Synthesis

© 2025 Damien Davison & Michael Maillet & Sacha Davison

Recursive AI Devs

Licensed under the Apache License 2.0 psychcoherence@gmail.com OR therealmichaelmaillet@gmail.com

SSGS (Spectral-State Guided Synthesis) is a hybrid generative audio model that combines linear predictive coding (LPC), hidden Markov models (HMMs), spectral clustering, global graph search, and a physically-inspired Karplus–Strong excitation engine into a single synthesis framework.

Unlike classical systems that commit to a single paradigm (neural, statistical, or physical), SSGS deliberately smashes algorithms together and extracts the useful computational primitives from each.
The result is a model that:

  • learns spectral structure from real signals,
  • organizes that structure into Markov states,
  • decodes the optimal state trajectory over time via A* search,
  • generates excitation signals using string-model dynamics, and
  • reconstructs audio by filtering excitation through LPC envelopes.

SSGS is not meant to imitate any existing model family.
It is a new hybrid class: detailed like DSP, global like HMMs, and expressive like physical models.


Key Idea: “Algorithm Deconstruction”

SSGS is built using the same philosophy behind the Symbo project:

Break apart many algorithms, strip them to their mathematical essence, recombine only the primitives that are actually useful.

Instead of inheriting algorithms whole (e.g., a standard HMM, standard LPC vocoder, standard Karplus–Strong), SSGS takes:

  • LPC → spectral envelope representation
  • HMM → temporal clustering + statistical state transitions
  • EM → unsupervised parameter refinement
  • A* search → globally optimal state sequence
  • Karplus–Strong → natural resonance and excitation noise
  • Heuristics → spectral smoothness constraints
  • Graph theory → prune invalid or degenerate state structures

SSGS then recomposes these pieces into a single generative pipeline that didn’t exist before.

This approach is extremely flexible: swap the envelope model, swap the excitation, modify the search heuristic — the system keeps working.


Pipeline Overview

1. LPC Analysis

The training signal is segmented and analyzed with LPC to extract:

  • LPC coefficients
  • excitation residual
  • power envelope

These become the feature vectors for clustering.


2. Spectral Clustering via HMM Initialization

Frames are embedded into a feature space (typically derived from LPC spectra). States are initialized using k-means, then upgraded into a full HMM with:

  • initial state distribution
  • transition matrix
  • mean vectors and covariance matrices for each state

3. Expectation-Maximization (EM) Refinement

A custom EM implementation updates:

  • state responsibilities (γ)
  • pairwise transitions (ξ)
  • transition probabilities
  • Gaussian parameters (means, covariances)

Result: the HMM becomes a structured map of repeating spectral “modes.”


4. Graph Constraint Pruning

To prevent degenerate solutions, SSGS analyzes the transition graph:

  • identifies strongly connected components (SCCs)
  • removes invalid or isolated states
  • ensures state sequences remain musically plausible

5. Global State Decoding (A Search)*

Instead of using Viterbi (which is greedy and purely local), SSGS uses A*:

  • cost function = negative log-likelihood + spectral smoothness heuristic
  • ensures global consistency in the decoded state path
  • supports long-range structure better than classical decoding

6. Physically Inspired Excitation

The decoded state sequence modulates a Karplus–Strong string model, producing a dynamic excitation that is:

  • rich in overtones
  • noisy where appropriate
  • resonant and evolving

7. LPC Synthesis

Finally, excitation is passed through the LPC filters of each decoded state:

  • reconstructs spectral envelopes
  • restores formants and resonances
  • yields new but structurally coherent audio

This is how SSGS generates novel signals even from short or simple training data.


Model Export, Compression, and Checkpointing

After training you can persist the learned parameters in a compact archive:

ssgs.export_model("models/ssgs_baseline", use_compression=True, pack_covariances=True)
  • .npz output uses ZIP compression by default when use_compression=True.
  • Covariance matrices are serialized as lower-triangular slices to halve their stored size.
  • Set include_training_artifacts=True if you also need cached LPC and residual buffers for later analysis.
  • Reload the model with SpectralStateGuidedSynthesis.load_model(path).

For long-running training, SSGS also supports .safetensors checkpoints:

ssgs.save_checkpoint("models/checkpoint_epoch_5.safetensors")
ssgs=SpectralStateGuidedSynthesis.load_checkpoint("models/checkpoint_epoch_5.safetensors")

Checkpoints can optionally persist adaptive statistics so continuous learning can resume without reprocessing previous audio.


Generative Capability Levels

SSGS now offers enhanced generative capabilities through configurable state counts:

  • Standard (16 states): Original baseline configuration
  • Enhanced (34 states): 112% increase - recommended for most applications
  • Maximum (42 states): 163% increase - highest generative capacity
# Enhanced generative capabilities (default)ssgs=SpectralStateGuidedSynthesis(n_states=34)
# Maximum generative capabilitiesssgs=SpectralStateGuidedSynthesis(n_states=42)
# Original baselinessgs=SpectralStateGuidedSynthesis(n_states=16)

More states = more spectral patterns the model can learn and generate = richer, more diverse output.


Example Output

SSGS produces time-domain and spectral visualizations like this:

SSGS demo output

For a deeper comparison diagnostic, see ssgs_analysis.png.

The generated signal is not a copy —
it is a new trajectory through learned spectral states.


Quick Start

Install dependencies:

pip install -r requirements.txt

Run the demo training + generation script:

python test_ssgs.py

Run the enhanced capabilities demo (defaults to 34 states):

python demo_enhanced_capabilities.py

Train on folders of audio with checkpointing:

python train_on_folder.py --folders training_001 training_002

Adaptive Learning

SSGS can adapt to new audio using persistent statistics:

ssgs.adapt_to_audio(new_audio, sample_rate=16000, adaptation_rate=0.5, memory_blend=0.3)

Adaptation can be exported through .npz or .safetensors artifacts and reloaded later.

Feature Indexing & Retrieval

Use built-in nearest-neighbor search over frame features for exploration and analysis:

index=ssgs.build_feature_index(feature_space="lpc", normalize=True)
result=ssgs.query_similar_frames(frame_idx=42, k=8, feature_space="lpc")
vector_result=ssgs.search_by_feature_vector(
vector=index.tree.data[0],
k=5,
feature_space="lpc",
)

The returned FrameQueryResult includes indices, optional distances, and frame metadata.

Performance: Native Compiled Backend

The numeric hot loops are also implemented as a small, dependency-light C library (ssgs_native.c) and exposed through ssgs_native.py. ssgs.py automatically uses them when libssgs_native.so is loadable, and falls back to the existing SciPy/NumPy paths otherwise — so the package still works on a machine with no compiler installed.

Kernels covered: LPC (Levinson–Durbin), Gaussian emission (hand Cholesky + forward solve), EM forward–backward, and the smoothness-first Viterbi decode. The decode is the dominant synthesis cost (a pure-Python triple loop in the SciPy build); the compiled kernel is ~170× faster there, with identical output:

kernelSciPy/NumPyCspeedup
LPC (Levinson–Durbin)0.172 s0.172 s1.0×
Gaussian emission0.0014 s0.0012 s1.2×
decode (2000 frames)0.4998 s0.0029 s170×

The C kernels are numerically identical to ssgs.py (enforced by test_regression_guard.py). Build the library with:

make # produces libssgs_native.so (requires gcc)

A pure-NumPy reference implementation lives in ssgs_core.py and is used by the regression guard to prove equivalence even when the compiled library is absent.


Testing

Run individual checks as standalone scripts:

python test_ssgs.py
python test_checkpoints.py
python test_adaptive_persistence.py
python test_data_indexing.py
python test_gain_control.py
python test_generative_capabilities.py
python test_utils.py

Numeric-equivalence and performance gates:

python ssgs_core.py # pure-NumPy core self-test (equivalence oracle)
python ssgs_native.py # compiled-kernel verify + benchmark
python test_regression_guard.py # native load + core/pipeline equivalence

Requirements

numpy>=1.21.0
scipy>=1.7.0
matplotlib>=3.4.0
soundfile>=0.10.0
PyWavelets>=1.1.0
safetensors>=0.4.0

Optional: gcc and make to build the native backend (libssgs_native.so). If unavailable, SSGS falls back to the SciPy/NumPy paths automatically.

Project Structure

ssgs.py # Full model implementation
ssgs_core.py # Pure-NumPy reference kernels (equivalence oracle)
ssgs_native.c # Compiled C kernels (Levinson-Durbin, emission, EM, decode)
ssgs_native.py # ctypes wrapper + benchmark for the C kernels
train_on_folder.py # Folder-based training with checkpointing
demo_enhanced_capabilities.py # State-count showcase
example_checkpoints.py # Checkpoint usage example
example_fidelity.py # Fidelity-focused generation example
test_ssgs.py # Training + generation demo
test_checkpoints.py # Checkpoint save/load tests
test_adaptive_persistence.py # Adaptive learning tests
test_data_indexing.py # Feature index tests
test_gain_control.py # Gain control tests
test_generative_capabilities.py # State-count tests
test_regression_guard.py # Native load + core/pipeline equivalence guard
Makefile # Builds libssgs_native.so
requirements.txt # Dependencies
ssgs_demo.png # Example output
ssgs_analysis.png # Diagnostic output
README.md # This document \

Authors

Damien Davison & Michael Maillet & Sacha Davison Recursive AI Devs

We build hybrid-symbolic neural, statistical, and physical AI systems by algorithm decomposition — extracting the useful primitives and recombining them into new model classes.

If you use SSGS in research or production, please cite the authors.

License — Apache 2.0

This project is licensed under the Apache License, Version 2.0.

Copyright 2025 Damien Davison & Michael Maillet Recursive AI Devs

Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy at:

http://www.apache.org/licenses/LICENSE-2.0

NOTICE

This product includes original work by:
Damien Davison & Michael Maillet & Sacha Davison (Recursive AI Devs)
Additional details can be found in the project's LICENSE and source headers.

About

Spectral-State Guided Synthesis

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

83 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SSGS — Spectral-State Guided Synthesis

© 2025 Damien Davison & Michael Maillet & Sacha Davison

Recursive AI Devs

Licensed under the Apache License 2.0 psychcoherence@gmail.com OR therealmichaelmaillet@gmail.com

SSGS (Spectral-State Guided Synthesis) is a hybrid generative audio model that combines linear predictive coding (LPC), hidden Markov models (HMMs), spectral clustering, global graph search, and a physically-inspired Karplus–Strong excitation engine into a single synthesis framework.

Unlike classical systems that commit to a single paradigm (neural, statistical, or physical), SSGS deliberately smashes algorithms together and extracts the useful computational primitives from each.
The result is a model that:

  • learns spectral structure from real signals,
  • organizes that structure into Markov states,
  • decodes the optimal state trajectory over time via A* search,
  • generates excitation signals using string-model dynamics, and
  • reconstructs audio by filtering excitation through LPC envelopes.

SSGS is not meant to imitate any existing model family.
It is a new hybrid class: detailed like DSP, global like HMMs, and expressive like physical models.


Key Idea: “Algorithm Deconstruction”

SSGS is built using the same philosophy behind the Symbo project:

Break apart many algorithms, strip them to their mathematical essence, recombine only the primitives that are actually useful.

Instead of inheriting algorithms whole (e.g., a standard HMM, standard LPC vocoder, standard Karplus–Strong), SSGS takes:

  • LPC → spectral envelope representation
  • HMM → temporal clustering + statistical state transitions
  • EM → unsupervised parameter refinement
  • A* search → globally optimal state sequence
  • Karplus–Strong → natural resonance and excitation noise
  • Heuristics → spectral smoothness constraints
  • Graph theory → prune invalid or degenerate state structures

SSGS then recomposes these pieces into a single generative pipeline that didn’t exist before.

This approach is extremely flexible: swap the envelope model, swap the excitation, modify the search heuristic — the system keeps working.


Pipeline Overview

1. LPC Analysis

The training signal is segmented and analyzed with LPC to extract:

  • LPC coefficients
  • excitation residual
  • power envelope

These become the feature vectors for clustering.


2. Spectral Clustering via HMM Initialization

Frames are embedded into a feature space (typically derived from LPC spectra). States are initialized using k-means, then upgraded into a full HMM with:

  • initial state distribution
  • transition matrix
  • mean vectors and covariance matrices for each state

3. Expectation-Maximization (EM) Refinement

A custom EM implementation updates:

  • state responsibilities (γ)
  • pairwise transitions (ξ)
  • transition probabilities
  • Gaussian parameters (means, covariances)

Result: the HMM becomes a structured map of repeating spectral “modes.”


4. Graph Constraint Pruning

To prevent degenerate solutions, SSGS analyzes the transition graph:

  • identifies strongly connected components (SCCs)
  • removes invalid or isolated states
  • ensures state sequences remain musically plausible

5. Global State Decoding (A Search)*

Instead of using Viterbi (which is greedy and purely local), SSGS uses A*:

  • cost function = negative log-likelihood + spectral smoothness heuristic
  • ensures global consistency in the decoded state path
  • supports long-range structure better than classical decoding

6. Physically Inspired Excitation

The decoded state sequence modulates a Karplus–Strong string model, producing a dynamic excitation that is:

  • rich in overtones
  • noisy where appropriate
  • resonant and evolving

7. LPC Synthesis

Finally, excitation is passed through the LPC filters of each decoded state:

  • reconstructs spectral envelopes
  • restores formants and resonances
  • yields new but structurally coherent audio

This is how SSGS generates novel signals even from short or simple training data.


Model Export, Compression, and Checkpointing

After training you can persist the learned parameters in a compact archive:

ssgs.export_model("models/ssgs_baseline", use_compression=True, pack_covariances=True)
  • .npz output uses ZIP compression by default when use_compression=True.
  • Covariance matrices are serialized as lower-triangular slices to halve their stored size.
  • Set include_training_artifacts=True if you also need cached LPC and residual buffers for later analysis.
  • Reload the model with SpectralStateGuidedSynthesis.load_model(path).

For long-running training, SSGS also supports .safetensors checkpoints:

ssgs.save_checkpoint("models/checkpoint_epoch_5.safetensors")
ssgs=SpectralStateGuidedSynthesis.load_checkpoint("models/checkpoint_epoch_5.safetensors")

Checkpoints can optionally persist adaptive statistics so continuous learning can resume without reprocessing previous audio.


Generative Capability Levels

SSGS now offers enhanced generative capabilities through configurable state counts:

  • Standard (16 states): Original baseline configuration
  • Enhanced (34 states): 112% increase - recommended for most applications
  • Maximum (42 states): 163% increase - highest generative capacity
# Enhanced generative capabilities (default)ssgs=SpectralStateGuidedSynthesis(n_states=34)
# Maximum generative capabilitiesssgs=SpectralStateGuidedSynthesis(n_states=42)
# Original baselinessgs=SpectralStateGuidedSynthesis(n_states=16)

More states = more spectral patterns the model can learn and generate = richer, more diverse output.


Example Output

SSGS produces time-domain and spectral visualizations like this:

SSGS demo output

For a deeper comparison diagnostic, see ssgs_analysis.png.

The generated signal is not a copy —
it is a new trajectory through learned spectral states.


Quick Start

Install dependencies:

pip install -r requirements.txt

Run the demo training + generation script:

python test_ssgs.py

Run the enhanced capabilities demo (defaults to 34 states):

python demo_enhanced_capabilities.py

Train on folders of audio with checkpointing:

python train_on_folder.py --folders training_001 training_002

Adaptive Learning

SSGS can adapt to new audio using persistent statistics:

ssgs.adapt_to_audio(new_audio, sample_rate=16000, adaptation_rate=0.5, memory_blend=0.3)

Adaptation can be exported through .npz or .safetensors artifacts and reloaded later.

Feature Indexing & Retrieval

Use built-in nearest-neighbor search over frame features for exploration and analysis:

index=ssgs.build_feature_index(feature_space="lpc", normalize=True)
result=ssgs.query_similar_frames(frame_idx=42, k=8, feature_space="lpc")
vector_result=ssgs.search_by_feature_vector(
vector=index.tree.data[0],
k=5,
feature_space="lpc",
)

The returned FrameQueryResult includes indices, optional distances, and frame metadata.

Performance: Native Compiled Backend

The numeric hot loops are also implemented as a small, dependency-light C library (ssgs_native.c) and exposed through ssgs_native.py. ssgs.py automatically uses them when libssgs_native.so is loadable, and falls back to the existing SciPy/NumPy paths otherwise — so the package still works on a machine with no compiler installed.

Kernels covered: LPC (Levinson–Durbin), Gaussian emission (hand Cholesky + forward solve), EM forward–backward, and the smoothness-first Viterbi decode. The decode is the dominant synthesis cost (a pure-Python triple loop in the SciPy build); the compiled kernel is ~170× faster there, with identical output:

kernelSciPy/NumPyCspeedup
LPC (Levinson–Durbin)0.172 s0.172 s1.0×
Gaussian emission0.0014 s0.0012 s1.2×
decode (2000 frames)0.4998 s0.0029 s170×

The C kernels are numerically identical to ssgs.py (enforced by test_regression_guard.py). Build the library with:

make # produces libssgs_native.so (requires gcc)

A pure-NumPy reference implementation lives in ssgs_core.py and is used by the regression guard to prove equivalence even when the compiled library is absent.


Testing

Run individual checks as standalone scripts:

python test_ssgs.py
python test_checkpoints.py
python test_adaptive_persistence.py
python test_data_indexing.py
python test_gain_control.py
python test_generative_capabilities.py
python test_utils.py

Numeric-equivalence and performance gates:

python ssgs_core.py # pure-NumPy core self-test (equivalence oracle)
python ssgs_native.py # compiled-kernel verify + benchmark
python test_regression_guard.py # native load + core/pipeline equivalence

Requirements

numpy>=1.21.0
scipy>=1.7.0
matplotlib>=3.4.0
soundfile>=0.10.0
PyWavelets>=1.1.0
safetensors>=0.4.0

Optional: gcc and make to build the native backend (libssgs_native.so). If unavailable, SSGS falls back to the SciPy/NumPy paths automatically.

Project Structure

ssgs.py # Full model implementation
ssgs_core.py # Pure-NumPy reference kernels (equivalence oracle)
ssgs_native.c # Compiled C kernels (Levinson-Durbin, emission, EM, decode)
ssgs_native.py # ctypes wrapper + benchmark for the C kernels
train_on_folder.py # Folder-based training with checkpointing
demo_enhanced_capabilities.py # State-count showcase
example_checkpoints.py # Checkpoint usage example
example_fidelity.py # Fidelity-focused generation example
test_ssgs.py # Training + generation demo
test_checkpoints.py # Checkpoint save/load tests
test_adaptive_persistence.py # Adaptive learning tests
test_data_indexing.py # Feature index tests
test_gain_control.py # Gain control tests
test_generative_capabilities.py # State-count tests
test_regression_guard.py # Native load + core/pipeline equivalence guard
Makefile # Builds libssgs_native.so
requirements.txt # Dependencies
ssgs_demo.png # Example output
ssgs_analysis.png # Diagnostic output
README.md # This document \

Authors

Damien Davison & Michael Maillet & Sacha Davison Recursive AI Devs

We build hybrid-symbolic neural, statistical, and physical AI systems by algorithm decomposition — extracting the useful primitives and recombining them into new model classes.

If you use SSGS in research or production, please cite the authors.

License — Apache 2.0

This project is licensed under the Apache License, Version 2.0.

Copyright 2025 Damien Davison & Michael Maillet Recursive AI Devs

Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy at:

http://www.apache.org/licenses/LICENSE-2.0

NOTICE

This product includes original work by:
Damien Davison & Michael Maillet & Sacha Davison (Recursive AI Devs)
Additional details can be found in the project's LICENSE and source headers.

About

Spectral-State Guided Synthesis

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

83 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SSGS — Spectral-State Guided Synthesis

© 2025 Damien Davison & Michael Maillet & Sacha Davison

Recursive AI Devs

Licensed under the Apache License 2.0 psychcoherence@gmail.com OR therealmichaelmaillet@gmail.com

SSGS (Spectral-State Guided Synthesis) is a hybrid generative audio model that combines linear predictive coding (LPC), hidden Markov models (HMMs), spectral clustering, global graph search, and a physically-inspired Karplus–Strong excitation engine into a single synthesis framework.

Unlike classical systems that commit to a single paradigm (neural, statistical, or physical), SSGS deliberately smashes algorithms together and extracts the useful computational primitives from each.
The result is a model that:

  • learns spectral structure from real signals,
  • organizes that structure into Markov states,
  • decodes the optimal state trajectory over time via A* search,
  • generates excitation signals using string-model dynamics, and
  • reconstructs audio by filtering excitation through LPC envelopes.

SSGS is not meant to imitate any existing model family.
It is a new hybrid class: detailed like DSP, global like HMMs, and expressive like physical models.


Key Idea: “Algorithm Deconstruction”

SSGS is built using the same philosophy behind the Symbo project:

Break apart many algorithms, strip them to their mathematical essence, recombine only the primitives that are actually useful.

Instead of inheriting algorithms whole (e.g., a standard HMM, standard LPC vocoder, standard Karplus–Strong), SSGS takes:

  • LPC → spectral envelope representation
  • HMM → temporal clustering + statistical state transitions
  • EM → unsupervised parameter refinement
  • A* search → globally optimal state sequence
  • Karplus–Strong → natural resonance and excitation noise
  • Heuristics → spectral smoothness constraints
  • Graph theory → prune invalid or degenerate state structures

SSGS then recomposes these pieces into a single generative pipeline that didn’t exist before.

This approach is extremely flexible: swap the envelope model, swap the excitation, modify the search heuristic — the system keeps working.


Pipeline Overview

1. LPC Analysis

The training signal is segmented and analyzed with LPC to extract:

  • LPC coefficients
  • excitation residual
  • power envelope

These become the feature vectors for clustering.


2. Spectral Clustering via HMM Initialization

Frames are embedded into a feature space (typically derived from LPC spectra). States are initialized using k-means, then upgraded into a full HMM with:

  • initial state distribution
  • transition matrix
  • mean vectors and covariance matrices for each state

3. Expectation-Maximization (EM) Refinement

A custom EM implementation updates:

  • state responsibilities (γ)
  • pairwise transitions (ξ)
  • transition probabilities
  • Gaussian parameters (means, covariances)

Result: the HMM becomes a structured map of repeating spectral “modes.”


4. Graph Constraint Pruning

To prevent degenerate solutions, SSGS analyzes the transition graph:

  • identifies strongly connected components (SCCs)
  • removes invalid or isolated states
  • ensures state sequences remain musically plausible

5. Global State Decoding (A Search)*

Instead of using Viterbi (which is greedy and purely local), SSGS uses A*:

  • cost function = negative log-likelihood + spectral smoothness heuristic
  • ensures global consistency in the decoded state path
  • supports long-range structure better than classical decoding

6. Physically Inspired Excitation

The decoded state sequence modulates a Karplus–Strong string model, producing a dynamic excitation that is:

  • rich in overtones
  • noisy where appropriate
  • resonant and evolving

7. LPC Synthesis

Finally, excitation is passed through the LPC filters of each decoded state:

  • reconstructs spectral envelopes
  • restores formants and resonances
  • yields new but structurally coherent audio

This is how SSGS generates novel signals even from short or simple training data.


Model Export, Compression, and Checkpointing

After training you can persist the learned parameters in a compact archive:

ssgs.export_model("models/ssgs_baseline", use_compression=True, pack_covariances=True)
  • .npz output uses ZIP compression by default when use_compression=True.
  • Covariance matrices are serialized as lower-triangular slices to halve their stored size.
  • Set include_training_artifacts=True if you also need cached LPC and residual buffers for later analysis.
  • Reload the model with SpectralStateGuidedSynthesis.load_model(path).

For long-running training, SSGS also supports .safetensors checkpoints:

ssgs.save_checkpoint("models/checkpoint_epoch_5.safetensors")
ssgs=SpectralStateGuidedSynthesis.load_checkpoint("models/checkpoint_epoch_5.safetensors")

Checkpoints can optionally persist adaptive statistics so continuous learning can resume without reprocessing previous audio.


Generative Capability Levels

SSGS now offers enhanced generative capabilities through configurable state counts:

  • Standard (16 states): Original baseline configuration
  • Enhanced (34 states): 112% increase - recommended for most applications
  • Maximum (42 states): 163% increase - highest generative capacity
# Enhanced generative capabilities (default)ssgs=SpectralStateGuidedSynthesis(n_states=34)
# Maximum generative capabilitiesssgs=SpectralStateGuidedSynthesis(n_states=42)
# Original baselinessgs=SpectralStateGuidedSynthesis(n_states=16)

More states = more spectral patterns the model can learn and generate = richer, more diverse output.


Example Output

SSGS produces time-domain and spectral visualizations like this:

SSGS demo output

For a deeper comparison diagnostic, see ssgs_analysis.png.

The generated signal is not a copy —
it is a new trajectory through learned spectral states.


Quick Start

Install dependencies:

pip install -r requirements.txt

Run the demo training + generation script:

python test_ssgs.py

Run the enhanced capabilities demo (defaults to 34 states):

python demo_enhanced_capabilities.py

Train on folders of audio with checkpointing:

python train_on_folder.py --folders training_001 training_002

Adaptive Learning

SSGS can adapt to new audio using persistent statistics:

ssgs.adapt_to_audio(new_audio, sample_rate=16000, adaptation_rate=0.5, memory_blend=0.3)

Adaptation can be exported through .npz or .safetensors artifacts and reloaded later.

Feature Indexing & Retrieval

Use built-in nearest-neighbor search over frame features for exploration and analysis:

index=ssgs.build_feature_index(feature_space="lpc", normalize=True)
result=ssgs.query_similar_frames(frame_idx=42, k=8, feature_space="lpc")
vector_result=ssgs.search_by_feature_vector(
vector=index.tree.data[0],
k=5,
feature_space="lpc",
)

The returned FrameQueryResult includes indices, optional distances, and frame metadata.

Performance: Native Compiled Backend

The numeric hot loops are also implemented as a small, dependency-light C library (ssgs_native.c) and exposed through ssgs_native.py. ssgs.py automatically uses them when libssgs_native.so is loadable, and falls back to the existing SciPy/NumPy paths otherwise — so the package still works on a machine with no compiler installed.

Kernels covered: LPC (Levinson–Durbin), Gaussian emission (hand Cholesky + forward solve), EM forward–backward, and the smoothness-first Viterbi decode. The decode is the dominant synthesis cost (a pure-Python triple loop in the SciPy build); the compiled kernel is ~170× faster there, with identical output:

kernelSciPy/NumPyCspeedup
LPC (Levinson–Durbin)0.172 s0.172 s1.0×
Gaussian emission0.0014 s0.0012 s1.2×
decode (2000 frames)0.4998 s0.0029 s170×

The C kernels are numerically identical to ssgs.py (enforced by test_regression_guard.py). Build the library with:

make # produces libssgs_native.so (requires gcc)

A pure-NumPy reference implementation lives in ssgs_core.py and is used by the regression guard to prove equivalence even when the compiled library is absent.


Testing

Run individual checks as standalone scripts:

python test_ssgs.py
python test_checkpoints.py
python test_adaptive_persistence.py
python test_data_indexing.py
python test_gain_control.py
python test_generative_capabilities.py
python test_utils.py

Numeric-equivalence and performance gates:

python ssgs_core.py # pure-NumPy core self-test (equivalence oracle)
python ssgs_native.py # compiled-kernel verify + benchmark
python test_regression_guard.py # native load + core/pipeline equivalence

Requirements

numpy>=1.21.0
scipy>=1.7.0
matplotlib>=3.4.0
soundfile>=0.10.0
PyWavelets>=1.1.0
safetensors>=0.4.0

Optional: gcc and make to build the native backend (libssgs_native.so). If unavailable, SSGS falls back to the SciPy/NumPy paths automatically.

Project Structure

ssgs.py # Full model implementation
ssgs_core.py # Pure-NumPy reference kernels (equivalence oracle)
ssgs_native.c # Compiled C kernels (Levinson-Durbin, emission, EM, decode)
ssgs_native.py # ctypes wrapper + benchmark for the C kernels
train_on_folder.py # Folder-based training with checkpointing
demo_enhanced_capabilities.py # State-count showcase
example_checkpoints.py # Checkpoint usage example
example_fidelity.py # Fidelity-focused generation example
test_ssgs.py # Training + generation demo
test_checkpoints.py # Checkpoint save/load tests
test_adaptive_persistence.py # Adaptive learning tests
test_data_indexing.py # Feature index tests
test_gain_control.py # Gain control tests
test_generative_capabilities.py # State-count tests
test_regression_guard.py # Native load + core/pipeline equivalence guard
Makefile # Builds libssgs_native.so
requirements.txt # Dependencies
ssgs_demo.png # Example output
ssgs_analysis.png # Diagnostic output
README.md # This document \

Authors

Damien Davison & Michael Maillet & Sacha Davison Recursive AI Devs

We build hybrid-symbolic neural, statistical, and physical AI systems by algorithm decomposition — extracting the useful primitives and recombining them into new model classes.

If you use SSGS in research or production, please cite the authors.

License — Apache 2.0

This project is licensed under the Apache License, Version 2.0.

Copyright 2025 Damien Davison & Michael Maillet Recursive AI Devs

Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy at:

http://www.apache.org/licenses/LICENSE-2.0

NOTICE

This product includes original work by:
Damien Davison & Michael Maillet & Sacha Davison (Recursive AI Devs)
Additional details can be found in the project's LICENSE and source headers.

About

Spectral-State Guided Synthesis

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

83 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SSGS — Spectral-State Guided Synthesis

© 2025 Damien Davison & Michael Maillet & Sacha Davison

Recursive AI Devs

Licensed under the Apache License 2.0 psychcoherence@gmail.com OR therealmichaelmaillet@gmail.com

SSGS (Spectral-State Guided Synthesis) is a hybrid generative audio model that combines linear predictive coding (LPC), hidden Markov models (HMMs), spectral clustering, global graph search, and a physically-inspired Karplus–Strong excitation engine into a single synthesis framework.

Unlike classical systems that commit to a single paradigm (neural, statistical, or physical), SSGS deliberately smashes algorithms together and extracts the useful computational primitives from each.
The result is a model that:

  • learns spectral structure from real signals,
  • organizes that structure into Markov states,
  • decodes the optimal state trajectory over time via A* search,
  • generates excitation signals using string-model dynamics, and
  • reconstructs audio by filtering excitation through LPC envelopes.

SSGS is not meant to imitate any existing model family.
It is a new hybrid class: detailed like DSP, global like HMMs, and expressive like physical models.


Key Idea: “Algorithm Deconstruction”

SSGS is built using the same philosophy behind the Symbo project:

Break apart many algorithms, strip them to their mathematical essence, recombine only the primitives that are actually useful.

Instead of inheriting algorithms whole (e.g., a standard HMM, standard LPC vocoder, standard Karplus–Strong), SSGS takes:

  • LPC → spectral envelope representation
  • HMM → temporal clustering + statistical state transitions
  • EM → unsupervised parameter refinement
  • A* search → globally optimal state sequence
  • Karplus–Strong → natural resonance and excitation noise
  • Heuristics → spectral smoothness constraints
  • Graph theory → prune invalid or degenerate state structures

SSGS then recomposes these pieces into a single generative pipeline that didn’t exist before.

This approach is extremely flexible: swap the envelope model, swap the excitation, modify the search heuristic — the system keeps working.


Pipeline Overview

1. LPC Analysis

The training signal is segmented and analyzed with LPC to extract:

  • LPC coefficients
  • excitation residual
  • power envelope

These become the feature vectors for clustering.


2. Spectral Clustering via HMM Initialization

Frames are embedded into a feature space (typically derived from LPC spectra). States are initialized using k-means, then upgraded into a full HMM with:

  • initial state distribution
  • transition matrix
  • mean vectors and covariance matrices for each state

3. Expectation-Maximization (EM) Refinement

A custom EM implementation updates:

  • state responsibilities (γ)
  • pairwise transitions (ξ)
  • transition probabilities
  • Gaussian parameters (means, covariances)

Result: the HMM becomes a structured map of repeating spectral “modes.”


4. Graph Constraint Pruning

To prevent degenerate solutions, SSGS analyzes the transition graph:

  • identifies strongly connected components (SCCs)
  • removes invalid or isolated states
  • ensures state sequences remain musically plausible

5. Global State Decoding (A Search)*

Instead of using Viterbi (which is greedy and purely local), SSGS uses A*:

  • cost function = negative log-likelihood + spectral smoothness heuristic
  • ensures global consistency in the decoded state path
  • supports long-range structure better than classical decoding

6. Physically Inspired Excitation

The decoded state sequence modulates a Karplus–Strong string model, producing a dynamic excitation that is:

  • rich in overtones
  • noisy where appropriate
  • resonant and evolving

7. LPC Synthesis

Finally, excitation is passed through the LPC filters of each decoded state:

  • reconstructs spectral envelopes
  • restores formants and resonances
  • yields new but structurally coherent audio

This is how SSGS generates novel signals even from short or simple training data.


Model Export, Compression, and Checkpointing

After training you can persist the learned parameters in a compact archive:

ssgs.export_model("models/ssgs_baseline", use_compression=True, pack_covariances=True)
  • .npz output uses ZIP compression by default when use_compression=True.
  • Covariance matrices are serialized as lower-triangular slices to halve their stored size.
  • Set include_training_artifacts=True if you also need cached LPC and residual buffers for later analysis.
  • Reload the model with SpectralStateGuidedSynthesis.load_model(path).

For long-running training, SSGS also supports .safetensors checkpoints:

ssgs.save_checkpoint("models/checkpoint_epoch_5.safetensors")
ssgs=SpectralStateGuidedSynthesis.load_checkpoint("models/checkpoint_epoch_5.safetensors")

Checkpoints can optionally persist adaptive statistics so continuous learning can resume without reprocessing previous audio.


Generative Capability Levels

SSGS now offers enhanced generative capabilities through configurable state counts:

  • Standard (16 states): Original baseline configuration
  • Enhanced (34 states): 112% increase - recommended for most applications
  • Maximum (42 states): 163% increase - highest generative capacity
# Enhanced generative capabilities (default)ssgs=SpectralStateGuidedSynthesis(n_states=34)
# Maximum generative capabilitiesssgs=SpectralStateGuidedSynthesis(n_states=42)
# Original baselinessgs=SpectralStateGuidedSynthesis(n_states=16)

More states = more spectral patterns the model can learn and generate = richer, more diverse output.


Example Output

SSGS produces time-domain and spectral visualizations like this:

SSGS demo output

For a deeper comparison diagnostic, see ssgs_analysis.png.

The generated signal is not a copy —
it is a new trajectory through learned spectral states.


Quick Start

Install dependencies:

pip install -r requirements.txt

Run the demo training + generation script:

python test_ssgs.py

Run the enhanced capabilities demo (defaults to 34 states):

python demo_enhanced_capabilities.py

Train on folders of audio with checkpointing:

python train_on_folder.py --folders training_001 training_002

Adaptive Learning

SSGS can adapt to new audio using persistent statistics:

ssgs.adapt_to_audio(new_audio, sample_rate=16000, adaptation_rate=0.5, memory_blend=0.3)

Adaptation can be exported through .npz or .safetensors artifacts and reloaded later.

Feature Indexing & Retrieval

Use built-in nearest-neighbor search over frame features for exploration and analysis:

index=ssgs.build_feature_index(feature_space="lpc", normalize=True)
result=ssgs.query_similar_frames(frame_idx=42, k=8, feature_space="lpc")
vector_result=ssgs.search_by_feature_vector(
vector=index.tree.data[0],
k=5,
feature_space="lpc",
)

The returned FrameQueryResult includes indices, optional distances, and frame metadata.

Performance: Native Compiled Backend

The numeric hot loops are also implemented as a small, dependency-light C library (ssgs_native.c) and exposed through ssgs_native.py. ssgs.py automatically uses them when libssgs_native.so is loadable, and falls back to the existing SciPy/NumPy paths otherwise — so the package still works on a machine with no compiler installed.

Kernels covered: LPC (Levinson–Durbin), Gaussian emission (hand Cholesky + forward solve), EM forward–backward, and the smoothness-first Viterbi decode. The decode is the dominant synthesis cost (a pure-Python triple loop in the SciPy build); the compiled kernel is ~170× faster there, with identical output:

kernelSciPy/NumPyCspeedup
LPC (Levinson–Durbin)0.172 s0.172 s1.0×
Gaussian emission0.0014 s0.0012 s1.2×
decode (2000 frames)0.4998 s0.0029 s170×

The C kernels are numerically identical to ssgs.py (enforced by test_regression_guard.py). Build the library with:

make # produces libssgs_native.so (requires gcc)

A pure-NumPy reference implementation lives in ssgs_core.py and is used by the regression guard to prove equivalence even when the compiled library is absent.


Testing

Run individual checks as standalone scripts:

python test_ssgs.py
python test_checkpoints.py
python test_adaptive_persistence.py
python test_data_indexing.py
python test_gain_control.py
python test_generative_capabilities.py
python test_utils.py

Numeric-equivalence and performance gates:

python ssgs_core.py # pure-NumPy core self-test (equivalence oracle)
python ssgs_native.py # compiled-kernel verify + benchmark
python test_regression_guard.py # native load + core/pipeline equivalence

Requirements

numpy>=1.21.0
scipy>=1.7.0
matplotlib>=3.4.0
soundfile>=0.10.0
PyWavelets>=1.1.0
safetensors>=0.4.0

Optional: gcc and make to build the native backend (libssgs_native.so). If unavailable, SSGS falls back to the SciPy/NumPy paths automatically.

Project Structure

ssgs.py # Full model implementation
ssgs_core.py # Pure-NumPy reference kernels (equivalence oracle)
ssgs_native.c # Compiled C kernels (Levinson-Durbin, emission, EM, decode)
ssgs_native.py # ctypes wrapper + benchmark for the C kernels
train_on_folder.py # Folder-based training with checkpointing
demo_enhanced_capabilities.py # State-count showcase
example_checkpoints.py # Checkpoint usage example
example_fidelity.py # Fidelity-focused generation example
test_ssgs.py # Training + generation demo
test_checkpoints.py # Checkpoint save/load tests
test_adaptive_persistence.py # Adaptive learning tests
test_data_indexing.py # Feature index tests
test_gain_control.py # Gain control tests
test_generative_capabilities.py # State-count tests
test_regression_guard.py # Native load + core/pipeline equivalence guard
Makefile # Builds libssgs_native.so
requirements.txt # Dependencies
ssgs_demo.png # Example output
ssgs_analysis.png # Diagnostic output
README.md # This document \

Authors

Damien Davison & Michael Maillet & Sacha Davison Recursive AI Devs

We build hybrid-symbolic neural, statistical, and physical AI systems by algorithm decomposition — extracting the useful primitives and recombining them into new model classes.

If you use SSGS in research or production, please cite the authors.

License — Apache 2.0

This project is licensed under the Apache License, Version 2.0.

Copyright 2025 Damien Davison & Michael Maillet Recursive AI Devs

Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy at:

http://www.apache.org/licenses/LICENSE-2.0

NOTICE

This product includes original work by:
Damien Davison & Michael Maillet & Sacha Davison (Recursive AI Devs)
Additional details can be found in the project's LICENSE and source headers.

About

Spectral-State Guided Synthesis

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages