Repository files navigation

python-playground

Playground for developing / demonstrating python dev techniques.

This repository hosts Acoustic Reference Book — Phase 1: a schema-driven pipeline that turns acoustic calculation output into validated XML Acoustic Dataset, built as a learning project where every decision is documented and defensible.

Quick start

Open in a GitHub Codespace (Code → Codespaces → Create codespace). It provisions Python 3.9.4 (matching the target system) and runs make bootstrap automatically. Then:

make verify # lint + tests — confirm the environment is green
make docs-serve # browse the docs at http://localhost:8000

Working locally instead? You need Python 3.9.x and make, then make bootstrap. Run make to see all targets.

Self-guided adventures

New here? Work through these in order. Each is a short, hands-on loop — run a command, look at what it produced, then peek at the code that did it. Everything below has been run end-to-end, so the output you see should match.

All commands assume make bootstrap has finished (it runs automatically in a Codespace). If acoustic isn't found, the package is exposed as a module too: replace acoustic with python -m acoustic_dataset.cli.

Adventure 1 — Run the whole pipeline and watch the gates

make verify # lint + type-check + tests: is the env green?
acoustic pipeline # input XML -> typed objects -> XML -> validate

You should see something like:

pipeline ok: 10 band(s) -> build/acoustic_dataset.xml (schema-valid, round-trip-equal)

Open the file it wrote (build/acoustic_dataset.xml) and compare it to the input (examples/calculation_input.xml). Then run the structural gate on its own and the migration-safety diff against the known-good reference:

acoustic validate --xml build/acoustic_dataset.xml
acoustic compare build/acoustic_dataset.xml examples/reference/trial_known_good.xml

Now make it your own.examples/calculation_input.xml is your sandbox — edit a value and re-run to watch the validated output change. Bump <BandCount>10 to 20, or <BaseLevelDb>140.0 to 155.0, then:

acoustic pipeline

The band count in the message and the generated XML move with your edit. This is the tightest loop in the project: input XML in, validated dataset out.

Then break it on purpose to watch the gates earn their keep. Set <BandCount> to a negative number, or put text where a number belongs, and re-run:

error: build rejected a value before serialisation: ... schema: value must be positive

The pipeline rejects it, prints the reason, exits non-zero, and does not write a stale artifact. (Restore the file afterward — git checkout examples/calculation_input.xml.)

Now read why:docs/concepts/two-verification-gates.md explains why "schema-valid" and "correct" are two different checks. The code lives in src/acoustic_dataset/validate.py and compare.py, dispatched from cli.py.

Adventure 2 — Explore a typed Python data class

The pipeline never carries data in a loose dict; it builds one typed object that meets the schema. Feel the difference yourself.

First, experience editor autocomplete. Open examples/explore_platform.py in VS Code (the Codespace ships Pylance). Find the # platform.radiated_noise. line, delete the #, put your cursor right after the dot, and type: VS Code pops up the object's declared attributes (band, …) — statically, before the code has even run. Try platform. too for the top-level attributes (radiated_noise, sensors, …). Keep drilling (platform.radiated_noise.band[0].) and it completes all the way down. That instant, schema-shaped autocomplete is the real payoff of declared fields. Run the file any time with:

python examples/explore_platform.py

If autocomplete doesn't appear, give Pylance a moment to index, and make sure VS Code's Python interpreter is set to the Codespace's 3.9 (bottom-right status bar / Python: Select Interpreter).

Then poke at it in the REPL:

python
>>>fromacoustic_datasetimportbuild, serialize>>>platform=build.build_platform_from_file("examples/calculation_input.xml")
>>>type(platform).__name__'Platform'>>>platform.radiated_noise.band[0].centre_frequency# a Decimal, not a string>>>print(serialize.to_xml(platform)[:500]) # the same object, as XML>>>fromacoustic_dataset.models.acoustic_datasetimportSector>>>Sector(bering=1, level=2) # a typo is a TypeError, not a silent new key

The REPL can also complete attributes, but that's a separate readline/rlcompleter feature that introspects the live object at runtime — not the static, declared-field completion you saw in the editor. If <TAB> doesn't complete, enable it for the session with import readline, rlcompleter; readline.parse_and_bind("tab: complete").

Read docs/concepts/typed-vs-dicts.md for what declared fields buy you over a dictionary, then look at how the builder rejects out-of-range values in src/acoustic_dataset/build.py.

Adventure 3 — Reverse-engineer the data classes from the schema

Those data classes aren't hand-written — they're generated from the XSD by xsdata. The XSD is the single source of truth; the Python is a build artifact (note the # DO NOT EDIT BY HAND header on every model file).

acoustic generate # regenerate models from every schema/*.xsd

Trace the chain for one element:

  1. Open schema/acoustic_dataset.xsd and find an element, e.g. Sector.
  2. Open the generated src/acoustic_dataset/models/acoustic_dataset.py and find the matching @dataclass. Notice how XSD types, ranges, and docs became Python field metadata.
  3. Read src/acoustic_dataset/generate.py to see the exact xsdata invocation — and the two post-processing steps that keep the output deterministic and Python-3.9-compatible.

Want to see generation in action? Change a <xs:documentation> string in the XSD, run acoustic generate, and git diff the models — the docstring tracks the schema. (Revert the XSD edit afterward; CI fails if committed models drift from the schema — ADR 0008.)

Adventure 4 — Generate and view the XSD documentation

Two complementary docs come out of this repo. First, a standalone HTML schema reference rendered straight from the XSD (via the vendored xs3p stylesheet):

acoustic gen-schema-docs --out build/schema-preview.html

Open build/schema-preview.html (in a Codespace: right-click the file → Download, or use the Live Preview extension) to browse every element, type, and constraint of the contract. The pre-built version is committed at docs/reference/schema/index.html. The generator lives in src/acoustic_dataset/schema_html.py.

Second, the full project site — tutorials, concepts, ADRs, and a Mermaid ERD — served by MkDocs Material:

make docs-serve # browse at http://localhost:8000

In a Codespace, when the port-forward notification appears, click Open in Browser.

Adventure 5 — Point it at the real (private) schema

Everything above runs against the placeholder schema committed in schema/. The real, proprietary XSD and corpus are meant to live under private/ — a directory that is entirely gitignored (see .gitignore) so it never reaches git, CI, or the internet.

One constraint shapes the workflow: src/acoustic_dataset/models/ is committed (the placeholder-generated models that CI drift-checks), so you must not regenerate over it from a real schema — an accidental commit would leak that structure. Generate real models into private/ instead. The same CLI takes explicit paths, so once the real material is in place (real XSD in private/schema/, inputs in private/examples/, known-good XML in private/reference/):

# Generate typed models from the real XSD into the gitignored output dir
acoustic generate --schema private/schema/<real>.xsd --out private/models
# Structural gate (XSD + round-trip) on a real file against the real schema
acoustic validate --xml private/examples/<real>.xml --schema private/schema/<real>.xsd
# Migration-safety diff: generated vs known-good reference
acoustic compare private/<generated>.xml private/reference/<known_good>.xml

Caveat — the full pipeline command.acoustic pipeline imports acoustic_dataset.models (the committed package), so it always builds against the placeholder-derived bindings even when you pass --schema private/.... Running the end-to-end pipeline on real-generated models would mean repointing that import at private/models/ — a deliberate change that touches committed code, so raise it for review first. With the real schema you can use generate, validate, and compare against private/ paths immediately; full pipeline needs that reviewed change.

Where to go next

Follow the guided tutorial that ties all of this together: docs/tutorials/01-start-here.md, then skim the decision records in docs/decisions/ — each says why a choice was made and what was rejected.

Documentation

The full, navigable documentation — tutorials, how-to guides, concepts, decision records (ADRs), and a Mermaid schema ERD — lives in docs/ and renders as an attractive HTML site via MkDocs Material (make docs).

Start here:docs/tutorials/01-start-here.md.

Spec-driven development

This project uses GitHub Spec Kit. The active feature's specification, plan, and design artifacts are under specs/001-codespace-xml-scaffold/. Use the /speckit-* skills (specify, plan, tasks, implement) to drive development.

Layout

PathWhat it is
.devcontainer/Codespaces environment (Python 3.9.4)
schema/The XSD contract
src/acoustic_dataset/The pipeline package (models/ is generated, never hand-edited)
examples/Example calculation input + reference/ known-good files
tests/Unit, integration, and golden-file tests
docs/The documentation site (MkDocs Material)
specs/Spec Kit feature specs, plans, and contracts

About

Playground to investigate/develop python tooling in support of complex XML Schema

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

python-playground

Playground for developing / demonstrating python dev techniques.

This repository hosts Acoustic Reference Book — Phase 1: a schema-driven pipeline that turns acoustic calculation output into validated XML Acoustic Dataset, built as a learning project where every decision is documented and defensible.

Quick start

Open in a GitHub Codespace (Code → Codespaces → Create codespace). It provisions Python 3.9.4 (matching the target system) and runs make bootstrap automatically. Then:

make verify # lint + tests — confirm the environment is green
make docs-serve # browse the docs at http://localhost:8000

Working locally instead? You need Python 3.9.x and make, then make bootstrap. Run make to see all targets.

Self-guided adventures

New here? Work through these in order. Each is a short, hands-on loop — run a command, look at what it produced, then peek at the code that did it. Everything below has been run end-to-end, so the output you see should match.

All commands assume make bootstrap has finished (it runs automatically in a Codespace). If acoustic isn't found, the package is exposed as a module too: replace acoustic with python -m acoustic_dataset.cli.

Adventure 1 — Run the whole pipeline and watch the gates

make verify # lint + type-check + tests: is the env green?
acoustic pipeline # input XML -> typed objects -> XML -> validate

You should see something like:

pipeline ok: 10 band(s) -> build/acoustic_dataset.xml (schema-valid, round-trip-equal)

Open the file it wrote (build/acoustic_dataset.xml) and compare it to the input (examples/calculation_input.xml). Then run the structural gate on its own and the migration-safety diff against the known-good reference:

acoustic validate --xml build/acoustic_dataset.xml
acoustic compare build/acoustic_dataset.xml examples/reference/trial_known_good.xml

Now make it your own.examples/calculation_input.xml is your sandbox — edit a value and re-run to watch the validated output change. Bump <BandCount>10 to 20, or <BaseLevelDb>140.0 to 155.0, then:

acoustic pipeline

The band count in the message and the generated XML move with your edit. This is the tightest loop in the project: input XML in, validated dataset out.

Then break it on purpose to watch the gates earn their keep. Set <BandCount> to a negative number, or put text where a number belongs, and re-run:

error: build rejected a value before serialisation: ... schema: value must be positive

The pipeline rejects it, prints the reason, exits non-zero, and does not write a stale artifact. (Restore the file afterward — git checkout examples/calculation_input.xml.)

Now read why:docs/concepts/two-verification-gates.md explains why "schema-valid" and "correct" are two different checks. The code lives in src/acoustic_dataset/validate.py and compare.py, dispatched from cli.py.

Adventure 2 — Explore a typed Python data class

The pipeline never carries data in a loose dict; it builds one typed object that meets the schema. Feel the difference yourself.

First, experience editor autocomplete. Open examples/explore_platform.py in VS Code (the Codespace ships Pylance). Find the # platform.radiated_noise. line, delete the #, put your cursor right after the dot, and type: VS Code pops up the object's declared attributes (band, …) — statically, before the code has even run. Try platform. too for the top-level attributes (radiated_noise, sensors, …). Keep drilling (platform.radiated_noise.band[0].) and it completes all the way down. That instant, schema-shaped autocomplete is the real payoff of declared fields. Run the file any time with:

python examples/explore_platform.py

If autocomplete doesn't appear, give Pylance a moment to index, and make sure VS Code's Python interpreter is set to the Codespace's 3.9 (bottom-right status bar / Python: Select Interpreter).

Then poke at it in the REPL:

python
>>>fromacoustic_datasetimportbuild, serialize>>>platform=build.build_platform_from_file("examples/calculation_input.xml")
>>>type(platform).__name__'Platform'>>>platform.radiated_noise.band[0].centre_frequency# a Decimal, not a string>>>print(serialize.to_xml(platform)[:500]) # the same object, as XML>>>fromacoustic_dataset.models.acoustic_datasetimportSector>>>Sector(bering=1, level=2) # a typo is a TypeError, not a silent new key

The REPL can also complete attributes, but that's a separate readline/rlcompleter feature that introspects the live object at runtime — not the static, declared-field completion you saw in the editor. If <TAB> doesn't complete, enable it for the session with import readline, rlcompleter; readline.parse_and_bind("tab: complete").

Read docs/concepts/typed-vs-dicts.md for what declared fields buy you over a dictionary, then look at how the builder rejects out-of-range values in src/acoustic_dataset/build.py.

Adventure 3 — Reverse-engineer the data classes from the schema

Those data classes aren't hand-written — they're generated from the XSD by xsdata. The XSD is the single source of truth; the Python is a build artifact (note the # DO NOT EDIT BY HAND header on every model file).

acoustic generate # regenerate models from every schema/*.xsd

Trace the chain for one element:

  1. Open schema/acoustic_dataset.xsd and find an element, e.g. Sector.
  2. Open the generated src/acoustic_dataset/models/acoustic_dataset.py and find the matching @dataclass. Notice how XSD types, ranges, and docs became Python field metadata.
  3. Read src/acoustic_dataset/generate.py to see the exact xsdata invocation — and the two post-processing steps that keep the output deterministic and Python-3.9-compatible.

Want to see generation in action? Change a <xs:documentation> string in the XSD, run acoustic generate, and git diff the models — the docstring tracks the schema. (Revert the XSD edit afterward; CI fails if committed models drift from the schema — ADR 0008.)

Adventure 4 — Generate and view the XSD documentation

Two complementary docs come out of this repo. First, a standalone HTML schema reference rendered straight from the XSD (via the vendored xs3p stylesheet):

acoustic gen-schema-docs --out build/schema-preview.html

Open build/schema-preview.html (in a Codespace: right-click the file → Download, or use the Live Preview extension) to browse every element, type, and constraint of the contract. The pre-built version is committed at docs/reference/schema/index.html. The generator lives in src/acoustic_dataset/schema_html.py.

Second, the full project site — tutorials, concepts, ADRs, and a Mermaid ERD — served by MkDocs Material:

make docs-serve # browse at http://localhost:8000

In a Codespace, when the port-forward notification appears, click Open in Browser.

Adventure 5 — Point it at the real (private) schema

Everything above runs against the placeholder schema committed in schema/. The real, proprietary XSD and corpus are meant to live under private/ — a directory that is entirely gitignored (see .gitignore) so it never reaches git, CI, or the internet.

One constraint shapes the workflow: src/acoustic_dataset/models/ is committed (the placeholder-generated models that CI drift-checks), so you must not regenerate over it from a real schema — an accidental commit would leak that structure. Generate real models into private/ instead. The same CLI takes explicit paths, so once the real material is in place (real XSD in private/schema/, inputs in private/examples/, known-good XML in private/reference/):

# Generate typed models from the real XSD into the gitignored output dir
acoustic generate --schema private/schema/<real>.xsd --out private/models
# Structural gate (XSD + round-trip) on a real file against the real schema
acoustic validate --xml private/examples/<real>.xml --schema private/schema/<real>.xsd
# Migration-safety diff: generated vs known-good reference
acoustic compare private/<generated>.xml private/reference/<known_good>.xml

Caveat — the full pipeline command.acoustic pipeline imports acoustic_dataset.models (the committed package), so it always builds against the placeholder-derived bindings even when you pass --schema private/.... Running the end-to-end pipeline on real-generated models would mean repointing that import at private/models/ — a deliberate change that touches committed code, so raise it for review first. With the real schema you can use generate, validate, and compare against private/ paths immediately; full pipeline needs that reviewed change.

Where to go next

Follow the guided tutorial that ties all of this together: docs/tutorials/01-start-here.md, then skim the decision records in docs/decisions/ — each says why a choice was made and what was rejected.

Documentation

The full, navigable documentation — tutorials, how-to guides, concepts, decision records (ADRs), and a Mermaid schema ERD — lives in docs/ and renders as an attractive HTML site via MkDocs Material (make docs).

Start here:docs/tutorials/01-start-here.md.

Spec-driven development

This project uses GitHub Spec Kit. The active feature's specification, plan, and design artifacts are under specs/001-codespace-xml-scaffold/. Use the /speckit-* skills (specify, plan, tasks, implement) to drive development.

Layout

PathWhat it is
.devcontainer/Codespaces environment (Python 3.9.4)
schema/The XSD contract
src/acoustic_dataset/The pipeline package (models/ is generated, never hand-edited)
examples/Example calculation input + reference/ known-good files
tests/Unit, integration, and golden-file tests
docs/The documentation site (MkDocs Material)
specs/Spec Kit feature specs, plans, and contracts

About

Playground to investigate/develop python tooling in support of complex XML Schema

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

python-playground

Playground for developing / demonstrating python dev techniques.

This repository hosts Acoustic Reference Book — Phase 1: a schema-driven pipeline that turns acoustic calculation output into validated XML Acoustic Dataset, built as a learning project where every decision is documented and defensible.

Quick start

Open in a GitHub Codespace (Code → Codespaces → Create codespace). It provisions Python 3.9.4 (matching the target system) and runs make bootstrap automatically. Then:

make verify # lint + tests — confirm the environment is green
make docs-serve # browse the docs at http://localhost:8000

Working locally instead? You need Python 3.9.x and make, then make bootstrap. Run make to see all targets.

Self-guided adventures

New here? Work through these in order. Each is a short, hands-on loop — run a command, look at what it produced, then peek at the code that did it. Everything below has been run end-to-end, so the output you see should match.

All commands assume make bootstrap has finished (it runs automatically in a Codespace). If acoustic isn't found, the package is exposed as a module too: replace acoustic with python -m acoustic_dataset.cli.

Adventure 1 — Run the whole pipeline and watch the gates

make verify # lint + type-check + tests: is the env green?
acoustic pipeline # input XML -> typed objects -> XML -> validate

You should see something like:

pipeline ok: 10 band(s) -> build/acoustic_dataset.xml (schema-valid, round-trip-equal)

Open the file it wrote (build/acoustic_dataset.xml) and compare it to the input (examples/calculation_input.xml). Then run the structural gate on its own and the migration-safety diff against the known-good reference:

acoustic validate --xml build/acoustic_dataset.xml
acoustic compare build/acoustic_dataset.xml examples/reference/trial_known_good.xml

Now make it your own.examples/calculation_input.xml is your sandbox — edit a value and re-run to watch the validated output change. Bump <BandCount>10 to 20, or <BaseLevelDb>140.0 to 155.0, then:

acoustic pipeline

The band count in the message and the generated XML move with your edit. This is the tightest loop in the project: input XML in, validated dataset out.

Then break it on purpose to watch the gates earn their keep. Set <BandCount> to a negative number, or put text where a number belongs, and re-run:

error: build rejected a value before serialisation: ... schema: value must be positive

The pipeline rejects it, prints the reason, exits non-zero, and does not write a stale artifact. (Restore the file afterward — git checkout examples/calculation_input.xml.)

Now read why:docs/concepts/two-verification-gates.md explains why "schema-valid" and "correct" are two different checks. The code lives in src/acoustic_dataset/validate.py and compare.py, dispatched from cli.py.

Adventure 2 — Explore a typed Python data class

The pipeline never carries data in a loose dict; it builds one typed object that meets the schema. Feel the difference yourself.

First, experience editor autocomplete. Open examples/explore_platform.py in VS Code (the Codespace ships Pylance). Find the # platform.radiated_noise. line, delete the #, put your cursor right after the dot, and type: VS Code pops up the object's declared attributes (band, …) — statically, before the code has even run. Try platform. too for the top-level attributes (radiated_noise, sensors, …). Keep drilling (platform.radiated_noise.band[0].) and it completes all the way down. That instant, schema-shaped autocomplete is the real payoff of declared fields. Run the file any time with:

python examples/explore_platform.py

If autocomplete doesn't appear, give Pylance a moment to index, and make sure VS Code's Python interpreter is set to the Codespace's 3.9 (bottom-right status bar / Python: Select Interpreter).

Then poke at it in the REPL:

python
>>>fromacoustic_datasetimportbuild, serialize>>>platform=build.build_platform_from_file("examples/calculation_input.xml")
>>>type(platform).__name__'Platform'>>>platform.radiated_noise.band[0].centre_frequency# a Decimal, not a string>>>print(serialize.to_xml(platform)[:500]) # the same object, as XML>>>fromacoustic_dataset.models.acoustic_datasetimportSector>>>Sector(bering=1, level=2) # a typo is a TypeError, not a silent new key

The REPL can also complete attributes, but that's a separate readline/rlcompleter feature that introspects the live object at runtime — not the static, declared-field completion you saw in the editor. If <TAB> doesn't complete, enable it for the session with import readline, rlcompleter; readline.parse_and_bind("tab: complete").

Read docs/concepts/typed-vs-dicts.md for what declared fields buy you over a dictionary, then look at how the builder rejects out-of-range values in src/acoustic_dataset/build.py.

Adventure 3 — Reverse-engineer the data classes from the schema

Those data classes aren't hand-written — they're generated from the XSD by xsdata. The XSD is the single source of truth; the Python is a build artifact (note the # DO NOT EDIT BY HAND header on every model file).

acoustic generate # regenerate models from every schema/*.xsd

Trace the chain for one element:

  1. Open schema/acoustic_dataset.xsd and find an element, e.g. Sector.
  2. Open the generated src/acoustic_dataset/models/acoustic_dataset.py and find the matching @dataclass. Notice how XSD types, ranges, and docs became Python field metadata.
  3. Read src/acoustic_dataset/generate.py to see the exact xsdata invocation — and the two post-processing steps that keep the output deterministic and Python-3.9-compatible.

Want to see generation in action? Change a <xs:documentation> string in the XSD, run acoustic generate, and git diff the models — the docstring tracks the schema. (Revert the XSD edit afterward; CI fails if committed models drift from the schema — ADR 0008.)

Adventure 4 — Generate and view the XSD documentation

Two complementary docs come out of this repo. First, a standalone HTML schema reference rendered straight from the XSD (via the vendored xs3p stylesheet):

acoustic gen-schema-docs --out build/schema-preview.html

Open build/schema-preview.html (in a Codespace: right-click the file → Download, or use the Live Preview extension) to browse every element, type, and constraint of the contract. The pre-built version is committed at docs/reference/schema/index.html. The generator lives in src/acoustic_dataset/schema_html.py.

Second, the full project site — tutorials, concepts, ADRs, and a Mermaid ERD — served by MkDocs Material:

make docs-serve # browse at http://localhost:8000

In a Codespace, when the port-forward notification appears, click Open in Browser.

Adventure 5 — Point it at the real (private) schema

Everything above runs against the placeholder schema committed in schema/. The real, proprietary XSD and corpus are meant to live under private/ — a directory that is entirely gitignored (see .gitignore) so it never reaches git, CI, or the internet.

One constraint shapes the workflow: src/acoustic_dataset/models/ is committed (the placeholder-generated models that CI drift-checks), so you must not regenerate over it from a real schema — an accidental commit would leak that structure. Generate real models into private/ instead. The same CLI takes explicit paths, so once the real material is in place (real XSD in private/schema/, inputs in private/examples/, known-good XML in private/reference/):

# Generate typed models from the real XSD into the gitignored output dir
acoustic generate --schema private/schema/<real>.xsd --out private/models
# Structural gate (XSD + round-trip) on a real file against the real schema
acoustic validate --xml private/examples/<real>.xml --schema private/schema/<real>.xsd
# Migration-safety diff: generated vs known-good reference
acoustic compare private/<generated>.xml private/reference/<known_good>.xml

Caveat — the full pipeline command.acoustic pipeline imports acoustic_dataset.models (the committed package), so it always builds against the placeholder-derived bindings even when you pass --schema private/.... Running the end-to-end pipeline on real-generated models would mean repointing that import at private/models/ — a deliberate change that touches committed code, so raise it for review first. With the real schema you can use generate, validate, and compare against private/ paths immediately; full pipeline needs that reviewed change.

Where to go next

Follow the guided tutorial that ties all of this together: docs/tutorials/01-start-here.md, then skim the decision records in docs/decisions/ — each says why a choice was made and what was rejected.

Documentation

The full, navigable documentation — tutorials, how-to guides, concepts, decision records (ADRs), and a Mermaid schema ERD — lives in docs/ and renders as an attractive HTML site via MkDocs Material (make docs).

Start here:docs/tutorials/01-start-here.md.

Spec-driven development

This project uses GitHub Spec Kit. The active feature's specification, plan, and design artifacts are under specs/001-codespace-xml-scaffold/. Use the /speckit-* skills (specify, plan, tasks, implement) to drive development.

Layout

PathWhat it is
.devcontainer/Codespaces environment (Python 3.9.4)
schema/The XSD contract
src/acoustic_dataset/The pipeline package (models/ is generated, never hand-edited)
examples/Example calculation input + reference/ known-good files
tests/Unit, integration, and golden-file tests
docs/The documentation site (MkDocs Material)
specs/Spec Kit feature specs, plans, and contracts

About

Playground to investigate/develop python tooling in support of complex XML Schema

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

python-playground

Playground for developing / demonstrating python dev techniques.

This repository hosts Acoustic Reference Book — Phase 1: a schema-driven pipeline that turns acoustic calculation output into validated XML Acoustic Dataset, built as a learning project where every decision is documented and defensible.

Quick start

Open in a GitHub Codespace (Code → Codespaces → Create codespace). It provisions Python 3.9.4 (matching the target system) and runs make bootstrap automatically. Then:

make verify # lint + tests — confirm the environment is green
make docs-serve # browse the docs at http://localhost:8000

Working locally instead? You need Python 3.9.x and make, then make bootstrap. Run make to see all targets.

Self-guided adventures

New here? Work through these in order. Each is a short, hands-on loop — run a command, look at what it produced, then peek at the code that did it. Everything below has been run end-to-end, so the output you see should match.

All commands assume make bootstrap has finished (it runs automatically in a Codespace). If acoustic isn't found, the package is exposed as a module too: replace acoustic with python -m acoustic_dataset.cli.

Adventure 1 — Run the whole pipeline and watch the gates

make verify # lint + type-check + tests: is the env green?
acoustic pipeline # input XML -> typed objects -> XML -> validate

You should see something like:

pipeline ok: 10 band(s) -> build/acoustic_dataset.xml (schema-valid, round-trip-equal)

Open the file it wrote (build/acoustic_dataset.xml) and compare it to the input (examples/calculation_input.xml). Then run the structural gate on its own and the migration-safety diff against the known-good reference:

acoustic validate --xml build/acoustic_dataset.xml
acoustic compare build/acoustic_dataset.xml examples/reference/trial_known_good.xml

Now make it your own.examples/calculation_input.xml is your sandbox — edit a value and re-run to watch the validated output change. Bump <BandCount>10 to 20, or <BaseLevelDb>140.0 to 155.0, then:

acoustic pipeline

The band count in the message and the generated XML move with your edit. This is the tightest loop in the project: input XML in, validated dataset out.

Then break it on purpose to watch the gates earn their keep. Set <BandCount> to a negative number, or put text where a number belongs, and re-run:

error: build rejected a value before serialisation: ... schema: value must be positive

The pipeline rejects it, prints the reason, exits non-zero, and does not write a stale artifact. (Restore the file afterward — git checkout examples/calculation_input.xml.)

Now read why:docs/concepts/two-verification-gates.md explains why "schema-valid" and "correct" are two different checks. The code lives in src/acoustic_dataset/validate.py and compare.py, dispatched from cli.py.

Adventure 2 — Explore a typed Python data class

The pipeline never carries data in a loose dict; it builds one typed object that meets the schema. Feel the difference yourself.

First, experience editor autocomplete. Open examples/explore_platform.py in VS Code (the Codespace ships Pylance). Find the # platform.radiated_noise. line, delete the #, put your cursor right after the dot, and type: VS Code pops up the object's declared attributes (band, …) — statically, before the code has even run. Try platform. too for the top-level attributes (radiated_noise, sensors, …). Keep drilling (platform.radiated_noise.band[0].) and it completes all the way down. That instant, schema-shaped autocomplete is the real payoff of declared fields. Run the file any time with:

python examples/explore_platform.py

If autocomplete doesn't appear, give Pylance a moment to index, and make sure VS Code's Python interpreter is set to the Codespace's 3.9 (bottom-right status bar / Python: Select Interpreter).

Then poke at it in the REPL:

python
>>>fromacoustic_datasetimportbuild, serialize>>>platform=build.build_platform_from_file("examples/calculation_input.xml")
>>>type(platform).__name__'Platform'>>>platform.radiated_noise.band[0].centre_frequency# a Decimal, not a string>>>print(serialize.to_xml(platform)[:500]) # the same object, as XML>>>fromacoustic_dataset.models.acoustic_datasetimportSector>>>Sector(bering=1, level=2) # a typo is a TypeError, not a silent new key

The REPL can also complete attributes, but that's a separate readline/rlcompleter feature that introspects the live object at runtime — not the static, declared-field completion you saw in the editor. If <TAB> doesn't complete, enable it for the session with import readline, rlcompleter; readline.parse_and_bind("tab: complete").

Read docs/concepts/typed-vs-dicts.md for what declared fields buy you over a dictionary, then look at how the builder rejects out-of-range values in src/acoustic_dataset/build.py.

Adventure 3 — Reverse-engineer the data classes from the schema

Those data classes aren't hand-written — they're generated from the XSD by xsdata. The XSD is the single source of truth; the Python is a build artifact (note the # DO NOT EDIT BY HAND header on every model file).

acoustic generate # regenerate models from every schema/*.xsd

Trace the chain for one element:

  1. Open schema/acoustic_dataset.xsd and find an element, e.g. Sector.
  2. Open the generated src/acoustic_dataset/models/acoustic_dataset.py and find the matching @dataclass. Notice how XSD types, ranges, and docs became Python field metadata.
  3. Read src/acoustic_dataset/generate.py to see the exact xsdata invocation — and the two post-processing steps that keep the output deterministic and Python-3.9-compatible.

Want to see generation in action? Change a <xs:documentation> string in the XSD, run acoustic generate, and git diff the models — the docstring tracks the schema. (Revert the XSD edit afterward; CI fails if committed models drift from the schema — ADR 0008.)

Adventure 4 — Generate and view the XSD documentation

Two complementary docs come out of this repo. First, a standalone HTML schema reference rendered straight from the XSD (via the vendored xs3p stylesheet):

acoustic gen-schema-docs --out build/schema-preview.html

Open build/schema-preview.html (in a Codespace: right-click the file → Download, or use the Live Preview extension) to browse every element, type, and constraint of the contract. The pre-built version is committed at docs/reference/schema/index.html. The generator lives in src/acoustic_dataset/schema_html.py.

Second, the full project site — tutorials, concepts, ADRs, and a Mermaid ERD — served by MkDocs Material:

make docs-serve # browse at http://localhost:8000

In a Codespace, when the port-forward notification appears, click Open in Browser.

Adventure 5 — Point it at the real (private) schema

Everything above runs against the placeholder schema committed in schema/. The real, proprietary XSD and corpus are meant to live under private/ — a directory that is entirely gitignored (see .gitignore) so it never reaches git, CI, or the internet.

One constraint shapes the workflow: src/acoustic_dataset/models/ is committed (the placeholder-generated models that CI drift-checks), so you must not regenerate over it from a real schema — an accidental commit would leak that structure. Generate real models into private/ instead. The same CLI takes explicit paths, so once the real material is in place (real XSD in private/schema/, inputs in private/examples/, known-good XML in private/reference/):

# Generate typed models from the real XSD into the gitignored output dir
acoustic generate --schema private/schema/<real>.xsd --out private/models
# Structural gate (XSD + round-trip) on a real file against the real schema
acoustic validate --xml private/examples/<real>.xml --schema private/schema/<real>.xsd
# Migration-safety diff: generated vs known-good reference
acoustic compare private/<generated>.xml private/reference/<known_good>.xml

Caveat — the full pipeline command.acoustic pipeline imports acoustic_dataset.models (the committed package), so it always builds against the placeholder-derived bindings even when you pass --schema private/.... Running the end-to-end pipeline on real-generated models would mean repointing that import at private/models/ — a deliberate change that touches committed code, so raise it for review first. With the real schema you can use generate, validate, and compare against private/ paths immediately; full pipeline needs that reviewed change.

Where to go next

Follow the guided tutorial that ties all of this together: docs/tutorials/01-start-here.md, then skim the decision records in docs/decisions/ — each says why a choice was made and what was rejected.

Documentation

The full, navigable documentation — tutorials, how-to guides, concepts, decision records (ADRs), and a Mermaid schema ERD — lives in docs/ and renders as an attractive HTML site via MkDocs Material (make docs).

Start here:docs/tutorials/01-start-here.md.

Spec-driven development

This project uses GitHub Spec Kit. The active feature's specification, plan, and design artifacts are under specs/001-codespace-xml-scaffold/. Use the /speckit-* skills (specify, plan, tasks, implement) to drive development.

Layout

PathWhat it is
.devcontainer/Codespaces environment (Python 3.9.4)
schema/The XSD contract
src/acoustic_dataset/The pipeline package (models/ is generated, never hand-edited)
examples/Example calculation input + reference/ known-good files
tests/Unit, integration, and golden-file tests
docs/The documentation site (MkDocs Material)
specs/Spec Kit feature specs, plans, and contracts

About

Playground to investigate/develop python tooling in support of complex XML Schema

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

python-playground

Playground for developing / demonstrating python dev techniques.

This repository hosts Acoustic Reference Book — Phase 1: a schema-driven pipeline that turns acoustic calculation output into validated XML Acoustic Dataset, built as a learning project where every decision is documented and defensible.

Quick start

Open in a GitHub Codespace (Code → Codespaces → Create codespace). It provisions Python 3.9.4 (matching the target system) and runs make bootstrap automatically. Then:

make verify # lint + tests — confirm the environment is green
make docs-serve # browse the docs at http://localhost:8000

Working locally instead? You need Python 3.9.x and make, then make bootstrap. Run make to see all targets.

Self-guided adventures

New here? Work through these in order. Each is a short, hands-on loop — run a command, look at what it produced, then peek at the code that did it. Everything below has been run end-to-end, so the output you see should match.

All commands assume make bootstrap has finished (it runs automatically in a Codespace). If acoustic isn't found, the package is exposed as a module too: replace acoustic with python -m acoustic_dataset.cli.

Adventure 1 — Run the whole pipeline and watch the gates

make verify # lint + type-check + tests: is the env green?
acoustic pipeline # input XML -> typed objects -> XML -> validate

You should see something like:

pipeline ok: 10 band(s) -> build/acoustic_dataset.xml (schema-valid, round-trip-equal)

Open the file it wrote (build/acoustic_dataset.xml) and compare it to the input (examples/calculation_input.xml). Then run the structural gate on its own and the migration-safety diff against the known-good reference:

acoustic validate --xml build/acoustic_dataset.xml
acoustic compare build/acoustic_dataset.xml examples/reference/trial_known_good.xml

Now make it your own.examples/calculation_input.xml is your sandbox — edit a value and re-run to watch the validated output change. Bump <BandCount>10 to 20, or <BaseLevelDb>140.0 to 155.0, then:

acoustic pipeline

The band count in the message and the generated XML move with your edit. This is the tightest loop in the project: input XML in, validated dataset out.

Then break it on purpose to watch the gates earn their keep. Set <BandCount> to a negative number, or put text where a number belongs, and re-run:

error: build rejected a value before serialisation: ... schema: value must be positive

The pipeline rejects it, prints the reason, exits non-zero, and does not write a stale artifact. (Restore the file afterward — git checkout examples/calculation_input.xml.)

Now read why:docs/concepts/two-verification-gates.md explains why "schema-valid" and "correct" are two different checks. The code lives in src/acoustic_dataset/validate.py and compare.py, dispatched from cli.py.

Adventure 2 — Explore a typed Python data class

The pipeline never carries data in a loose dict; it builds one typed object that meets the schema. Feel the difference yourself.

First, experience editor autocomplete. Open examples/explore_platform.py in VS Code (the Codespace ships Pylance). Find the # platform.radiated_noise. line, delete the #, put your cursor right after the dot, and type: VS Code pops up the object's declared attributes (band, …) — statically, before the code has even run. Try platform. too for the top-level attributes (radiated_noise, sensors, …). Keep drilling (platform.radiated_noise.band[0].) and it completes all the way down. That instant, schema-shaped autocomplete is the real payoff of declared fields. Run the file any time with:

python examples/explore_platform.py

If autocomplete doesn't appear, give Pylance a moment to index, and make sure VS Code's Python interpreter is set to the Codespace's 3.9 (bottom-right status bar / Python: Select Interpreter).

Then poke at it in the REPL:

python
>>>fromacoustic_datasetimportbuild, serialize>>>platform=build.build_platform_from_file("examples/calculation_input.xml")
>>>type(platform).__name__'Platform'>>>platform.radiated_noise.band[0].centre_frequency# a Decimal, not a string>>>print(serialize.to_xml(platform)[:500]) # the same object, as XML>>>fromacoustic_dataset.models.acoustic_datasetimportSector>>>Sector(bering=1, level=2) # a typo is a TypeError, not a silent new key

The REPL can also complete attributes, but that's a separate readline/rlcompleter feature that introspects the live object at runtime — not the static, declared-field completion you saw in the editor. If <TAB> doesn't complete, enable it for the session with import readline, rlcompleter; readline.parse_and_bind("tab: complete").

Read docs/concepts/typed-vs-dicts.md for what declared fields buy you over a dictionary, then look at how the builder rejects out-of-range values in src/acoustic_dataset/build.py.

Adventure 3 — Reverse-engineer the data classes from the schema

Those data classes aren't hand-written — they're generated from the XSD by xsdata. The XSD is the single source of truth; the Python is a build artifact (note the # DO NOT EDIT BY HAND header on every model file).

acoustic generate # regenerate models from every schema/*.xsd

Trace the chain for one element:

  1. Open schema/acoustic_dataset.xsd and find an element, e.g. Sector.
  2. Open the generated src/acoustic_dataset/models/acoustic_dataset.py and find the matching @dataclass. Notice how XSD types, ranges, and docs became Python field metadata.
  3. Read src/acoustic_dataset/generate.py to see the exact xsdata invocation — and the two post-processing steps that keep the output deterministic and Python-3.9-compatible.

Want to see generation in action? Change a <xs:documentation> string in the XSD, run acoustic generate, and git diff the models — the docstring tracks the schema. (Revert the XSD edit afterward; CI fails if committed models drift from the schema — ADR 0008.)

Adventure 4 — Generate and view the XSD documentation

Two complementary docs come out of this repo. First, a standalone HTML schema reference rendered straight from the XSD (via the vendored xs3p stylesheet):

acoustic gen-schema-docs --out build/schema-preview.html

Open build/schema-preview.html (in a Codespace: right-click the file → Download, or use the Live Preview extension) to browse every element, type, and constraint of the contract. The pre-built version is committed at docs/reference/schema/index.html. The generator lives in src/acoustic_dataset/schema_html.py.

Second, the full project site — tutorials, concepts, ADRs, and a Mermaid ERD — served by MkDocs Material:

make docs-serve # browse at http://localhost:8000

In a Codespace, when the port-forward notification appears, click Open in Browser.

Adventure 5 — Point it at the real (private) schema

Everything above runs against the placeholder schema committed in schema/. The real, proprietary XSD and corpus are meant to live under private/ — a directory that is entirely gitignored (see .gitignore) so it never reaches git, CI, or the internet.

One constraint shapes the workflow: src/acoustic_dataset/models/ is committed (the placeholder-generated models that CI drift-checks), so you must not regenerate over it from a real schema — an accidental commit would leak that structure. Generate real models into private/ instead. The same CLI takes explicit paths, so once the real material is in place (real XSD in private/schema/, inputs in private/examples/, known-good XML in private/reference/):

# Generate typed models from the real XSD into the gitignored output dir
acoustic generate --schema private/schema/<real>.xsd --out private/models
# Structural gate (XSD + round-trip) on a real file against the real schema
acoustic validate --xml private/examples/<real>.xml --schema private/schema/<real>.xsd
# Migration-safety diff: generated vs known-good reference
acoustic compare private/<generated>.xml private/reference/<known_good>.xml

Caveat — the full pipeline command.acoustic pipeline imports acoustic_dataset.models (the committed package), so it always builds against the placeholder-derived bindings even when you pass --schema private/.... Running the end-to-end pipeline on real-generated models would mean repointing that import at private/models/ — a deliberate change that touches committed code, so raise it for review first. With the real schema you can use generate, validate, and compare against private/ paths immediately; full pipeline needs that reviewed change.

Where to go next

Follow the guided tutorial that ties all of this together: docs/tutorials/01-start-here.md, then skim the decision records in docs/decisions/ — each says why a choice was made and what was rejected.

Documentation

The full, navigable documentation — tutorials, how-to guides, concepts, decision records (ADRs), and a Mermaid schema ERD — lives in docs/ and renders as an attractive HTML site via MkDocs Material (make docs).

Start here:docs/tutorials/01-start-here.md.

Spec-driven development

This project uses GitHub Spec Kit. The active feature's specification, plan, and design artifacts are under specs/001-codespace-xml-scaffold/. Use the /speckit-* skills (specify, plan, tasks, implement) to drive development.

Layout

PathWhat it is
.devcontainer/Codespaces environment (Python 3.9.4)
schema/The XSD contract
src/acoustic_dataset/The pipeline package (models/ is generated, never hand-edited)
examples/Example calculation input + reference/ known-good files
tests/Unit, integration, and golden-file tests
docs/The documentation site (MkDocs Material)
specs/Spec Kit feature specs, plans, and contracts

About

Playground to investigate/develop python tooling in support of complex XML Schema

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

python-playground

Playground for developing / demonstrating python dev techniques.

This repository hosts Acoustic Reference Book — Phase 1: a schema-driven pipeline that turns acoustic calculation output into validated XML Acoustic Dataset, built as a learning project where every decision is documented and defensible.

Quick start

Open in a GitHub Codespace (Code → Codespaces → Create codespace). It provisions Python 3.9.4 (matching the target system) and runs make bootstrap automatically. Then:

make verify # lint + tests — confirm the environment is green
make docs-serve # browse the docs at http://localhost:8000

Working locally instead? You need Python 3.9.x and make, then make bootstrap. Run make to see all targets.

Self-guided adventures

New here? Work through these in order. Each is a short, hands-on loop — run a command, look at what it produced, then peek at the code that did it. Everything below has been run end-to-end, so the output you see should match.

All commands assume make bootstrap has finished (it runs automatically in a Codespace). If acoustic isn't found, the package is exposed as a module too: replace acoustic with python -m acoustic_dataset.cli.

Adventure 1 — Run the whole pipeline and watch the gates

make verify # lint + type-check + tests: is the env green?
acoustic pipeline # input XML -> typed objects -> XML -> validate

You should see something like:

pipeline ok: 10 band(s) -> build/acoustic_dataset.xml (schema-valid, round-trip-equal)

Open the file it wrote (build/acoustic_dataset.xml) and compare it to the input (examples/calculation_input.xml). Then run the structural gate on its own and the migration-safety diff against the known-good reference:

acoustic validate --xml build/acoustic_dataset.xml
acoustic compare build/acoustic_dataset.xml examples/reference/trial_known_good.xml

Now make it your own.examples/calculation_input.xml is your sandbox — edit a value and re-run to watch the validated output change. Bump <BandCount>10 to 20, or <BaseLevelDb>140.0 to 155.0, then:

acoustic pipeline

The band count in the message and the generated XML move with your edit. This is the tightest loop in the project: input XML in, validated dataset out.

Then break it on purpose to watch the gates earn their keep. Set <BandCount> to a negative number, or put text where a number belongs, and re-run:

error: build rejected a value before serialisation: ... schema: value must be positive

The pipeline rejects it, prints the reason, exits non-zero, and does not write a stale artifact. (Restore the file afterward — git checkout examples/calculation_input.xml.)

Now read why:docs/concepts/two-verification-gates.md explains why "schema-valid" and "correct" are two different checks. The code lives in src/acoustic_dataset/validate.py and compare.py, dispatched from cli.py.

Adventure 2 — Explore a typed Python data class

The pipeline never carries data in a loose dict; it builds one typed object that meets the schema. Feel the difference yourself.

First, experience editor autocomplete. Open examples/explore_platform.py in VS Code (the Codespace ships Pylance). Find the # platform.radiated_noise. line, delete the #, put your cursor right after the dot, and type: VS Code pops up the object's declared attributes (band, …) — statically, before the code has even run. Try platform. too for the top-level attributes (radiated_noise, sensors, …). Keep drilling (platform.radiated_noise.band[0].) and it completes all the way down. That instant, schema-shaped autocomplete is the real payoff of declared fields. Run the file any time with:

python examples/explore_platform.py

If autocomplete doesn't appear, give Pylance a moment to index, and make sure VS Code's Python interpreter is set to the Codespace's 3.9 (bottom-right status bar / Python: Select Interpreter).

Then poke at it in the REPL:

python
>>>fromacoustic_datasetimportbuild, serialize>>>platform=build.build_platform_from_file("examples/calculation_input.xml")
>>>type(platform).__name__'Platform'>>>platform.radiated_noise.band[0].centre_frequency# a Decimal, not a string>>>print(serialize.to_xml(platform)[:500]) # the same object, as XML>>>fromacoustic_dataset.models.acoustic_datasetimportSector>>>Sector(bering=1, level=2) # a typo is a TypeError, not a silent new key

The REPL can also complete attributes, but that's a separate readline/rlcompleter feature that introspects the live object at runtime — not the static, declared-field completion you saw in the editor. If <TAB> doesn't complete, enable it for the session with import readline, rlcompleter; readline.parse_and_bind("tab: complete").

Read docs/concepts/typed-vs-dicts.md for what declared fields buy you over a dictionary, then look at how the builder rejects out-of-range values in src/acoustic_dataset/build.py.

Adventure 3 — Reverse-engineer the data classes from the schema

Those data classes aren't hand-written — they're generated from the XSD by xsdata. The XSD is the single source of truth; the Python is a build artifact (note the # DO NOT EDIT BY HAND header on every model file).

acoustic generate # regenerate models from every schema/*.xsd

Trace the chain for one element:

  1. Open schema/acoustic_dataset.xsd and find an element, e.g. Sector.
  2. Open the generated src/acoustic_dataset/models/acoustic_dataset.py and find the matching @dataclass. Notice how XSD types, ranges, and docs became Python field metadata.
  3. Read src/acoustic_dataset/generate.py to see the exact xsdata invocation — and the two post-processing steps that keep the output deterministic and Python-3.9-compatible.

Want to see generation in action? Change a <xs:documentation> string in the XSD, run acoustic generate, and git diff the models — the docstring tracks the schema. (Revert the XSD edit afterward; CI fails if committed models drift from the schema — ADR 0008.)

Adventure 4 — Generate and view the XSD documentation

Two complementary docs come out of this repo. First, a standalone HTML schema reference rendered straight from the XSD (via the vendored xs3p stylesheet):

acoustic gen-schema-docs --out build/schema-preview.html

Open build/schema-preview.html (in a Codespace: right-click the file → Download, or use the Live Preview extension) to browse every element, type, and constraint of the contract. The pre-built version is committed at docs/reference/schema/index.html. The generator lives in src/acoustic_dataset/schema_html.py.

Second, the full project site — tutorials, concepts, ADRs, and a Mermaid ERD — served by MkDocs Material:

make docs-serve # browse at http://localhost:8000

In a Codespace, when the port-forward notification appears, click Open in Browser.

Adventure 5 — Point it at the real (private) schema

Everything above runs against the placeholder schema committed in schema/. The real, proprietary XSD and corpus are meant to live under private/ — a directory that is entirely gitignored (see .gitignore) so it never reaches git, CI, or the internet.

One constraint shapes the workflow: src/acoustic_dataset/models/ is committed (the placeholder-generated models that CI drift-checks), so you must not regenerate over it from a real schema — an accidental commit would leak that structure. Generate real models into private/ instead. The same CLI takes explicit paths, so once the real material is in place (real XSD in private/schema/, inputs in private/examples/, known-good XML in private/reference/):

# Generate typed models from the real XSD into the gitignored output dir
acoustic generate --schema private/schema/<real>.xsd --out private/models
# Structural gate (XSD + round-trip) on a real file against the real schema
acoustic validate --xml private/examples/<real>.xml --schema private/schema/<real>.xsd
# Migration-safety diff: generated vs known-good reference
acoustic compare private/<generated>.xml private/reference/<known_good>.xml

Caveat — the full pipeline command.acoustic pipeline imports acoustic_dataset.models (the committed package), so it always builds against the placeholder-derived bindings even when you pass --schema private/.... Running the end-to-end pipeline on real-generated models would mean repointing that import at private/models/ — a deliberate change that touches committed code, so raise it for review first. With the real schema you can use generate, validate, and compare against private/ paths immediately; full pipeline needs that reviewed change.

Where to go next

Follow the guided tutorial that ties all of this together: docs/tutorials/01-start-here.md, then skim the decision records in docs/decisions/ — each says why a choice was made and what was rejected.

Documentation

The full, navigable documentation — tutorials, how-to guides, concepts, decision records (ADRs), and a Mermaid schema ERD — lives in docs/ and renders as an attractive HTML site via MkDocs Material (make docs).

Start here:docs/tutorials/01-start-here.md.

Spec-driven development

This project uses GitHub Spec Kit. The active feature's specification, plan, and design artifacts are under specs/001-codespace-xml-scaffold/. Use the /speckit-* skills (specify, plan, tasks, implement) to drive development.

Layout

PathWhat it is
.devcontainer/Codespaces environment (Python 3.9.4)
schema/The XSD contract
src/acoustic_dataset/The pipeline package (models/ is generated, never hand-edited)
examples/Example calculation input + reference/ known-good files
tests/Unit, integration, and golden-file tests
docs/The documentation site (MkDocs Material)
specs/Spec Kit feature specs, plans, and contracts

About

Playground to investigate/develop python tooling in support of complex XML Schema

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

python-playground

Playground for developing / demonstrating python dev techniques.

This repository hosts Acoustic Reference Book — Phase 1: a schema-driven pipeline that turns acoustic calculation output into validated XML Acoustic Dataset, built as a learning project where every decision is documented and defensible.

Quick start

Open in a GitHub Codespace (Code → Codespaces → Create codespace). It provisions Python 3.9.4 (matching the target system) and runs make bootstrap automatically. Then:

make verify # lint + tests — confirm the environment is green
make docs-serve # browse the docs at http://localhost:8000

Working locally instead? You need Python 3.9.x and make, then make bootstrap. Run make to see all targets.

Self-guided adventures

New here? Work through these in order. Each is a short, hands-on loop — run a command, look at what it produced, then peek at the code that did it. Everything below has been run end-to-end, so the output you see should match.

All commands assume make bootstrap has finished (it runs automatically in a Codespace). If acoustic isn't found, the package is exposed as a module too: replace acoustic with python -m acoustic_dataset.cli.

Adventure 1 — Run the whole pipeline and watch the gates

make verify # lint + type-check + tests: is the env green?
acoustic pipeline # input XML -> typed objects -> XML -> validate

You should see something like:

pipeline ok: 10 band(s) -> build/acoustic_dataset.xml (schema-valid, round-trip-equal)

Open the file it wrote (build/acoustic_dataset.xml) and compare it to the input (examples/calculation_input.xml). Then run the structural gate on its own and the migration-safety diff against the known-good reference:

acoustic validate --xml build/acoustic_dataset.xml
acoustic compare build/acoustic_dataset.xml examples/reference/trial_known_good.xml

Now make it your own.examples/calculation_input.xml is your sandbox — edit a value and re-run to watch the validated output change. Bump <BandCount>10 to 20, or <BaseLevelDb>140.0 to 155.0, then:

acoustic pipeline

The band count in the message and the generated XML move with your edit. This is the tightest loop in the project: input XML in, validated dataset out.

Then break it on purpose to watch the gates earn their keep. Set <BandCount> to a negative number, or put text where a number belongs, and re-run:

error: build rejected a value before serialisation: ... schema: value must be positive

The pipeline rejects it, prints the reason, exits non-zero, and does not write a stale artifact. (Restore the file afterward — git checkout examples/calculation_input.xml.)

Now read why:docs/concepts/two-verification-gates.md explains why "schema-valid" and "correct" are two different checks. The code lives in src/acoustic_dataset/validate.py and compare.py, dispatched from cli.py.

Adventure 2 — Explore a typed Python data class

The pipeline never carries data in a loose dict; it builds one typed object that meets the schema. Feel the difference yourself.

First, experience editor autocomplete. Open examples/explore_platform.py in VS Code (the Codespace ships Pylance). Find the # platform.radiated_noise. line, delete the #, put your cursor right after the dot, and type: VS Code pops up the object's declared attributes (band, …) — statically, before the code has even run. Try platform. too for the top-level attributes (radiated_noise, sensors, …). Keep drilling (platform.radiated_noise.band[0].) and it completes all the way down. That instant, schema-shaped autocomplete is the real payoff of declared fields. Run the file any time with:

python examples/explore_platform.py

If autocomplete doesn't appear, give Pylance a moment to index, and make sure VS Code's Python interpreter is set to the Codespace's 3.9 (bottom-right status bar / Python: Select Interpreter).

Then poke at it in the REPL:

python
>>>fromacoustic_datasetimportbuild, serialize>>>platform=build.build_platform_from_file("examples/calculation_input.xml")
>>>type(platform).__name__'Platform'>>>platform.radiated_noise.band[0].centre_frequency# a Decimal, not a string>>>print(serialize.to_xml(platform)[:500]) # the same object, as XML>>>fromacoustic_dataset.models.acoustic_datasetimportSector>>>Sector(bering=1, level=2) # a typo is a TypeError, not a silent new key

The REPL can also complete attributes, but that's a separate readline/rlcompleter feature that introspects the live object at runtime — not the static, declared-field completion you saw in the editor. If <TAB> doesn't complete, enable it for the session with import readline, rlcompleter; readline.parse_and_bind("tab: complete").

Read docs/concepts/typed-vs-dicts.md for what declared fields buy you over a dictionary, then look at how the builder rejects out-of-range values in src/acoustic_dataset/build.py.

Adventure 3 — Reverse-engineer the data classes from the schema

Those data classes aren't hand-written — they're generated from the XSD by xsdata. The XSD is the single source of truth; the Python is a build artifact (note the # DO NOT EDIT BY HAND header on every model file).

acoustic generate # regenerate models from every schema/*.xsd

Trace the chain for one element:

  1. Open schema/acoustic_dataset.xsd and find an element, e.g. Sector.
  2. Open the generated src/acoustic_dataset/models/acoustic_dataset.py and find the matching @dataclass. Notice how XSD types, ranges, and docs became Python field metadata.
  3. Read src/acoustic_dataset/generate.py to see the exact xsdata invocation — and the two post-processing steps that keep the output deterministic and Python-3.9-compatible.

Want to see generation in action? Change a <xs:documentation> string in the XSD, run acoustic generate, and git diff the models — the docstring tracks the schema. (Revert the XSD edit afterward; CI fails if committed models drift from the schema — ADR 0008.)

Adventure 4 — Generate and view the XSD documentation

Two complementary docs come out of this repo. First, a standalone HTML schema reference rendered straight from the XSD (via the vendored xs3p stylesheet):

acoustic gen-schema-docs --out build/schema-preview.html

Open build/schema-preview.html (in a Codespace: right-click the file → Download, or use the Live Preview extension) to browse every element, type, and constraint of the contract. The pre-built version is committed at docs/reference/schema/index.html. The generator lives in src/acoustic_dataset/schema_html.py.

Second, the full project site — tutorials, concepts, ADRs, and a Mermaid ERD — served by MkDocs Material:

make docs-serve # browse at http://localhost:8000

In a Codespace, when the port-forward notification appears, click Open in Browser.

Adventure 5 — Point it at the real (private) schema

Everything above runs against the placeholder schema committed in schema/. The real, proprietary XSD and corpus are meant to live under private/ — a directory that is entirely gitignored (see .gitignore) so it never reaches git, CI, or the internet.

One constraint shapes the workflow: src/acoustic_dataset/models/ is committed (the placeholder-generated models that CI drift-checks), so you must not regenerate over it from a real schema — an accidental commit would leak that structure. Generate real models into private/ instead. The same CLI takes explicit paths, so once the real material is in place (real XSD in private/schema/, inputs in private/examples/, known-good XML in private/reference/):

# Generate typed models from the real XSD into the gitignored output dir
acoustic generate --schema private/schema/<real>.xsd --out private/models
# Structural gate (XSD + round-trip) on a real file against the real schema
acoustic validate --xml private/examples/<real>.xml --schema private/schema/<real>.xsd
# Migration-safety diff: generated vs known-good reference
acoustic compare private/<generated>.xml private/reference/<known_good>.xml

Caveat — the full pipeline command.acoustic pipeline imports acoustic_dataset.models (the committed package), so it always builds against the placeholder-derived bindings even when you pass --schema private/.... Running the end-to-end pipeline on real-generated models would mean repointing that import at private/models/ — a deliberate change that touches committed code, so raise it for review first. With the real schema you can use generate, validate, and compare against private/ paths immediately; full pipeline needs that reviewed change.

Where to go next

Follow the guided tutorial that ties all of this together: docs/tutorials/01-start-here.md, then skim the decision records in docs/decisions/ — each says why a choice was made and what was rejected.

Documentation

The full, navigable documentation — tutorials, how-to guides, concepts, decision records (ADRs), and a Mermaid schema ERD — lives in docs/ and renders as an attractive HTML site via MkDocs Material (make docs).

Start here:docs/tutorials/01-start-here.md.

Spec-driven development

This project uses GitHub Spec Kit. The active feature's specification, plan, and design artifacts are under specs/001-codespace-xml-scaffold/. Use the /speckit-* skills (specify, plan, tasks, implement) to drive development.

Layout

PathWhat it is
.devcontainer/Codespaces environment (Python 3.9.4)
schema/The XSD contract
src/acoustic_dataset/The pipeline package (models/ is generated, never hand-edited)
examples/Example calculation input + reference/ known-good files
tests/Unit, integration, and golden-file tests
docs/The documentation site (MkDocs Material)
specs/Spec Kit feature specs, plans, and contracts

About

Playground to investigate/develop python tooling in support of complex XML Schema

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

python-playground

Playground for developing / demonstrating python dev techniques.

This repository hosts Acoustic Reference Book — Phase 1: a schema-driven pipeline that turns acoustic calculation output into validated XML Acoustic Dataset, built as a learning project where every decision is documented and defensible.

Quick start

Open in a GitHub Codespace (Code → Codespaces → Create codespace). It provisions Python 3.9.4 (matching the target system) and runs make bootstrap automatically. Then:

make verify # lint + tests — confirm the environment is green
make docs-serve # browse the docs at http://localhost:8000

Working locally instead? You need Python 3.9.x and make, then make bootstrap. Run make to see all targets.

Self-guided adventures

New here? Work through these in order. Each is a short, hands-on loop — run a command, look at what it produced, then peek at the code that did it. Everything below has been run end-to-end, so the output you see should match.

All commands assume make bootstrap has finished (it runs automatically in a Codespace). If acoustic isn't found, the package is exposed as a module too: replace acoustic with python -m acoustic_dataset.cli.

Adventure 1 — Run the whole pipeline and watch the gates

make verify # lint + type-check + tests: is the env green?
acoustic pipeline # input XML -> typed objects -> XML -> validate

You should see something like:

pipeline ok: 10 band(s) -> build/acoustic_dataset.xml (schema-valid, round-trip-equal)

Open the file it wrote (build/acoustic_dataset.xml) and compare it to the input (examples/calculation_input.xml). Then run the structural gate on its own and the migration-safety diff against the known-good reference:

acoustic validate --xml build/acoustic_dataset.xml
acoustic compare build/acoustic_dataset.xml examples/reference/trial_known_good.xml

Now make it your own.examples/calculation_input.xml is your sandbox — edit a value and re-run to watch the validated output change. Bump <BandCount>10 to 20, or <BaseLevelDb>140.0 to 155.0, then:

acoustic pipeline

The band count in the message and the generated XML move with your edit. This is the tightest loop in the project: input XML in, validated dataset out.

Then break it on purpose to watch the gates earn their keep. Set <BandCount> to a negative number, or put text where a number belongs, and re-run:

error: build rejected a value before serialisation: ... schema: value must be positive

The pipeline rejects it, prints the reason, exits non-zero, and does not write a stale artifact. (Restore the file afterward — git checkout examples/calculation_input.xml.)

Now read why:docs/concepts/two-verification-gates.md explains why "schema-valid" and "correct" are two different checks. The code lives in src/acoustic_dataset/validate.py and compare.py, dispatched from cli.py.

Adventure 2 — Explore a typed Python data class

The pipeline never carries data in a loose dict; it builds one typed object that meets the schema. Feel the difference yourself.

First, experience editor autocomplete. Open examples/explore_platform.py in VS Code (the Codespace ships Pylance). Find the # platform.radiated_noise. line, delete the #, put your cursor right after the dot, and type: VS Code pops up the object's declared attributes (band, …) — statically, before the code has even run. Try platform. too for the top-level attributes (radiated_noise, sensors, …). Keep drilling (platform.radiated_noise.band[0].) and it completes all the way down. That instant, schema-shaped autocomplete is the real payoff of declared fields. Run the file any time with:

python examples/explore_platform.py

If autocomplete doesn't appear, give Pylance a moment to index, and make sure VS Code's Python interpreter is set to the Codespace's 3.9 (bottom-right status bar / Python: Select Interpreter).

Then poke at it in the REPL:

python
>>>fromacoustic_datasetimportbuild, serialize>>>platform=build.build_platform_from_file("examples/calculation_input.xml")
>>>type(platform).__name__'Platform'>>>platform.radiated_noise.band[0].centre_frequency# a Decimal, not a string>>>print(serialize.to_xml(platform)[:500]) # the same object, as XML>>>fromacoustic_dataset.models.acoustic_datasetimportSector>>>Sector(bering=1, level=2) # a typo is a TypeError, not a silent new key

The REPL can also complete attributes, but that's a separate readline/rlcompleter feature that introspects the live object at runtime — not the static, declared-field completion you saw in the editor. If <TAB> doesn't complete, enable it for the session with import readline, rlcompleter; readline.parse_and_bind("tab: complete").

Read docs/concepts/typed-vs-dicts.md for what declared fields buy you over a dictionary, then look at how the builder rejects out-of-range values in src/acoustic_dataset/build.py.

Adventure 3 — Reverse-engineer the data classes from the schema

Those data classes aren't hand-written — they're generated from the XSD by xsdata. The XSD is the single source of truth; the Python is a build artifact (note the # DO NOT EDIT BY HAND header on every model file).

acoustic generate # regenerate models from every schema/*.xsd

Trace the chain for one element:

  1. Open schema/acoustic_dataset.xsd and find an element, e.g. Sector.
  2. Open the generated src/acoustic_dataset/models/acoustic_dataset.py and find the matching @dataclass. Notice how XSD types, ranges, and docs became Python field metadata.
  3. Read src/acoustic_dataset/generate.py to see the exact xsdata invocation — and the two post-processing steps that keep the output deterministic and Python-3.9-compatible.

Want to see generation in action? Change a <xs:documentation> string in the XSD, run acoustic generate, and git diff the models — the docstring tracks the schema. (Revert the XSD edit afterward; CI fails if committed models drift from the schema — ADR 0008.)

Adventure 4 — Generate and view the XSD documentation

Two complementary docs come out of this repo. First, a standalone HTML schema reference rendered straight from the XSD (via the vendored xs3p stylesheet):

acoustic gen-schema-docs --out build/schema-preview.html

Open build/schema-preview.html (in a Codespace: right-click the file → Download, or use the Live Preview extension) to browse every element, type, and constraint of the contract. The pre-built version is committed at docs/reference/schema/index.html. The generator lives in src/acoustic_dataset/schema_html.py.

Second, the full project site — tutorials, concepts, ADRs, and a Mermaid ERD — served by MkDocs Material:

make docs-serve # browse at http://localhost:8000

In a Codespace, when the port-forward notification appears, click Open in Browser.

Adventure 5 — Point it at the real (private) schema

Everything above runs against the placeholder schema committed in schema/. The real, proprietary XSD and corpus are meant to live under private/ — a directory that is entirely gitignored (see .gitignore) so it never reaches git, CI, or the internet.

One constraint shapes the workflow: src/acoustic_dataset/models/ is committed (the placeholder-generated models that CI drift-checks), so you must not regenerate over it from a real schema — an accidental commit would leak that structure. Generate real models into private/ instead. The same CLI takes explicit paths, so once the real material is in place (real XSD in private/schema/, inputs in private/examples/, known-good XML in private/reference/):

# Generate typed models from the real XSD into the gitignored output dir
acoustic generate --schema private/schema/<real>.xsd --out private/models
# Structural gate (XSD + round-trip) on a real file against the real schema
acoustic validate --xml private/examples/<real>.xml --schema private/schema/<real>.xsd
# Migration-safety diff: generated vs known-good reference
acoustic compare private/<generated>.xml private/reference/<known_good>.xml

Caveat — the full pipeline command.acoustic pipeline imports acoustic_dataset.models (the committed package), so it always builds against the placeholder-derived bindings even when you pass --schema private/.... Running the end-to-end pipeline on real-generated models would mean repointing that import at private/models/ — a deliberate change that touches committed code, so raise it for review first. With the real schema you can use generate, validate, and compare against private/ paths immediately; full pipeline needs that reviewed change.

Where to go next

Follow the guided tutorial that ties all of this together: docs/tutorials/01-start-here.md, then skim the decision records in docs/decisions/ — each says why a choice was made and what was rejected.

Documentation

The full, navigable documentation — tutorials, how-to guides, concepts, decision records (ADRs), and a Mermaid schema ERD — lives in docs/ and renders as an attractive HTML site via MkDocs Material (make docs).

Start here:docs/tutorials/01-start-here.md.

Spec-driven development

This project uses GitHub Spec Kit. The active feature's specification, plan, and design artifacts are under specs/001-codespace-xml-scaffold/. Use the /speckit-* skills (specify, plan, tasks, implement) to drive development.

Layout

PathWhat it is
.devcontainer/Codespaces environment (Python 3.9.4)
schema/The XSD contract
src/acoustic_dataset/The pipeline package (models/ is generated, never hand-edited)
examples/Example calculation input + reference/ known-good files
tests/Unit, integration, and golden-file tests
docs/The documentation site (MkDocs Material)
specs/Spec Kit feature specs, plans, and contracts

About

Playground to investigate/develop python tooling in support of complex XML Schema

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages