Repository files navigation

Babel Validator

This repository has several tools for validating the outputs from Babel runs, which are the underlying data used for the Translator Node Normalization and Name Resolver services.

PyTest

The best tests in this repository are Python tests stored in the ./tests folder. This includes both unit tests as well as "Google Sheet"-based tests, which use the shared Babel Validation Google Sheet containing facts that we can use to test a NodeNorm instance. The sheet's ID is deliberately not checked in: copy env.default to .env and fill in BABEL_VALIDATION_SHEET_ID (ask a maintainer for the ID; in GitHub Actions it comes from a repository secret of the same name). env.default documents every variable this repository reads, and what each one turns on.

To run these tests, you need to install uv. You can then use uv to run the tests. The file tests/targets.ini allows you to control which NodeNorm instance is tested. The [DEFAULT] section applies defaults for all the environments. For example, to run all the tests on the dev instance, you can use --target:

$ pytest --target dev
============================= test session starts ==============================
platform darwin -- Python 3.13.3, pytest-8.3.3, pluggy-1.5.0
testing target 'dev': {'nodenormurl': 'https://nodenormalization-sri.renci.org/', 'nameresurl': 'https://name-resolution-sri.renci.org/', 'namereslimit': '20', 'nameresxfailifintop': '5'}
included categories: set()
excluded categories: set()
rootdir: /Users/gaurav/Developer/translator/babel-validation
collected 4338 items [...]

Google Tests have a Category column. To filter based on this column, you can specify a --category on the command line.

$ pytest --target dev --category "Unit Tests" tests/nodenorm/test_nodenorm_from_gsheet.py
==================================================================== test session starts ====================================================================
platform darwin -- Python 3.13.3, pytest-8.3.3, pluggy-1.5.0
testing target 'dev': {'nodenormurl': 'https://nodenormalization-sri.renci.org/', 'nameresurl': 'https://name-resolution-sri.renci.org/', 'namereslimit': '20', 'nameresxfailifintop': '5'}
included categories: {'Unit Tests'}
excluded categories: set()
rootdir: /Users/gaurav/Developer/translator/babel-validation/tests
configfile: pytest.ini
collected 2010 items tests/nodenorm/test_nodenorm_from_gsheet.py sssssxsssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss.ss.x.....sssssssssssssssssssss.ssssss [ 5%]
ssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss...........ssss.....ss.........s...x..sxsssssssssss.ssssss..sssssssssssssssssss [ 12%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 20%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 27%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 34%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 42%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 49%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 57%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 64%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 71%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 79%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 86%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 94%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [100%]
======================================================= 41 passed, 1965 skipped, 4 xfailed in 10.11s ========================================================

GitHub issue tests

Assertions can also be embedded directly in GitHub issue bodies — see src/babel_validation/assertions/README.md for the syntax and the available assertion types. The repositories scanned for them are listed under Repositories in the [DEFAULT] section of tests/targets.ini.

Issue bodies are untrusted input, so the harness caps what one issue may contain — 100 assertions, 1,000 params lists, 1,000 parameters, 1,000 characters per parameter — and rejects YAML anchors, aliases and duplicate keys. An issue over a cap fails loudly rather than running part of itself; split it into several issues. The caps are listed in src/babel_validation/assertions/README.md. --issue resolves only within the configured Repositories, so a run can never be pointed at assertions from somewhere else.

Beware when discussing the syntax in an issue: a complete {{BabelTest|...}} marker is picked up wherever it appears, backticks included, and an unrecognised assertion name fails the run rather than being ignored. Quote a partial marker instead — the pattern needs the closing }} to match.

$ pytest tests/github_issues --target dev # every issue carrying assertions
$ pytest tests/github_issues --target dev --issue 'org/repo#42'# just one (also 'repo#42' or '42')

These tests need a GITHUB_TOKEN, in the environment or in a .env file. Without one they skip rather than fail, so a run can look green having tested nothing. Generate a personal access token; inside a GitHub Action, use the automatic GITHUB_TOKEN instead.

The token is not needed for authentication as such — every repository we scan is public, and both the single-issue and search endpoints answer unauthenticated requests. It is needed for the rate limits:

UnauthenticatedWith a token
Core60 / hour, per IP5,000 / hour
Search10 / minute30 / minute

Discovery is search-bound, not core-bound: two searches per configured repository (one per trigger keyword, plus a request per extra page of results), and then no core request at all, because a search result already carries the issue body and html_url the harness needs. Scanning the five configured repositories currently finds 96 issues for zero core requests.

Core requests are spent re-hydrating issues one at a time, which happens whenever the cached ID list is reused instead of the search being repeated — notably in every pytest-xdist worker after the first. That path costs one request per issue per worker, so an unauthenticated run would exhaust the 60/hour core budget well before finishing.

GET /rate_limit reports what is left without itself counting against the limit (docs). Note that the search window resets every 60 seconds, so its counter is often back at zero by the time you look:

$ curl -s -H "Authorization: Bearer $GITHUB_TOKEN" https://api.github.com/rate_limit

Log Analysis

The Jupyter Notebook in log-analysis/ contains some basic analysis of the logs from NodeNorm (and, someday, NameRes) instances.

The dashboard website

The Astro site in website/ is deployed to https://translatorsri.github.io/babel-validation/. It shows the results of running this test suite against every environment in tests/targets.ini, alongside each environment's /status information (Babel version, database sizes, NameRes latency). Because test expectations are pinned to the environment where a new Babel version lands first, environments are not expected to all be green — the dashboard's purpose is to show which issues are visible in which environment.

The .github/workflows/dashboard.yaml workflow regenerates and deploys it daily (or on manual dispatch): it runs pytest per target with --report-jsonl, turns the raw outcomes and /status responses into report.json and history.jsonl with src.babel_validation.tools.generate_report, and publishes the built site to the gh-pages branch.

The Vue components' client-side logic (URL round-tripping, filtering, pagination, the odd-one-out shading, the drift grouping, and the withholding of blocklist detail) has vitest tests in website/test/, run by npm test and by the Tests workflow.

To work on the site against the data the live dashboard is showing, download the published report.json and history.jsonl instead of generating them (both land in website/public/data/, which is gitignored):

$ cd website && npm install && npm run fetch-data && npm run dev

To regenerate it locally against a couple of environments:

$ uv run pytest tests/nodenorm/test_nodenorm_from_gsheet.py tests/nameres/test_nameres_from_gsheet.py \
--target dev --target prod -n 8 --report-jsonl raw/local.jsonl
$ uv run python -m src.babel_validation.tools.generate_report --raw-dir raw \
--targets-ini tests/targets.ini --out-dir website/public/data
$ cd website && npm install && npm run dev

The Babel Validator in Scala

An initial version of the Babel Validator was written in Scala, but this is no longer being maintained. It is available in the scala-validation/ directory.

Subcommands supported by Babel Validator

The main Babel Validator

diff

$ sbt diff {latest Babel output} {earlier Babel output} --n-cores {number of cores} --output {output directory for Diff files}

Generates a list of differences between two versions of Babel outputs.

About

Programs for validating compendia generated by Babel

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Babel Validator

This repository has several tools for validating the outputs from Babel runs, which are the underlying data used for the Translator Node Normalization and Name Resolver services.

PyTest

The best tests in this repository are Python tests stored in the ./tests folder. This includes both unit tests as well as "Google Sheet"-based tests, which use the shared Babel Validation Google Sheet containing facts that we can use to test a NodeNorm instance. The sheet's ID is deliberately not checked in: copy env.default to .env and fill in BABEL_VALIDATION_SHEET_ID (ask a maintainer for the ID; in GitHub Actions it comes from a repository secret of the same name). env.default documents every variable this repository reads, and what each one turns on.

To run these tests, you need to install uv. You can then use uv to run the tests. The file tests/targets.ini allows you to control which NodeNorm instance is tested. The [DEFAULT] section applies defaults for all the environments. For example, to run all the tests on the dev instance, you can use --target:

$ pytest --target dev
============================= test session starts ==============================
platform darwin -- Python 3.13.3, pytest-8.3.3, pluggy-1.5.0
testing target 'dev': {'nodenormurl': 'https://nodenormalization-sri.renci.org/', 'nameresurl': 'https://name-resolution-sri.renci.org/', 'namereslimit': '20', 'nameresxfailifintop': '5'}
included categories: set()
excluded categories: set()
rootdir: /Users/gaurav/Developer/translator/babel-validation
collected 4338 items [...]

Google Tests have a Category column. To filter based on this column, you can specify a --category on the command line.

$ pytest --target dev --category "Unit Tests" tests/nodenorm/test_nodenorm_from_gsheet.py
==================================================================== test session starts ====================================================================
platform darwin -- Python 3.13.3, pytest-8.3.3, pluggy-1.5.0
testing target 'dev': {'nodenormurl': 'https://nodenormalization-sri.renci.org/', 'nameresurl': 'https://name-resolution-sri.renci.org/', 'namereslimit': '20', 'nameresxfailifintop': '5'}
included categories: {'Unit Tests'}
excluded categories: set()
rootdir: /Users/gaurav/Developer/translator/babel-validation/tests
configfile: pytest.ini
collected 2010 items tests/nodenorm/test_nodenorm_from_gsheet.py sssssxsssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss.ss.x.....sssssssssssssssssssss.ssssss [ 5%]
ssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss...........ssss.....ss.........s...x..sxsssssssssss.ssssss..sssssssssssssssssss [ 12%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 20%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 27%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 34%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 42%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 49%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 57%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 64%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 71%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 79%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 86%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 94%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [100%]
======================================================= 41 passed, 1965 skipped, 4 xfailed in 10.11s ========================================================

GitHub issue tests

Assertions can also be embedded directly in GitHub issue bodies — see src/babel_validation/assertions/README.md for the syntax and the available assertion types. The repositories scanned for them are listed under Repositories in the [DEFAULT] section of tests/targets.ini.

Issue bodies are untrusted input, so the harness caps what one issue may contain — 100 assertions, 1,000 params lists, 1,000 parameters, 1,000 characters per parameter — and rejects YAML anchors, aliases and duplicate keys. An issue over a cap fails loudly rather than running part of itself; split it into several issues. The caps are listed in src/babel_validation/assertions/README.md. --issue resolves only within the configured Repositories, so a run can never be pointed at assertions from somewhere else.

Beware when discussing the syntax in an issue: a complete {{BabelTest|...}} marker is picked up wherever it appears, backticks included, and an unrecognised assertion name fails the run rather than being ignored. Quote a partial marker instead — the pattern needs the closing }} to match.

$ pytest tests/github_issues --target dev # every issue carrying assertions
$ pytest tests/github_issues --target dev --issue 'org/repo#42'# just one (also 'repo#42' or '42')

These tests need a GITHUB_TOKEN, in the environment or in a .env file. Without one they skip rather than fail, so a run can look green having tested nothing. Generate a personal access token; inside a GitHub Action, use the automatic GITHUB_TOKEN instead.

The token is not needed for authentication as such — every repository we scan is public, and both the single-issue and search endpoints answer unauthenticated requests. It is needed for the rate limits:

UnauthenticatedWith a token
Core60 / hour, per IP5,000 / hour
Search10 / minute30 / minute

Discovery is search-bound, not core-bound: two searches per configured repository (one per trigger keyword, plus a request per extra page of results), and then no core request at all, because a search result already carries the issue body and html_url the harness needs. Scanning the five configured repositories currently finds 96 issues for zero core requests.

Core requests are spent re-hydrating issues one at a time, which happens whenever the cached ID list is reused instead of the search being repeated — notably in every pytest-xdist worker after the first. That path costs one request per issue per worker, so an unauthenticated run would exhaust the 60/hour core budget well before finishing.

GET /rate_limit reports what is left without itself counting against the limit (docs). Note that the search window resets every 60 seconds, so its counter is often back at zero by the time you look:

$ curl -s -H "Authorization: Bearer $GITHUB_TOKEN" https://api.github.com/rate_limit

Log Analysis

The Jupyter Notebook in log-analysis/ contains some basic analysis of the logs from NodeNorm (and, someday, NameRes) instances.

The dashboard website

The Astro site in website/ is deployed to https://translatorsri.github.io/babel-validation/. It shows the results of running this test suite against every environment in tests/targets.ini, alongside each environment's /status information (Babel version, database sizes, NameRes latency). Because test expectations are pinned to the environment where a new Babel version lands first, environments are not expected to all be green — the dashboard's purpose is to show which issues are visible in which environment.

The .github/workflows/dashboard.yaml workflow regenerates and deploys it daily (or on manual dispatch): it runs pytest per target with --report-jsonl, turns the raw outcomes and /status responses into report.json and history.jsonl with src.babel_validation.tools.generate_report, and publishes the built site to the gh-pages branch.

The Vue components' client-side logic (URL round-tripping, filtering, pagination, the odd-one-out shading, the drift grouping, and the withholding of blocklist detail) has vitest tests in website/test/, run by npm test and by the Tests workflow.

To work on the site against the data the live dashboard is showing, download the published report.json and history.jsonl instead of generating them (both land in website/public/data/, which is gitignored):

$ cd website && npm install && npm run fetch-data && npm run dev

To regenerate it locally against a couple of environments:

$ uv run pytest tests/nodenorm/test_nodenorm_from_gsheet.py tests/nameres/test_nameres_from_gsheet.py \
--target dev --target prod -n 8 --report-jsonl raw/local.jsonl
$ uv run python -m src.babel_validation.tools.generate_report --raw-dir raw \
--targets-ini tests/targets.ini --out-dir website/public/data
$ cd website && npm install && npm run dev

The Babel Validator in Scala

An initial version of the Babel Validator was written in Scala, but this is no longer being maintained. It is available in the scala-validation/ directory.

Subcommands supported by Babel Validator

The main Babel Validator

diff

$ sbt diff {latest Babel output} {earlier Babel output} --n-cores {number of cores} --output {output directory for Diff files}

Generates a list of differences between two versions of Babel outputs.

About

Programs for validating compendia generated by Babel

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Babel Validator

This repository has several tools for validating the outputs from Babel runs, which are the underlying data used for the Translator Node Normalization and Name Resolver services.

PyTest

The best tests in this repository are Python tests stored in the ./tests folder. This includes both unit tests as well as "Google Sheet"-based tests, which use the shared Babel Validation Google Sheet containing facts that we can use to test a NodeNorm instance. The sheet's ID is deliberately not checked in: copy env.default to .env and fill in BABEL_VALIDATION_SHEET_ID (ask a maintainer for the ID; in GitHub Actions it comes from a repository secret of the same name). env.default documents every variable this repository reads, and what each one turns on.

To run these tests, you need to install uv. You can then use uv to run the tests. The file tests/targets.ini allows you to control which NodeNorm instance is tested. The [DEFAULT] section applies defaults for all the environments. For example, to run all the tests on the dev instance, you can use --target:

$ pytest --target dev
============================= test session starts ==============================
platform darwin -- Python 3.13.3, pytest-8.3.3, pluggy-1.5.0
testing target 'dev': {'nodenormurl': 'https://nodenormalization-sri.renci.org/', 'nameresurl': 'https://name-resolution-sri.renci.org/', 'namereslimit': '20', 'nameresxfailifintop': '5'}
included categories: set()
excluded categories: set()
rootdir: /Users/gaurav/Developer/translator/babel-validation
collected 4338 items [...]

Google Tests have a Category column. To filter based on this column, you can specify a --category on the command line.

$ pytest --target dev --category "Unit Tests" tests/nodenorm/test_nodenorm_from_gsheet.py
==================================================================== test session starts ====================================================================
platform darwin -- Python 3.13.3, pytest-8.3.3, pluggy-1.5.0
testing target 'dev': {'nodenormurl': 'https://nodenormalization-sri.renci.org/', 'nameresurl': 'https://name-resolution-sri.renci.org/', 'namereslimit': '20', 'nameresxfailifintop': '5'}
included categories: {'Unit Tests'}
excluded categories: set()
rootdir: /Users/gaurav/Developer/translator/babel-validation/tests
configfile: pytest.ini
collected 2010 items tests/nodenorm/test_nodenorm_from_gsheet.py sssssxsssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss.ss.x.....sssssssssssssssssssss.ssssss [ 5%]
ssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss...........ssss.....ss.........s...x..sxsssssssssss.ssssss..sssssssssssssssssss [ 12%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 20%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 27%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 34%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 42%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 49%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 57%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 64%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 71%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 79%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 86%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 94%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [100%]
======================================================= 41 passed, 1965 skipped, 4 xfailed in 10.11s ========================================================

GitHub issue tests

Assertions can also be embedded directly in GitHub issue bodies — see src/babel_validation/assertions/README.md for the syntax and the available assertion types. The repositories scanned for them are listed under Repositories in the [DEFAULT] section of tests/targets.ini.

Issue bodies are untrusted input, so the harness caps what one issue may contain — 100 assertions, 1,000 params lists, 1,000 parameters, 1,000 characters per parameter — and rejects YAML anchors, aliases and duplicate keys. An issue over a cap fails loudly rather than running part of itself; split it into several issues. The caps are listed in src/babel_validation/assertions/README.md. --issue resolves only within the configured Repositories, so a run can never be pointed at assertions from somewhere else.

Beware when discussing the syntax in an issue: a complete {{BabelTest|...}} marker is picked up wherever it appears, backticks included, and an unrecognised assertion name fails the run rather than being ignored. Quote a partial marker instead — the pattern needs the closing }} to match.

$ pytest tests/github_issues --target dev # every issue carrying assertions
$ pytest tests/github_issues --target dev --issue 'org/repo#42'# just one (also 'repo#42' or '42')

These tests need a GITHUB_TOKEN, in the environment or in a .env file. Without one they skip rather than fail, so a run can look green having tested nothing. Generate a personal access token; inside a GitHub Action, use the automatic GITHUB_TOKEN instead.

The token is not needed for authentication as such — every repository we scan is public, and both the single-issue and search endpoints answer unauthenticated requests. It is needed for the rate limits:

UnauthenticatedWith a token
Core60 / hour, per IP5,000 / hour
Search10 / minute30 / minute

Discovery is search-bound, not core-bound: two searches per configured repository (one per trigger keyword, plus a request per extra page of results), and then no core request at all, because a search result already carries the issue body and html_url the harness needs. Scanning the five configured repositories currently finds 96 issues for zero core requests.

Core requests are spent re-hydrating issues one at a time, which happens whenever the cached ID list is reused instead of the search being repeated — notably in every pytest-xdist worker after the first. That path costs one request per issue per worker, so an unauthenticated run would exhaust the 60/hour core budget well before finishing.

GET /rate_limit reports what is left without itself counting against the limit (docs). Note that the search window resets every 60 seconds, so its counter is often back at zero by the time you look:

$ curl -s -H "Authorization: Bearer $GITHUB_TOKEN" https://api.github.com/rate_limit

Log Analysis

The Jupyter Notebook in log-analysis/ contains some basic analysis of the logs from NodeNorm (and, someday, NameRes) instances.

The dashboard website

The Astro site in website/ is deployed to https://translatorsri.github.io/babel-validation/. It shows the results of running this test suite against every environment in tests/targets.ini, alongside each environment's /status information (Babel version, database sizes, NameRes latency). Because test expectations are pinned to the environment where a new Babel version lands first, environments are not expected to all be green — the dashboard's purpose is to show which issues are visible in which environment.

The .github/workflows/dashboard.yaml workflow regenerates and deploys it daily (or on manual dispatch): it runs pytest per target with --report-jsonl, turns the raw outcomes and /status responses into report.json and history.jsonl with src.babel_validation.tools.generate_report, and publishes the built site to the gh-pages branch.

The Vue components' client-side logic (URL round-tripping, filtering, pagination, the odd-one-out shading, the drift grouping, and the withholding of blocklist detail) has vitest tests in website/test/, run by npm test and by the Tests workflow.

To work on the site against the data the live dashboard is showing, download the published report.json and history.jsonl instead of generating them (both land in website/public/data/, which is gitignored):

$ cd website && npm install && npm run fetch-data && npm run dev

To regenerate it locally against a couple of environments:

$ uv run pytest tests/nodenorm/test_nodenorm_from_gsheet.py tests/nameres/test_nameres_from_gsheet.py \
--target dev --target prod -n 8 --report-jsonl raw/local.jsonl
$ uv run python -m src.babel_validation.tools.generate_report --raw-dir raw \
--targets-ini tests/targets.ini --out-dir website/public/data
$ cd website && npm install && npm run dev

The Babel Validator in Scala

An initial version of the Babel Validator was written in Scala, but this is no longer being maintained. It is available in the scala-validation/ directory.

Subcommands supported by Babel Validator

The main Babel Validator

diff

$ sbt diff {latest Babel output} {earlier Babel output} --n-cores {number of cores} --output {output directory for Diff files}

Generates a list of differences between two versions of Babel outputs.

About

Programs for validating compendia generated by Babel

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Babel Validator

This repository has several tools for validating the outputs from Babel runs, which are the underlying data used for the Translator Node Normalization and Name Resolver services.

PyTest

The best tests in this repository are Python tests stored in the ./tests folder. This includes both unit tests as well as "Google Sheet"-based tests, which use the shared Babel Validation Google Sheet containing facts that we can use to test a NodeNorm instance. The sheet's ID is deliberately not checked in: copy env.default to .env and fill in BABEL_VALIDATION_SHEET_ID (ask a maintainer for the ID; in GitHub Actions it comes from a repository secret of the same name). env.default documents every variable this repository reads, and what each one turns on.

To run these tests, you need to install uv. You can then use uv to run the tests. The file tests/targets.ini allows you to control which NodeNorm instance is tested. The [DEFAULT] section applies defaults for all the environments. For example, to run all the tests on the dev instance, you can use --target:

$ pytest --target dev
============================= test session starts ==============================
platform darwin -- Python 3.13.3, pytest-8.3.3, pluggy-1.5.0
testing target 'dev': {'nodenormurl': 'https://nodenormalization-sri.renci.org/', 'nameresurl': 'https://name-resolution-sri.renci.org/', 'namereslimit': '20', 'nameresxfailifintop': '5'}
included categories: set()
excluded categories: set()
rootdir: /Users/gaurav/Developer/translator/babel-validation
collected 4338 items [...]

Google Tests have a Category column. To filter based on this column, you can specify a --category on the command line.

$ pytest --target dev --category "Unit Tests" tests/nodenorm/test_nodenorm_from_gsheet.py
==================================================================== test session starts ====================================================================
platform darwin -- Python 3.13.3, pytest-8.3.3, pluggy-1.5.0
testing target 'dev': {'nodenormurl': 'https://nodenormalization-sri.renci.org/', 'nameresurl': 'https://name-resolution-sri.renci.org/', 'namereslimit': '20', 'nameresxfailifintop': '5'}
included categories: {'Unit Tests'}
excluded categories: set()
rootdir: /Users/gaurav/Developer/translator/babel-validation/tests
configfile: pytest.ini
collected 2010 items tests/nodenorm/test_nodenorm_from_gsheet.py sssssxsssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss.ss.x.....sssssssssssssssssssss.ssssss [ 5%]
ssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss...........ssss.....ss.........s...x..sxsssssssssss.ssssss..sssssssssssssssssss [ 12%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 20%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 27%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 34%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 42%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 49%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 57%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 64%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 71%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 79%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 86%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 94%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [100%]
======================================================= 41 passed, 1965 skipped, 4 xfailed in 10.11s ========================================================

GitHub issue tests

Assertions can also be embedded directly in GitHub issue bodies — see src/babel_validation/assertions/README.md for the syntax and the available assertion types. The repositories scanned for them are listed under Repositories in the [DEFAULT] section of tests/targets.ini.

Issue bodies are untrusted input, so the harness caps what one issue may contain — 100 assertions, 1,000 params lists, 1,000 parameters, 1,000 characters per parameter — and rejects YAML anchors, aliases and duplicate keys. An issue over a cap fails loudly rather than running part of itself; split it into several issues. The caps are listed in src/babel_validation/assertions/README.md. --issue resolves only within the configured Repositories, so a run can never be pointed at assertions from somewhere else.

Beware when discussing the syntax in an issue: a complete {{BabelTest|...}} marker is picked up wherever it appears, backticks included, and an unrecognised assertion name fails the run rather than being ignored. Quote a partial marker instead — the pattern needs the closing }} to match.

$ pytest tests/github_issues --target dev # every issue carrying assertions
$ pytest tests/github_issues --target dev --issue 'org/repo#42'# just one (also 'repo#42' or '42')

These tests need a GITHUB_TOKEN, in the environment or in a .env file. Without one they skip rather than fail, so a run can look green having tested nothing. Generate a personal access token; inside a GitHub Action, use the automatic GITHUB_TOKEN instead.

The token is not needed for authentication as such — every repository we scan is public, and both the single-issue and search endpoints answer unauthenticated requests. It is needed for the rate limits:

UnauthenticatedWith a token
Core60 / hour, per IP5,000 / hour
Search10 / minute30 / minute

Discovery is search-bound, not core-bound: two searches per configured repository (one per trigger keyword, plus a request per extra page of results), and then no core request at all, because a search result already carries the issue body and html_url the harness needs. Scanning the five configured repositories currently finds 96 issues for zero core requests.

Core requests are spent re-hydrating issues one at a time, which happens whenever the cached ID list is reused instead of the search being repeated — notably in every pytest-xdist worker after the first. That path costs one request per issue per worker, so an unauthenticated run would exhaust the 60/hour core budget well before finishing.

GET /rate_limit reports what is left without itself counting against the limit (docs). Note that the search window resets every 60 seconds, so its counter is often back at zero by the time you look:

$ curl -s -H "Authorization: Bearer $GITHUB_TOKEN" https://api.github.com/rate_limit

Log Analysis

The Jupyter Notebook in log-analysis/ contains some basic analysis of the logs from NodeNorm (and, someday, NameRes) instances.

The dashboard website

The Astro site in website/ is deployed to https://translatorsri.github.io/babel-validation/. It shows the results of running this test suite against every environment in tests/targets.ini, alongside each environment's /status information (Babel version, database sizes, NameRes latency). Because test expectations are pinned to the environment where a new Babel version lands first, environments are not expected to all be green — the dashboard's purpose is to show which issues are visible in which environment.

The .github/workflows/dashboard.yaml workflow regenerates and deploys it daily (or on manual dispatch): it runs pytest per target with --report-jsonl, turns the raw outcomes and /status responses into report.json and history.jsonl with src.babel_validation.tools.generate_report, and publishes the built site to the gh-pages branch.

The Vue components' client-side logic (URL round-tripping, filtering, pagination, the odd-one-out shading, the drift grouping, and the withholding of blocklist detail) has vitest tests in website/test/, run by npm test and by the Tests workflow.

To work on the site against the data the live dashboard is showing, download the published report.json and history.jsonl instead of generating them (both land in website/public/data/, which is gitignored):

$ cd website && npm install && npm run fetch-data && npm run dev

To regenerate it locally against a couple of environments:

$ uv run pytest tests/nodenorm/test_nodenorm_from_gsheet.py tests/nameres/test_nameres_from_gsheet.py \
--target dev --target prod -n 8 --report-jsonl raw/local.jsonl
$ uv run python -m src.babel_validation.tools.generate_report --raw-dir raw \
--targets-ini tests/targets.ini --out-dir website/public/data
$ cd website && npm install && npm run dev

The Babel Validator in Scala

An initial version of the Babel Validator was written in Scala, but this is no longer being maintained. It is available in the scala-validation/ directory.

Subcommands supported by Babel Validator

The main Babel Validator

diff

$ sbt diff {latest Babel output} {earlier Babel output} --n-cores {number of cores} --output {output directory for Diff files}

Generates a list of differences between two versions of Babel outputs.

About

Programs for validating compendia generated by Babel

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Babel Validator

This repository has several tools for validating the outputs from Babel runs, which are the underlying data used for the Translator Node Normalization and Name Resolver services.

PyTest

The best tests in this repository are Python tests stored in the ./tests folder. This includes both unit tests as well as "Google Sheet"-based tests, which use the shared Babel Validation Google Sheet containing facts that we can use to test a NodeNorm instance. The sheet's ID is deliberately not checked in: copy env.default to .env and fill in BABEL_VALIDATION_SHEET_ID (ask a maintainer for the ID; in GitHub Actions it comes from a repository secret of the same name). env.default documents every variable this repository reads, and what each one turns on.

To run these tests, you need to install uv. You can then use uv to run the tests. The file tests/targets.ini allows you to control which NodeNorm instance is tested. The [DEFAULT] section applies defaults for all the environments. For example, to run all the tests on the dev instance, you can use --target:

$ pytest --target dev
============================= test session starts ==============================
platform darwin -- Python 3.13.3, pytest-8.3.3, pluggy-1.5.0
testing target 'dev': {'nodenormurl': 'https://nodenormalization-sri.renci.org/', 'nameresurl': 'https://name-resolution-sri.renci.org/', 'namereslimit': '20', 'nameresxfailifintop': '5'}
included categories: set()
excluded categories: set()
rootdir: /Users/gaurav/Developer/translator/babel-validation
collected 4338 items [...]

Google Tests have a Category column. To filter based on this column, you can specify a --category on the command line.

$ pytest --target dev --category "Unit Tests" tests/nodenorm/test_nodenorm_from_gsheet.py
==================================================================== test session starts ====================================================================
platform darwin -- Python 3.13.3, pytest-8.3.3, pluggy-1.5.0
testing target 'dev': {'nodenormurl': 'https://nodenormalization-sri.renci.org/', 'nameresurl': 'https://name-resolution-sri.renci.org/', 'namereslimit': '20', 'nameresxfailifintop': '5'}
included categories: {'Unit Tests'}
excluded categories: set()
rootdir: /Users/gaurav/Developer/translator/babel-validation/tests
configfile: pytest.ini
collected 2010 items tests/nodenorm/test_nodenorm_from_gsheet.py sssssxsssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss.ss.x.....sssssssssssssssssssss.ssssss [ 5%]
ssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss...........ssss.....ss.........s...x..sxsssssssssss.ssssss..sssssssssssssssssss [ 12%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 20%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 27%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 34%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 42%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 49%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 57%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 64%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 71%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 79%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 86%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 94%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [100%]
======================================================= 41 passed, 1965 skipped, 4 xfailed in 10.11s ========================================================

GitHub issue tests

Assertions can also be embedded directly in GitHub issue bodies — see src/babel_validation/assertions/README.md for the syntax and the available assertion types. The repositories scanned for them are listed under Repositories in the [DEFAULT] section of tests/targets.ini.

Issue bodies are untrusted input, so the harness caps what one issue may contain — 100 assertions, 1,000 params lists, 1,000 parameters, 1,000 characters per parameter — and rejects YAML anchors, aliases and duplicate keys. An issue over a cap fails loudly rather than running part of itself; split it into several issues. The caps are listed in src/babel_validation/assertions/README.md. --issue resolves only within the configured Repositories, so a run can never be pointed at assertions from somewhere else.

Beware when discussing the syntax in an issue: a complete {{BabelTest|...}} marker is picked up wherever it appears, backticks included, and an unrecognised assertion name fails the run rather than being ignored. Quote a partial marker instead — the pattern needs the closing }} to match.

$ pytest tests/github_issues --target dev # every issue carrying assertions
$ pytest tests/github_issues --target dev --issue 'org/repo#42'# just one (also 'repo#42' or '42')

These tests need a GITHUB_TOKEN, in the environment or in a .env file. Without one they skip rather than fail, so a run can look green having tested nothing. Generate a personal access token; inside a GitHub Action, use the automatic GITHUB_TOKEN instead.

The token is not needed for authentication as such — every repository we scan is public, and both the single-issue and search endpoints answer unauthenticated requests. It is needed for the rate limits:

UnauthenticatedWith a token
Core60 / hour, per IP5,000 / hour
Search10 / minute30 / minute

Discovery is search-bound, not core-bound: two searches per configured repository (one per trigger keyword, plus a request per extra page of results), and then no core request at all, because a search result already carries the issue body and html_url the harness needs. Scanning the five configured repositories currently finds 96 issues for zero core requests.

Core requests are spent re-hydrating issues one at a time, which happens whenever the cached ID list is reused instead of the search being repeated — notably in every pytest-xdist worker after the first. That path costs one request per issue per worker, so an unauthenticated run would exhaust the 60/hour core budget well before finishing.

GET /rate_limit reports what is left without itself counting against the limit (docs). Note that the search window resets every 60 seconds, so its counter is often back at zero by the time you look:

$ curl -s -H "Authorization: Bearer $GITHUB_TOKEN" https://api.github.com/rate_limit

Log Analysis

The Jupyter Notebook in log-analysis/ contains some basic analysis of the logs from NodeNorm (and, someday, NameRes) instances.

The dashboard website

The Astro site in website/ is deployed to https://translatorsri.github.io/babel-validation/. It shows the results of running this test suite against every environment in tests/targets.ini, alongside each environment's /status information (Babel version, database sizes, NameRes latency). Because test expectations are pinned to the environment where a new Babel version lands first, environments are not expected to all be green — the dashboard's purpose is to show which issues are visible in which environment.

The .github/workflows/dashboard.yaml workflow regenerates and deploys it daily (or on manual dispatch): it runs pytest per target with --report-jsonl, turns the raw outcomes and /status responses into report.json and history.jsonl with src.babel_validation.tools.generate_report, and publishes the built site to the gh-pages branch.

The Vue components' client-side logic (URL round-tripping, filtering, pagination, the odd-one-out shading, the drift grouping, and the withholding of blocklist detail) has vitest tests in website/test/, run by npm test and by the Tests workflow.

To work on the site against the data the live dashboard is showing, download the published report.json and history.jsonl instead of generating them (both land in website/public/data/, which is gitignored):

$ cd website && npm install && npm run fetch-data && npm run dev

To regenerate it locally against a couple of environments:

$ uv run pytest tests/nodenorm/test_nodenorm_from_gsheet.py tests/nameres/test_nameres_from_gsheet.py \
--target dev --target prod -n 8 --report-jsonl raw/local.jsonl
$ uv run python -m src.babel_validation.tools.generate_report --raw-dir raw \
--targets-ini tests/targets.ini --out-dir website/public/data
$ cd website && npm install && npm run dev

The Babel Validator in Scala

An initial version of the Babel Validator was written in Scala, but this is no longer being maintained. It is available in the scala-validation/ directory.

Subcommands supported by Babel Validator

The main Babel Validator

diff

$ sbt diff {latest Babel output} {earlier Babel output} --n-cores {number of cores} --output {output directory for Diff files}

Generates a list of differences between two versions of Babel outputs.

About

Programs for validating compendia generated by Babel

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Babel Validator

This repository has several tools for validating the outputs from Babel runs, which are the underlying data used for the Translator Node Normalization and Name Resolver services.

PyTest

The best tests in this repository are Python tests stored in the ./tests folder. This includes both unit tests as well as "Google Sheet"-based tests, which use the shared Babel Validation Google Sheet containing facts that we can use to test a NodeNorm instance. The sheet's ID is deliberately not checked in: copy env.default to .env and fill in BABEL_VALIDATION_SHEET_ID (ask a maintainer for the ID; in GitHub Actions it comes from a repository secret of the same name). env.default documents every variable this repository reads, and what each one turns on.

To run these tests, you need to install uv. You can then use uv to run the tests. The file tests/targets.ini allows you to control which NodeNorm instance is tested. The [DEFAULT] section applies defaults for all the environments. For example, to run all the tests on the dev instance, you can use --target:

$ pytest --target dev
============================= test session starts ==============================
platform darwin -- Python 3.13.3, pytest-8.3.3, pluggy-1.5.0
testing target 'dev': {'nodenormurl': 'https://nodenormalization-sri.renci.org/', 'nameresurl': 'https://name-resolution-sri.renci.org/', 'namereslimit': '20', 'nameresxfailifintop': '5'}
included categories: set()
excluded categories: set()
rootdir: /Users/gaurav/Developer/translator/babel-validation
collected 4338 items [...]

Google Tests have a Category column. To filter based on this column, you can specify a --category on the command line.

$ pytest --target dev --category "Unit Tests" tests/nodenorm/test_nodenorm_from_gsheet.py
==================================================================== test session starts ====================================================================
platform darwin -- Python 3.13.3, pytest-8.3.3, pluggy-1.5.0
testing target 'dev': {'nodenormurl': 'https://nodenormalization-sri.renci.org/', 'nameresurl': 'https://name-resolution-sri.renci.org/', 'namereslimit': '20', 'nameresxfailifintop': '5'}
included categories: {'Unit Tests'}
excluded categories: set()
rootdir: /Users/gaurav/Developer/translator/babel-validation/tests
configfile: pytest.ini
collected 2010 items tests/nodenorm/test_nodenorm_from_gsheet.py sssssxsssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss.ss.x.....sssssssssssssssssssss.ssssss [ 5%]
ssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss...........ssss.....ss.........s...x..sxsssssssssss.ssssss..sssssssssssssssssss [ 12%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 20%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 27%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 34%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 42%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 49%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 57%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 64%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 71%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 79%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 86%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 94%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [100%]
======================================================= 41 passed, 1965 skipped, 4 xfailed in 10.11s ========================================================

GitHub issue tests

Assertions can also be embedded directly in GitHub issue bodies — see src/babel_validation/assertions/README.md for the syntax and the available assertion types. The repositories scanned for them are listed under Repositories in the [DEFAULT] section of tests/targets.ini.

Issue bodies are untrusted input, so the harness caps what one issue may contain — 100 assertions, 1,000 params lists, 1,000 parameters, 1,000 characters per parameter — and rejects YAML anchors, aliases and duplicate keys. An issue over a cap fails loudly rather than running part of itself; split it into several issues. The caps are listed in src/babel_validation/assertions/README.md. --issue resolves only within the configured Repositories, so a run can never be pointed at assertions from somewhere else.

Beware when discussing the syntax in an issue: a complete {{BabelTest|...}} marker is picked up wherever it appears, backticks included, and an unrecognised assertion name fails the run rather than being ignored. Quote a partial marker instead — the pattern needs the closing }} to match.

$ pytest tests/github_issues --target dev # every issue carrying assertions
$ pytest tests/github_issues --target dev --issue 'org/repo#42'# just one (also 'repo#42' or '42')

These tests need a GITHUB_TOKEN, in the environment or in a .env file. Without one they skip rather than fail, so a run can look green having tested nothing. Generate a personal access token; inside a GitHub Action, use the automatic GITHUB_TOKEN instead.

The token is not needed for authentication as such — every repository we scan is public, and both the single-issue and search endpoints answer unauthenticated requests. It is needed for the rate limits:

UnauthenticatedWith a token
Core60 / hour, per IP5,000 / hour
Search10 / minute30 / minute

Discovery is search-bound, not core-bound: two searches per configured repository (one per trigger keyword, plus a request per extra page of results), and then no core request at all, because a search result already carries the issue body and html_url the harness needs. Scanning the five configured repositories currently finds 96 issues for zero core requests.

Core requests are spent re-hydrating issues one at a time, which happens whenever the cached ID list is reused instead of the search being repeated — notably in every pytest-xdist worker after the first. That path costs one request per issue per worker, so an unauthenticated run would exhaust the 60/hour core budget well before finishing.

GET /rate_limit reports what is left without itself counting against the limit (docs). Note that the search window resets every 60 seconds, so its counter is often back at zero by the time you look:

$ curl -s -H "Authorization: Bearer $GITHUB_TOKEN" https://api.github.com/rate_limit

Log Analysis

The Jupyter Notebook in log-analysis/ contains some basic analysis of the logs from NodeNorm (and, someday, NameRes) instances.

The dashboard website

The Astro site in website/ is deployed to https://translatorsri.github.io/babel-validation/. It shows the results of running this test suite against every environment in tests/targets.ini, alongside each environment's /status information (Babel version, database sizes, NameRes latency). Because test expectations are pinned to the environment where a new Babel version lands first, environments are not expected to all be green — the dashboard's purpose is to show which issues are visible in which environment.

The .github/workflows/dashboard.yaml workflow regenerates and deploys it daily (or on manual dispatch): it runs pytest per target with --report-jsonl, turns the raw outcomes and /status responses into report.json and history.jsonl with src.babel_validation.tools.generate_report, and publishes the built site to the gh-pages branch.

The Vue components' client-side logic (URL round-tripping, filtering, pagination, the odd-one-out shading, the drift grouping, and the withholding of blocklist detail) has vitest tests in website/test/, run by npm test and by the Tests workflow.

To work on the site against the data the live dashboard is showing, download the published report.json and history.jsonl instead of generating them (both land in website/public/data/, which is gitignored):

$ cd website && npm install && npm run fetch-data && npm run dev

To regenerate it locally against a couple of environments:

$ uv run pytest tests/nodenorm/test_nodenorm_from_gsheet.py tests/nameres/test_nameres_from_gsheet.py \
--target dev --target prod -n 8 --report-jsonl raw/local.jsonl
$ uv run python -m src.babel_validation.tools.generate_report --raw-dir raw \
--targets-ini tests/targets.ini --out-dir website/public/data
$ cd website && npm install && npm run dev

The Babel Validator in Scala

An initial version of the Babel Validator was written in Scala, but this is no longer being maintained. It is available in the scala-validation/ directory.

Subcommands supported by Babel Validator

The main Babel Validator

diff

$ sbt diff {latest Babel output} {earlier Babel output} --n-cores {number of cores} --output {output directory for Diff files}

Generates a list of differences between two versions of Babel outputs.

About

Programs for validating compendia generated by Babel

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Babel Validator

This repository has several tools for validating the outputs from Babel runs, which are the underlying data used for the Translator Node Normalization and Name Resolver services.

PyTest

The best tests in this repository are Python tests stored in the ./tests folder. This includes both unit tests as well as "Google Sheet"-based tests, which use the shared Babel Validation Google Sheet containing facts that we can use to test a NodeNorm instance. The sheet's ID is deliberately not checked in: copy env.default to .env and fill in BABEL_VALIDATION_SHEET_ID (ask a maintainer for the ID; in GitHub Actions it comes from a repository secret of the same name). env.default documents every variable this repository reads, and what each one turns on.

To run these tests, you need to install uv. You can then use uv to run the tests. The file tests/targets.ini allows you to control which NodeNorm instance is tested. The [DEFAULT] section applies defaults for all the environments. For example, to run all the tests on the dev instance, you can use --target:

$ pytest --target dev
============================= test session starts ==============================
platform darwin -- Python 3.13.3, pytest-8.3.3, pluggy-1.5.0
testing target 'dev': {'nodenormurl': 'https://nodenormalization-sri.renci.org/', 'nameresurl': 'https://name-resolution-sri.renci.org/', 'namereslimit': '20', 'nameresxfailifintop': '5'}
included categories: set()
excluded categories: set()
rootdir: /Users/gaurav/Developer/translator/babel-validation
collected 4338 items [...]

Google Tests have a Category column. To filter based on this column, you can specify a --category on the command line.

$ pytest --target dev --category "Unit Tests" tests/nodenorm/test_nodenorm_from_gsheet.py
==================================================================== test session starts ====================================================================
platform darwin -- Python 3.13.3, pytest-8.3.3, pluggy-1.5.0
testing target 'dev': {'nodenormurl': 'https://nodenormalization-sri.renci.org/', 'nameresurl': 'https://name-resolution-sri.renci.org/', 'namereslimit': '20', 'nameresxfailifintop': '5'}
included categories: {'Unit Tests'}
excluded categories: set()
rootdir: /Users/gaurav/Developer/translator/babel-validation/tests
configfile: pytest.ini
collected 2010 items tests/nodenorm/test_nodenorm_from_gsheet.py sssssxsssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss.ss.x.....sssssssssssssssssssss.ssssss [ 5%]
ssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss...........ssss.....ss.........s...x..sxsssssssssss.ssssss..sssssssssssssssssss [ 12%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 20%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 27%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 34%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 42%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 49%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 57%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 64%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 71%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 79%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 86%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 94%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [100%]
======================================================= 41 passed, 1965 skipped, 4 xfailed in 10.11s ========================================================

GitHub issue tests

Assertions can also be embedded directly in GitHub issue bodies — see src/babel_validation/assertions/README.md for the syntax and the available assertion types. The repositories scanned for them are listed under Repositories in the [DEFAULT] section of tests/targets.ini.

Issue bodies are untrusted input, so the harness caps what one issue may contain — 100 assertions, 1,000 params lists, 1,000 parameters, 1,000 characters per parameter — and rejects YAML anchors, aliases and duplicate keys. An issue over a cap fails loudly rather than running part of itself; split it into several issues. The caps are listed in src/babel_validation/assertions/README.md. --issue resolves only within the configured Repositories, so a run can never be pointed at assertions from somewhere else.

Beware when discussing the syntax in an issue: a complete {{BabelTest|...}} marker is picked up wherever it appears, backticks included, and an unrecognised assertion name fails the run rather than being ignored. Quote a partial marker instead — the pattern needs the closing }} to match.

$ pytest tests/github_issues --target dev # every issue carrying assertions
$ pytest tests/github_issues --target dev --issue 'org/repo#42'# just one (also 'repo#42' or '42')

These tests need a GITHUB_TOKEN, in the environment or in a .env file. Without one they skip rather than fail, so a run can look green having tested nothing. Generate a personal access token; inside a GitHub Action, use the automatic GITHUB_TOKEN instead.

The token is not needed for authentication as such — every repository we scan is public, and both the single-issue and search endpoints answer unauthenticated requests. It is needed for the rate limits:

UnauthenticatedWith a token
Core60 / hour, per IP5,000 / hour
Search10 / minute30 / minute

Discovery is search-bound, not core-bound: two searches per configured repository (one per trigger keyword, plus a request per extra page of results), and then no core request at all, because a search result already carries the issue body and html_url the harness needs. Scanning the five configured repositories currently finds 96 issues for zero core requests.

Core requests are spent re-hydrating issues one at a time, which happens whenever the cached ID list is reused instead of the search being repeated — notably in every pytest-xdist worker after the first. That path costs one request per issue per worker, so an unauthenticated run would exhaust the 60/hour core budget well before finishing.

GET /rate_limit reports what is left without itself counting against the limit (docs). Note that the search window resets every 60 seconds, so its counter is often back at zero by the time you look:

$ curl -s -H "Authorization: Bearer $GITHUB_TOKEN" https://api.github.com/rate_limit

Log Analysis

The Jupyter Notebook in log-analysis/ contains some basic analysis of the logs from NodeNorm (and, someday, NameRes) instances.

The dashboard website

The Astro site in website/ is deployed to https://translatorsri.github.io/babel-validation/. It shows the results of running this test suite against every environment in tests/targets.ini, alongside each environment's /status information (Babel version, database sizes, NameRes latency). Because test expectations are pinned to the environment where a new Babel version lands first, environments are not expected to all be green — the dashboard's purpose is to show which issues are visible in which environment.

The .github/workflows/dashboard.yaml workflow regenerates and deploys it daily (or on manual dispatch): it runs pytest per target with --report-jsonl, turns the raw outcomes and /status responses into report.json and history.jsonl with src.babel_validation.tools.generate_report, and publishes the built site to the gh-pages branch.

The Vue components' client-side logic (URL round-tripping, filtering, pagination, the odd-one-out shading, the drift grouping, and the withholding of blocklist detail) has vitest tests in website/test/, run by npm test and by the Tests workflow.

To work on the site against the data the live dashboard is showing, download the published report.json and history.jsonl instead of generating them (both land in website/public/data/, which is gitignored):

$ cd website && npm install && npm run fetch-data && npm run dev

To regenerate it locally against a couple of environments:

$ uv run pytest tests/nodenorm/test_nodenorm_from_gsheet.py tests/nameres/test_nameres_from_gsheet.py \
--target dev --target prod -n 8 --report-jsonl raw/local.jsonl
$ uv run python -m src.babel_validation.tools.generate_report --raw-dir raw \
--targets-ini tests/targets.ini --out-dir website/public/data
$ cd website && npm install && npm run dev

The Babel Validator in Scala

An initial version of the Babel Validator was written in Scala, but this is no longer being maintained. It is available in the scala-validation/ directory.

Subcommands supported by Babel Validator

The main Babel Validator

diff

$ sbt diff {latest Babel output} {earlier Babel output} --n-cores {number of cores} --output {output directory for Diff files}

Generates a list of differences between two versions of Babel outputs.

About

Programs for validating compendia generated by Babel

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Babel Validator

This repository has several tools for validating the outputs from Babel runs, which are the underlying data used for the Translator Node Normalization and Name Resolver services.

PyTest

The best tests in this repository are Python tests stored in the ./tests folder. This includes both unit tests as well as "Google Sheet"-based tests, which use the shared Babel Validation Google Sheet containing facts that we can use to test a NodeNorm instance. The sheet's ID is deliberately not checked in: copy env.default to .env and fill in BABEL_VALIDATION_SHEET_ID (ask a maintainer for the ID; in GitHub Actions it comes from a repository secret of the same name). env.default documents every variable this repository reads, and what each one turns on.

To run these tests, you need to install uv. You can then use uv to run the tests. The file tests/targets.ini allows you to control which NodeNorm instance is tested. The [DEFAULT] section applies defaults for all the environments. For example, to run all the tests on the dev instance, you can use --target:

$ pytest --target dev
============================= test session starts ==============================
platform darwin -- Python 3.13.3, pytest-8.3.3, pluggy-1.5.0
testing target 'dev': {'nodenormurl': 'https://nodenormalization-sri.renci.org/', 'nameresurl': 'https://name-resolution-sri.renci.org/', 'namereslimit': '20', 'nameresxfailifintop': '5'}
included categories: set()
excluded categories: set()
rootdir: /Users/gaurav/Developer/translator/babel-validation
collected 4338 items [...]

Google Tests have a Category column. To filter based on this column, you can specify a --category on the command line.

$ pytest --target dev --category "Unit Tests" tests/nodenorm/test_nodenorm_from_gsheet.py
==================================================================== test session starts ====================================================================
platform darwin -- Python 3.13.3, pytest-8.3.3, pluggy-1.5.0
testing target 'dev': {'nodenormurl': 'https://nodenormalization-sri.renci.org/', 'nameresurl': 'https://name-resolution-sri.renci.org/', 'namereslimit': '20', 'nameresxfailifintop': '5'}
included categories: {'Unit Tests'}
excluded categories: set()
rootdir: /Users/gaurav/Developer/translator/babel-validation/tests
configfile: pytest.ini
collected 2010 items tests/nodenorm/test_nodenorm_from_gsheet.py sssssxsssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss.ss.x.....sssssssssssssssssssss.ssssss [ 5%]
ssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss...........ssss.....ss.........s...x..sxsssssssssss.ssssss..sssssssssssssssssss [ 12%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 20%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 27%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 34%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 42%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 49%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 57%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 64%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 71%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 79%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 86%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [ 94%]
sssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssssss [100%]
======================================================= 41 passed, 1965 skipped, 4 xfailed in 10.11s ========================================================

GitHub issue tests

Assertions can also be embedded directly in GitHub issue bodies — see src/babel_validation/assertions/README.md for the syntax and the available assertion types. The repositories scanned for them are listed under Repositories in the [DEFAULT] section of tests/targets.ini.

Issue bodies are untrusted input, so the harness caps what one issue may contain — 100 assertions, 1,000 params lists, 1,000 parameters, 1,000 characters per parameter — and rejects YAML anchors, aliases and duplicate keys. An issue over a cap fails loudly rather than running part of itself; split it into several issues. The caps are listed in src/babel_validation/assertions/README.md. --issue resolves only within the configured Repositories, so a run can never be pointed at assertions from somewhere else.

Beware when discussing the syntax in an issue: a complete {{BabelTest|...}} marker is picked up wherever it appears, backticks included, and an unrecognised assertion name fails the run rather than being ignored. Quote a partial marker instead — the pattern needs the closing }} to match.

$ pytest tests/github_issues --target dev # every issue carrying assertions
$ pytest tests/github_issues --target dev --issue 'org/repo#42'# just one (also 'repo#42' or '42')

These tests need a GITHUB_TOKEN, in the environment or in a .env file. Without one they skip rather than fail, so a run can look green having tested nothing. Generate a personal access token; inside a GitHub Action, use the automatic GITHUB_TOKEN instead.

The token is not needed for authentication as such — every repository we scan is public, and both the single-issue and search endpoints answer unauthenticated requests. It is needed for the rate limits:

UnauthenticatedWith a token
Core60 / hour, per IP5,000 / hour
Search10 / minute30 / minute

Discovery is search-bound, not core-bound: two searches per configured repository (one per trigger keyword, plus a request per extra page of results), and then no core request at all, because a search result already carries the issue body and html_url the harness needs. Scanning the five configured repositories currently finds 96 issues for zero core requests.

Core requests are spent re-hydrating issues one at a time, which happens whenever the cached ID list is reused instead of the search being repeated — notably in every pytest-xdist worker after the first. That path costs one request per issue per worker, so an unauthenticated run would exhaust the 60/hour core budget well before finishing.

GET /rate_limit reports what is left without itself counting against the limit (docs). Note that the search window resets every 60 seconds, so its counter is often back at zero by the time you look:

$ curl -s -H "Authorization: Bearer $GITHUB_TOKEN" https://api.github.com/rate_limit

Log Analysis

The Jupyter Notebook in log-analysis/ contains some basic analysis of the logs from NodeNorm (and, someday, NameRes) instances.

The dashboard website

The Astro site in website/ is deployed to https://translatorsri.github.io/babel-validation/. It shows the results of running this test suite against every environment in tests/targets.ini, alongside each environment's /status information (Babel version, database sizes, NameRes latency). Because test expectations are pinned to the environment where a new Babel version lands first, environments are not expected to all be green — the dashboard's purpose is to show which issues are visible in which environment.

The .github/workflows/dashboard.yaml workflow regenerates and deploys it daily (or on manual dispatch): it runs pytest per target with --report-jsonl, turns the raw outcomes and /status responses into report.json and history.jsonl with src.babel_validation.tools.generate_report, and publishes the built site to the gh-pages branch.

The Vue components' client-side logic (URL round-tripping, filtering, pagination, the odd-one-out shading, the drift grouping, and the withholding of blocklist detail) has vitest tests in website/test/, run by npm test and by the Tests workflow.

To work on the site against the data the live dashboard is showing, download the published report.json and history.jsonl instead of generating them (both land in website/public/data/, which is gitignored):

$ cd website && npm install && npm run fetch-data && npm run dev

To regenerate it locally against a couple of environments:

$ uv run pytest tests/nodenorm/test_nodenorm_from_gsheet.py tests/nameres/test_nameres_from_gsheet.py \
--target dev --target prod -n 8 --report-jsonl raw/local.jsonl
$ uv run python -m src.babel_validation.tools.generate_report --raw-dir raw \
--targets-ini tests/targets.ini --out-dir website/public/data
$ cd website && npm install && npm run dev

The Babel Validator in Scala

An initial version of the Babel Validator was written in Scala, but this is no longer being maintained. It is available in the scala-validation/ directory.

Subcommands supported by Babel Validator

The main Babel Validator

diff

$ sbt diff {latest Babel output} {earlier Babel output} --n-cores {number of cores} --output {output directory for Diff files}

Generates a list of differences between two versions of Babel outputs.

About

Programs for validating compendia generated by Babel

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages