Skip to content

benchmarks: create the .net cache gen_networks.py exists to fill (#492) - #494

Merged
wshlavacek merged 1 commit into
mainfrom
bench-gen-networks-cache-dir-492
Aug 27, 2026
Merged

benchmarks: create the .net cache gen_networks.py exists to fill (#492)#494
wshlavacek merged 1 commit into
mainfrom
bench-gen-networks-cache-dir-492

Conversation

@wshlavacek

Copy link
Copy Markdown
Collaborator

Closes#492.

Taken as asked — this one is on a recovery path, so the mkdir and the errno
check land together.

The fix

ensure_cache_dir(NETS) runs before the header, so the cache: ... line the
pass prints names a directory that exists rather than one every write is about
to fail against.

$ rm -rf benchmarks/suites/ode_fullnet/nets
$ python gen_networks.py --models Lang2024 --workers 1
cache: .../ode_fullnet/nets # now true when printed
[1/1] ok 0.41s Lang_2024.bngl N=73 rxn=332

The errno check, paired as you asked

Once the directory is created first, the old failure is unreachable on a normal
run. What is left is the case mkdir cannot cover — the directory removed
while a multi-hour sweep is in flight — and that is what
describe_write_failure() is for:

.../ode_fullnet/nets (the .net cache directory) no longer exists -- it is
created at the start of the pass, so it was removed mid-run; the network
itself generated fine

Three details, matching the three things you said conspired to make it hard to
read:

  1. It names the directory, not the file. And the row's status is now
    cache_write_failed rather than a bare error, so the console tag alone
    answers netgen-or-write without needing the detail at all. Nothing downstream
    enumerates statuses — run_timing.py only compares against "ok".

  2. The path comes first in the message, because the console truncates a
    row's detail at 70 characters. That truncation is why the path was cut off in
    your reproduction.

  3. One correction here, and it may save you time later: the manifest did
    carry the message, under detail rather than error.
    There is no error
    key in a manifest row, so .get("error") is None for every row including
    the healthy ones — the 585 ok rows have no error either. detail holds
    up to 400 characters and had the full path all along. Checking a
    netgen_timeout row in the current manifest:

    {"status": "netgen_timeout", "detail": "BNG2.pl netgen timed out after 180.0s"}

    So the report could have answered it; the key name was the obstacle, not the
    record. I have not renamed or aliased the key — that would rewrite the shape
    of a tracked 592-row manifest for a naming preference, which is your call to
    make, not a fix to slip into this PR.

The test

python/tests/test_gen_networks_cache_dir.py — 6 cases, no BNG2.pl needed,
which is the point: it drives main() with a model filter matching nothing, so
the run reaches its "nothing to do" exit without generating anything. That is
the ordering under test — #492 was not that the cache was never created, it was
that it was not created first.

Negative-verified precisely: removing only the ensure_cache_dir call (leaving
the helper and everything else in place) fails that one test on its assertion,
with the run still printing cache: <path> for a path that does not exist. The
other five cover idempotence on the 585-network resume, and the three
classification branches — vanished directory, a real ENOENT with the directory
present, and a non-ENOENT PermissionError (a full disk or read-only mount is
not a missing directory).

For the run machine

This is safe to take before the measurement phase, on your reasoning: it changes
what happens when the cache is missing, and nothing about what is generated
once it exists. The generated .net bytes are untouched — the only new code
paths are a mkdir and an error-message branch that a successful run never
reaches.

The mkdir -p workaround stays valid, so nothing forces the timing of this.

Phase 1 of the full-network ODE benchmark writes every generated network
into NETS = HERE / "nets" and never created that directory. On a machine
where it did not exist -- a fresh checkout, or one where the cache had
been cleaned to reclaim disk -- every model failed.
Expensively, and that is the part worth naming: the generation work
happens first and only the write fails, so six models spent 28 s to
produce nothing, and the full corpus would spend hours the same way. A
missing cache is also exactly the state in which this script is the
documented recovery path, so the one script whose job is to *build* the
cache was the one script that could not create it.
The pass now calls ensure_cache_dir(NETS) before the header, so the
"cache: ..." line it prints names a directory that exists rather than one
every write is about to fail against.
That makes the old failure unreachable on a normal run, which leaves the
case it cannot cover: the directory removed *while* a multi-hour sweep is
in flight. describe_write_failure() handles that one and says so in those
words. It matters because the old message named the wrong thing --
FileNotFoundError on a .net path reads as "this model could not be
generated", which is what a genuine BNG2.pl netgen failure looks like too
-- and because the console truncates a row's detail, so the path the
reader needed was cut off. The new detail leads with the directory, and
the row's status is cache_write_failed rather than a bare error, so the
console tag alone answers netgen-or-write without needing the detail at
all. Nothing downstream enumerates statuses; run_timing.py only compares
against "ok".
test_gen_networks_cache_dir.py drives main() with a model filter that
matches nothing, so it reaches the "nothing to do" exit without BNG2.pl
and without generating anything -- which is the ordering under test. #492
was not that the cache was never created, it was that it was not created
first. Removing only the ensure_cache_dir call fails that test on its
assertion, with the run still printing "cache: <path>" for a path that
does not exist.
One correction to the issue's third point: the manifest did carry the
message, under "detail" rather than "error" -- `.get("error")` is None
because there is no such key. The 400-char detail held the full path all
along; it was the 70-char console truncation that hid it.
@wshlavacek
wshlavacek merged commit f8be826 into mainAug 27, 2026
4 checks passed
@wshlavacek
wshlavacek deleted the bench-gen-networks-cache-dir-492 branch August 27, 2026 18:11
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

gen_networks.py does not create its own nets/ cache directory, so every model fails on a machine without it

1 participant

@wshlavacek
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
benchmarks: create the .net cache gen_networks.py exists to fill (#492) by wshlavacek · Pull Request #494 · lanl/bngsim · GitHub
Skip to content

benchmarks: create the .net cache gen_networks.py exists to fill (#492) - #494

Merged
wshlavacek merged 1 commit into
mainfrom
bench-gen-networks-cache-dir-492
Aug 27, 2026
Merged

benchmarks: create the .net cache gen_networks.py exists to fill (#492)#494
wshlavacek merged 1 commit into
mainfrom
bench-gen-networks-cache-dir-492

Conversation

@wshlavacek

Copy link
Copy Markdown
Collaborator

Closes#492.

Taken as asked — this one is on a recovery path, so the mkdir and the errno
check land together.

The fix

ensure_cache_dir(NETS) runs before the header, so the cache: ... line the
pass prints names a directory that exists rather than one every write is about
to fail against.

$ rm -rf benchmarks/suites/ode_fullnet/nets
$ python gen_networks.py --models Lang2024 --workers 1
cache: .../ode_fullnet/nets # now true when printed
[1/1] ok 0.41s Lang_2024.bngl N=73 rxn=332

The errno check, paired as you asked

Once the directory is created first, the old failure is unreachable on a normal
run. What is left is the case mkdir cannot cover — the directory removed
while a multi-hour sweep is in flight — and that is what
describe_write_failure() is for:

.../ode_fullnet/nets (the .net cache directory) no longer exists -- it is
created at the start of the pass, so it was removed mid-run; the network
itself generated fine

Three details, matching the three things you said conspired to make it hard to
read:

  1. It names the directory, not the file. And the row's status is now
    cache_write_failed rather than a bare error, so the console tag alone
    answers netgen-or-write without needing the detail at all. Nothing downstream
    enumerates statuses — run_timing.py only compares against "ok".

  2. The path comes first in the message, because the console truncates a
    row's detail at 70 characters. That truncation is why the path was cut off in
    your reproduction.

  3. One correction here, and it may save you time later: the manifest did
    carry the message, under detail rather than error.
    There is no error
    key in a manifest row, so .get("error") is None for every row including
    the healthy ones — the 585 ok rows have no error either. detail holds
    up to 400 characters and had the full path all along. Checking a
    netgen_timeout row in the current manifest:

    {"status": "netgen_timeout", "detail": "BNG2.pl netgen timed out after 180.0s"}

    So the report could have answered it; the key name was the obstacle, not the
    record. I have not renamed or aliased the key — that would rewrite the shape
    of a tracked 592-row manifest for a naming preference, which is your call to
    make, not a fix to slip into this PR.

The test

python/tests/test_gen_networks_cache_dir.py — 6 cases, no BNG2.pl needed,
which is the point: it drives main() with a model filter matching nothing, so
the run reaches its "nothing to do" exit without generating anything. That is
the ordering under test — #492 was not that the cache was never created, it was
that it was not created first.

Negative-verified precisely: removing only the ensure_cache_dir call (leaving
the helper and everything else in place) fails that one test on its assertion,
with the run still printing cache: <path> for a path that does not exist. The
other five cover idempotence on the 585-network resume, and the three
classification branches — vanished directory, a real ENOENT with the directory
present, and a non-ENOENT PermissionError (a full disk or read-only mount is
not a missing directory).

For the run machine

This is safe to take before the measurement phase, on your reasoning: it changes
what happens when the cache is missing, and nothing about what is generated
once it exists. The generated .net bytes are untouched — the only new code
paths are a mkdir and an error-message branch that a successful run never
reaches.

The mkdir -p workaround stays valid, so nothing forces the timing of this.

Phase 1 of the full-network ODE benchmark writes every generated network
into NETS = HERE / "nets" and never created that directory. On a machine
where it did not exist -- a fresh checkout, or one where the cache had
been cleaned to reclaim disk -- every model failed.
Expensively, and that is the part worth naming: the generation work
happens first and only the write fails, so six models spent 28 s to
produce nothing, and the full corpus would spend hours the same way. A
missing cache is also exactly the state in which this script is the
documented recovery path, so the one script whose job is to *build* the
cache was the one script that could not create it.
The pass now calls ensure_cache_dir(NETS) before the header, so the
"cache: ..." line it prints names a directory that exists rather than one
every write is about to fail against.
That makes the old failure unreachable on a normal run, which leaves the
case it cannot cover: the directory removed *while* a multi-hour sweep is
in flight. describe_write_failure() handles that one and says so in those
words. It matters because the old message named the wrong thing --
FileNotFoundError on a .net path reads as "this model could not be
generated", which is what a genuine BNG2.pl netgen failure looks like too
-- and because the console truncates a row's detail, so the path the
reader needed was cut off. The new detail leads with the directory, and
the row's status is cache_write_failed rather than a bare error, so the
console tag alone answers netgen-or-write without needing the detail at
all. Nothing downstream enumerates statuses; run_timing.py only compares
against "ok".
test_gen_networks_cache_dir.py drives main() with a model filter that
matches nothing, so it reaches the "nothing to do" exit without BNG2.pl
and without generating anything -- which is the ordering under test. #492
was not that the cache was never created, it was that it was not created
first. Removing only the ensure_cache_dir call fails that test on its
assertion, with the run still printing "cache: <path>" for a path that
does not exist.
One correction to the issue's third point: the manifest did carry the
message, under "detail" rather than "error" -- `.get("error")` is None
because there is no such key. The 400-char detail held the full path all
along; it was the 70-char console truncation that hid it.
@wshlavacek
wshlavacek merged commit f8be826 into mainAug 27, 2026
4 checks passed
@wshlavacek
wshlavacek deleted the bench-gen-networks-cache-dir-492 branch August 27, 2026 18:11
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

gen_networks.py does not create its own nets/ cache directory, so every model fails on a machine without it

1 participant

@wshlavacek
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' benchmarks: create the .net cache gen_networks.py exists to fill (#492) by wshlavacek · Pull Request #494 · lanl/bngsim · GitHub
Skip to content

benchmarks: create the .net cache gen_networks.py exists to fill (#492) - #494

Merged
wshlavacek merged 1 commit into
mainfrom
bench-gen-networks-cache-dir-492
Aug 27, 2026
Merged

benchmarks: create the .net cache gen_networks.py exists to fill (#492)#494
wshlavacek merged 1 commit into
mainfrom
bench-gen-networks-cache-dir-492

Conversation

@wshlavacek

Copy link
Copy Markdown
Collaborator

Closes#492.

Taken as asked — this one is on a recovery path, so the mkdir and the errno
check land together.

The fix

ensure_cache_dir(NETS) runs before the header, so the cache: ... line the
pass prints names a directory that exists rather than one every write is about
to fail against.

$ rm -rf benchmarks/suites/ode_fullnet/nets
$ python gen_networks.py --models Lang2024 --workers 1
cache: .../ode_fullnet/nets # now true when printed
[1/1] ok 0.41s Lang_2024.bngl N=73 rxn=332

The errno check, paired as you asked

Once the directory is created first, the old failure is unreachable on a normal
run. What is left is the case mkdir cannot cover — the directory removed
while a multi-hour sweep is in flight — and that is what
describe_write_failure() is for:

.../ode_fullnet/nets (the .net cache directory) no longer exists -- it is
created at the start of the pass, so it was removed mid-run; the network
itself generated fine

Three details, matching the three things you said conspired to make it hard to
read:

  1. It names the directory, not the file. And the row's status is now
    cache_write_failed rather than a bare error, so the console tag alone
    answers netgen-or-write without needing the detail at all. Nothing downstream
    enumerates statuses — run_timing.py only compares against "ok".

  2. The path comes first in the message, because the console truncates a
    row's detail at 70 characters. That truncation is why the path was cut off in
    your reproduction.

  3. One correction here, and it may save you time later: the manifest did
    carry the message, under detail rather than error.
    There is no error
    key in a manifest row, so .get("error") is None for every row including
    the healthy ones — the 585 ok rows have no error either. detail holds
    up to 400 characters and had the full path all along. Checking a
    netgen_timeout row in the current manifest:

    {"status": "netgen_timeout", "detail": "BNG2.pl netgen timed out after 180.0s"}

    So the report could have answered it; the key name was the obstacle, not the
    record. I have not renamed or aliased the key — that would rewrite the shape
    of a tracked 592-row manifest for a naming preference, which is your call to
    make, not a fix to slip into this PR.

The test

python/tests/test_gen_networks_cache_dir.py — 6 cases, no BNG2.pl needed,
which is the point: it drives main() with a model filter matching nothing, so
the run reaches its "nothing to do" exit without generating anything. That is
the ordering under test — #492 was not that the cache was never created, it was
that it was not created first.

Negative-verified precisely: removing only the ensure_cache_dir call (leaving
the helper and everything else in place) fails that one test on its assertion,
with the run still printing cache: <path> for a path that does not exist. The
other five cover idempotence on the 585-network resume, and the three
classification branches — vanished directory, a real ENOENT with the directory
present, and a non-ENOENT PermissionError (a full disk or read-only mount is
not a missing directory).

For the run machine

This is safe to take before the measurement phase, on your reasoning: it changes
what happens when the cache is missing, and nothing about what is generated
once it exists. The generated .net bytes are untouched — the only new code
paths are a mkdir and an error-message branch that a successful run never
reaches.

The mkdir -p workaround stays valid, so nothing forces the timing of this.

Phase 1 of the full-network ODE benchmark writes every generated network
into NETS = HERE / "nets" and never created that directory. On a machine
where it did not exist -- a fresh checkout, or one where the cache had
been cleaned to reclaim disk -- every model failed.
Expensively, and that is the part worth naming: the generation work
happens first and only the write fails, so six models spent 28 s to
produce nothing, and the full corpus would spend hours the same way. A
missing cache is also exactly the state in which this script is the
documented recovery path, so the one script whose job is to *build* the
cache was the one script that could not create it.
The pass now calls ensure_cache_dir(NETS) before the header, so the
"cache: ..." line it prints names a directory that exists rather than one
every write is about to fail against.
That makes the old failure unreachable on a normal run, which leaves the
case it cannot cover: the directory removed *while* a multi-hour sweep is
in flight. describe_write_failure() handles that one and says so in those
words. It matters because the old message named the wrong thing --
FileNotFoundError on a .net path reads as "this model could not be
generated", which is what a genuine BNG2.pl netgen failure looks like too
-- and because the console truncates a row's detail, so the path the
reader needed was cut off. The new detail leads with the directory, and
the row's status is cache_write_failed rather than a bare error, so the
console tag alone answers netgen-or-write without needing the detail at
all. Nothing downstream enumerates statuses; run_timing.py only compares
against "ok".
test_gen_networks_cache_dir.py drives main() with a model filter that
matches nothing, so it reaches the "nothing to do" exit without BNG2.pl
and without generating anything -- which is the ordering under test. #492
was not that the cache was never created, it was that it was not created
first. Removing only the ensure_cache_dir call fails that test on its
assertion, with the run still printing "cache: <path>" for a path that
does not exist.
One correction to the issue's third point: the manifest did carry the
message, under "detail" rather than "error" -- `.get("error")` is None
because there is no such key. The 400-char detail held the full path all
along; it was the 70-char console truncation that hid it.
@wshlavacek
wshlavacek merged commit f8be826 into mainAug 27, 2026
4 checks passed
@wshlavacek
wshlavacek deleted the bench-gen-networks-cache-dir-492 branch August 27, 2026 18:11
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

gen_networks.py does not create its own nets/ cache directory, so every model fails on a machine without it

1 participant

@wshlavacek
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' benchmarks: create the .net cache gen_networks.py exists to fill (#492) by wshlavacek · Pull Request #494 · lanl/bngsim · GitHub
Skip to content

benchmarks: create the .net cache gen_networks.py exists to fill (#492) - #494

Merged
wshlavacek merged 1 commit into
mainfrom
bench-gen-networks-cache-dir-492
Aug 27, 2026
Merged

benchmarks: create the .net cache gen_networks.py exists to fill (#492)#494
wshlavacek merged 1 commit into
mainfrom
bench-gen-networks-cache-dir-492

Conversation

@wshlavacek

Copy link
Copy Markdown
Collaborator

Closes#492.

Taken as asked — this one is on a recovery path, so the mkdir and the errno
check land together.

The fix

ensure_cache_dir(NETS) runs before the header, so the cache: ... line the
pass prints names a directory that exists rather than one every write is about
to fail against.

$ rm -rf benchmarks/suites/ode_fullnet/nets
$ python gen_networks.py --models Lang2024 --workers 1
cache: .../ode_fullnet/nets # now true when printed
[1/1] ok 0.41s Lang_2024.bngl N=73 rxn=332

The errno check, paired as you asked

Once the directory is created first, the old failure is unreachable on a normal
run. What is left is the case mkdir cannot cover — the directory removed
while a multi-hour sweep is in flight — and that is what
describe_write_failure() is for:

.../ode_fullnet/nets (the .net cache directory) no longer exists -- it is
created at the start of the pass, so it was removed mid-run; the network
itself generated fine

Three details, matching the three things you said conspired to make it hard to
read:

  1. It names the directory, not the file. And the row's status is now
    cache_write_failed rather than a bare error, so the console tag alone
    answers netgen-or-write without needing the detail at all. Nothing downstream
    enumerates statuses — run_timing.py only compares against "ok".

  2. The path comes first in the message, because the console truncates a
    row's detail at 70 characters. That truncation is why the path was cut off in
    your reproduction.

  3. One correction here, and it may save you time later: the manifest did
    carry the message, under detail rather than error.
    There is no error
    key in a manifest row, so .get("error") is None for every row including
    the healthy ones — the 585 ok rows have no error either. detail holds
    up to 400 characters and had the full path all along. Checking a
    netgen_timeout row in the current manifest:

    {"status": "netgen_timeout", "detail": "BNG2.pl netgen timed out after 180.0s"}

    So the report could have answered it; the key name was the obstacle, not the
    record. I have not renamed or aliased the key — that would rewrite the shape
    of a tracked 592-row manifest for a naming preference, which is your call to
    make, not a fix to slip into this PR.

The test

python/tests/test_gen_networks_cache_dir.py — 6 cases, no BNG2.pl needed,
which is the point: it drives main() with a model filter matching nothing, so
the run reaches its "nothing to do" exit without generating anything. That is
the ordering under test — #492 was not that the cache was never created, it was
that it was not created first.

Negative-verified precisely: removing only the ensure_cache_dir call (leaving
the helper and everything else in place) fails that one test on its assertion,
with the run still printing cache: <path> for a path that does not exist. The
other five cover idempotence on the 585-network resume, and the three
classification branches — vanished directory, a real ENOENT with the directory
present, and a non-ENOENT PermissionError (a full disk or read-only mount is
not a missing directory).

For the run machine

This is safe to take before the measurement phase, on your reasoning: it changes
what happens when the cache is missing, and nothing about what is generated
once it exists. The generated .net bytes are untouched — the only new code
paths are a mkdir and an error-message branch that a successful run never
reaches.

The mkdir -p workaround stays valid, so nothing forces the timing of this.

Phase 1 of the full-network ODE benchmark writes every generated network
into NETS = HERE / "nets" and never created that directory. On a machine
where it did not exist -- a fresh checkout, or one where the cache had
been cleaned to reclaim disk -- every model failed.
Expensively, and that is the part worth naming: the generation work
happens first and only the write fails, so six models spent 28 s to
produce nothing, and the full corpus would spend hours the same way. A
missing cache is also exactly the state in which this script is the
documented recovery path, so the one script whose job is to *build* the
cache was the one script that could not create it.
The pass now calls ensure_cache_dir(NETS) before the header, so the
"cache: ..." line it prints names a directory that exists rather than one
every write is about to fail against.
That makes the old failure unreachable on a normal run, which leaves the
case it cannot cover: the directory removed *while* a multi-hour sweep is
in flight. describe_write_failure() handles that one and says so in those
words. It matters because the old message named the wrong thing --
FileNotFoundError on a .net path reads as "this model could not be
generated", which is what a genuine BNG2.pl netgen failure looks like too
-- and because the console truncates a row's detail, so the path the
reader needed was cut off. The new detail leads with the directory, and
the row's status is cache_write_failed rather than a bare error, so the
console tag alone answers netgen-or-write without needing the detail at
all. Nothing downstream enumerates statuses; run_timing.py only compares
against "ok".
test_gen_networks_cache_dir.py drives main() with a model filter that
matches nothing, so it reaches the "nothing to do" exit without BNG2.pl
and without generating anything -- which is the ordering under test. #492
was not that the cache was never created, it was that it was not created
first. Removing only the ensure_cache_dir call fails that test on its
assertion, with the run still printing "cache: <path>" for a path that
does not exist.
One correction to the issue's third point: the manifest did carry the
message, under "detail" rather than "error" -- `.get("error")` is None
because there is no such key. The 400-char detail held the full path all
along; it was the 70-char console truncation that hid it.
@wshlavacek
wshlavacek merged commit f8be826 into mainAug 27, 2026
4 checks passed
@wshlavacek
wshlavacek deleted the bench-gen-networks-cache-dir-492 branch August 27, 2026 18:11
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

gen_networks.py does not create its own nets/ cache directory, so every model fails on a machine without it

1 participant

@wshlavacek
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' benchmarks: create the .net cache gen_networks.py exists to fill (#492) by wshlavacek · Pull Request #494 · lanl/bngsim · GitHub
Skip to content

benchmarks: create the .net cache gen_networks.py exists to fill (#492) - #494

Merged
wshlavacek merged 1 commit into
mainfrom
bench-gen-networks-cache-dir-492
Aug 27, 2026
Merged

benchmarks: create the .net cache gen_networks.py exists to fill (#492)#494
wshlavacek merged 1 commit into
mainfrom
bench-gen-networks-cache-dir-492

Conversation

@wshlavacek

Copy link
Copy Markdown
Collaborator

Closes#492.

Taken as asked — this one is on a recovery path, so the mkdir and the errno
check land together.

The fix

ensure_cache_dir(NETS) runs before the header, so the cache: ... line the
pass prints names a directory that exists rather than one every write is about
to fail against.

$ rm -rf benchmarks/suites/ode_fullnet/nets
$ python gen_networks.py --models Lang2024 --workers 1
cache: .../ode_fullnet/nets # now true when printed
[1/1] ok 0.41s Lang_2024.bngl N=73 rxn=332

The errno check, paired as you asked

Once the directory is created first, the old failure is unreachable on a normal
run. What is left is the case mkdir cannot cover — the directory removed
while a multi-hour sweep is in flight — and that is what
describe_write_failure() is for:

.../ode_fullnet/nets (the .net cache directory) no longer exists -- it is
created at the start of the pass, so it was removed mid-run; the network
itself generated fine

Three details, matching the three things you said conspired to make it hard to
read:

  1. It names the directory, not the file. And the row's status is now
    cache_write_failed rather than a bare error, so the console tag alone
    answers netgen-or-write without needing the detail at all. Nothing downstream
    enumerates statuses — run_timing.py only compares against "ok".

  2. The path comes first in the message, because the console truncates a
    row's detail at 70 characters. That truncation is why the path was cut off in
    your reproduction.

  3. One correction here, and it may save you time later: the manifest did
    carry the message, under detail rather than error.
    There is no error
    key in a manifest row, so .get("error") is None for every row including
    the healthy ones — the 585 ok rows have no error either. detail holds
    up to 400 characters and had the full path all along. Checking a
    netgen_timeout row in the current manifest:

    {"status": "netgen_timeout", "detail": "BNG2.pl netgen timed out after 180.0s"}

    So the report could have answered it; the key name was the obstacle, not the
    record. I have not renamed or aliased the key — that would rewrite the shape
    of a tracked 592-row manifest for a naming preference, which is your call to
    make, not a fix to slip into this PR.

The test

python/tests/test_gen_networks_cache_dir.py — 6 cases, no BNG2.pl needed,
which is the point: it drives main() with a model filter matching nothing, so
the run reaches its "nothing to do" exit without generating anything. That is
the ordering under test — #492 was not that the cache was never created, it was
that it was not created first.

Negative-verified precisely: removing only the ensure_cache_dir call (leaving
the helper and everything else in place) fails that one test on its assertion,
with the run still printing cache: <path> for a path that does not exist. The
other five cover idempotence on the 585-network resume, and the three
classification branches — vanished directory, a real ENOENT with the directory
present, and a non-ENOENT PermissionError (a full disk or read-only mount is
not a missing directory).

For the run machine

This is safe to take before the measurement phase, on your reasoning: it changes
what happens when the cache is missing, and nothing about what is generated
once it exists. The generated .net bytes are untouched — the only new code
paths are a mkdir and an error-message branch that a successful run never
reaches.

The mkdir -p workaround stays valid, so nothing forces the timing of this.

Phase 1 of the full-network ODE benchmark writes every generated network
into NETS = HERE / "nets" and never created that directory. On a machine
where it did not exist -- a fresh checkout, or one where the cache had
been cleaned to reclaim disk -- every model failed.
Expensively, and that is the part worth naming: the generation work
happens first and only the write fails, so six models spent 28 s to
produce nothing, and the full corpus would spend hours the same way. A
missing cache is also exactly the state in which this script is the
documented recovery path, so the one script whose job is to *build* the
cache was the one script that could not create it.
The pass now calls ensure_cache_dir(NETS) before the header, so the
"cache: ..." line it prints names a directory that exists rather than one
every write is about to fail against.
That makes the old failure unreachable on a normal run, which leaves the
case it cannot cover: the directory removed *while* a multi-hour sweep is
in flight. describe_write_failure() handles that one and says so in those
words. It matters because the old message named the wrong thing --
FileNotFoundError on a .net path reads as "this model could not be
generated", which is what a genuine BNG2.pl netgen failure looks like too
-- and because the console truncates a row's detail, so the path the
reader needed was cut off. The new detail leads with the directory, and
the row's status is cache_write_failed rather than a bare error, so the
console tag alone answers netgen-or-write without needing the detail at
all. Nothing downstream enumerates statuses; run_timing.py only compares
against "ok".
test_gen_networks_cache_dir.py drives main() with a model filter that
matches nothing, so it reaches the "nothing to do" exit without BNG2.pl
and without generating anything -- which is the ordering under test. #492
was not that the cache was never created, it was that it was not created
first. Removing only the ensure_cache_dir call fails that test on its
assertion, with the run still printing "cache: <path>" for a path that
does not exist.
One correction to the issue's third point: the manifest did carry the
message, under "detail" rather than "error" -- `.get("error")` is None
because there is no such key. The 400-char detail held the full path all
along; it was the 70-char console truncation that hid it.
@wshlavacek
wshlavacek merged commit f8be826 into mainAug 27, 2026
4 checks passed
@wshlavacek
wshlavacek deleted the bench-gen-networks-cache-dir-492 branch August 27, 2026 18:11
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

gen_networks.py does not create its own nets/ cache directory, so every model fails on a machine without it

1 participant

@wshlavacek
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' benchmarks: create the .net cache gen_networks.py exists to fill (#492) by wshlavacek · Pull Request #494 · lanl/bngsim · GitHub
Skip to content

benchmarks: create the .net cache gen_networks.py exists to fill (#492) - #494

Merged
wshlavacek merged 1 commit into
mainfrom
bench-gen-networks-cache-dir-492
Aug 27, 2026
Merged

benchmarks: create the .net cache gen_networks.py exists to fill (#492)#494
wshlavacek merged 1 commit into
mainfrom
bench-gen-networks-cache-dir-492

Conversation

@wshlavacek

Copy link
Copy Markdown
Collaborator

Closes#492.

Taken as asked — this one is on a recovery path, so the mkdir and the errno
check land together.

The fix

ensure_cache_dir(NETS) runs before the header, so the cache: ... line the
pass prints names a directory that exists rather than one every write is about
to fail against.

$ rm -rf benchmarks/suites/ode_fullnet/nets
$ python gen_networks.py --models Lang2024 --workers 1
cache: .../ode_fullnet/nets # now true when printed
[1/1] ok 0.41s Lang_2024.bngl N=73 rxn=332

The errno check, paired as you asked

Once the directory is created first, the old failure is unreachable on a normal
run. What is left is the case mkdir cannot cover — the directory removed
while a multi-hour sweep is in flight — and that is what
describe_write_failure() is for:

.../ode_fullnet/nets (the .net cache directory) no longer exists -- it is
created at the start of the pass, so it was removed mid-run; the network
itself generated fine

Three details, matching the three things you said conspired to make it hard to
read:

  1. It names the directory, not the file. And the row's status is now
    cache_write_failed rather than a bare error, so the console tag alone
    answers netgen-or-write without needing the detail at all. Nothing downstream
    enumerates statuses — run_timing.py only compares against "ok".

  2. The path comes first in the message, because the console truncates a
    row's detail at 70 characters. That truncation is why the path was cut off in
    your reproduction.

  3. One correction here, and it may save you time later: the manifest did
    carry the message, under detail rather than error.
    There is no error
    key in a manifest row, so .get("error") is None for every row including
    the healthy ones — the 585 ok rows have no error either. detail holds
    up to 400 characters and had the full path all along. Checking a
    netgen_timeout row in the current manifest:

    {"status": "netgen_timeout", "detail": "BNG2.pl netgen timed out after 180.0s"}

    So the report could have answered it; the key name was the obstacle, not the
    record. I have not renamed or aliased the key — that would rewrite the shape
    of a tracked 592-row manifest for a naming preference, which is your call to
    make, not a fix to slip into this PR.

The test

python/tests/test_gen_networks_cache_dir.py — 6 cases, no BNG2.pl needed,
which is the point: it drives main() with a model filter matching nothing, so
the run reaches its "nothing to do" exit without generating anything. That is
the ordering under test — #492 was not that the cache was never created, it was
that it was not created first.

Negative-verified precisely: removing only the ensure_cache_dir call (leaving
the helper and everything else in place) fails that one test on its assertion,
with the run still printing cache: <path> for a path that does not exist. The
other five cover idempotence on the 585-network resume, and the three
classification branches — vanished directory, a real ENOENT with the directory
present, and a non-ENOENT PermissionError (a full disk or read-only mount is
not a missing directory).

For the run machine

This is safe to take before the measurement phase, on your reasoning: it changes
what happens when the cache is missing, and nothing about what is generated
once it exists. The generated .net bytes are untouched — the only new code
paths are a mkdir and an error-message branch that a successful run never
reaches.

The mkdir -p workaround stays valid, so nothing forces the timing of this.

Phase 1 of the full-network ODE benchmark writes every generated network
into NETS = HERE / "nets" and never created that directory. On a machine
where it did not exist -- a fresh checkout, or one where the cache had
been cleaned to reclaim disk -- every model failed.
Expensively, and that is the part worth naming: the generation work
happens first and only the write fails, so six models spent 28 s to
produce nothing, and the full corpus would spend hours the same way. A
missing cache is also exactly the state in which this script is the
documented recovery path, so the one script whose job is to *build* the
cache was the one script that could not create it.
The pass now calls ensure_cache_dir(NETS) before the header, so the
"cache: ..." line it prints names a directory that exists rather than one
every write is about to fail against.
That makes the old failure unreachable on a normal run, which leaves the
case it cannot cover: the directory removed *while* a multi-hour sweep is
in flight. describe_write_failure() handles that one and says so in those
words. It matters because the old message named the wrong thing --
FileNotFoundError on a .net path reads as "this model could not be
generated", which is what a genuine BNG2.pl netgen failure looks like too
-- and because the console truncates a row's detail, so the path the
reader needed was cut off. The new detail leads with the directory, and
the row's status is cache_write_failed rather than a bare error, so the
console tag alone answers netgen-or-write without needing the detail at
all. Nothing downstream enumerates statuses; run_timing.py only compares
against "ok".
test_gen_networks_cache_dir.py drives main() with a model filter that
matches nothing, so it reaches the "nothing to do" exit without BNG2.pl
and without generating anything -- which is the ordering under test. #492
was not that the cache was never created, it was that it was not created
first. Removing only the ensure_cache_dir call fails that test on its
assertion, with the run still printing "cache: <path>" for a path that
does not exist.
One correction to the issue's third point: the manifest did carry the
message, under "detail" rather than "error" -- `.get("error")` is None
because there is no such key. The 400-char detail held the full path all
along; it was the 70-char console truncation that hid it.
@wshlavacek
wshlavacek merged commit f8be826 into mainAug 27, 2026
4 checks passed
@wshlavacek
wshlavacek deleted the bench-gen-networks-cache-dir-492 branch August 27, 2026 18:11
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

gen_networks.py does not create its own nets/ cache directory, so every model fails on a machine without it

1 participant

@wshlavacek
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' benchmarks: create the .net cache gen_networks.py exists to fill (#492) by wshlavacek · Pull Request #494 · lanl/bngsim · GitHub
Skip to content

benchmarks: create the .net cache gen_networks.py exists to fill (#492) - #494

Merged
wshlavacek merged 1 commit into
mainfrom
bench-gen-networks-cache-dir-492
Aug 27, 2026
Merged

benchmarks: create the .net cache gen_networks.py exists to fill (#492)#494
wshlavacek merged 1 commit into
mainfrom
bench-gen-networks-cache-dir-492

Conversation

@wshlavacek

Copy link
Copy Markdown
Collaborator

Closes#492.

Taken as asked — this one is on a recovery path, so the mkdir and the errno
check land together.

The fix

ensure_cache_dir(NETS) runs before the header, so the cache: ... line the
pass prints names a directory that exists rather than one every write is about
to fail against.

$ rm -rf benchmarks/suites/ode_fullnet/nets
$ python gen_networks.py --models Lang2024 --workers 1
cache: .../ode_fullnet/nets # now true when printed
[1/1] ok 0.41s Lang_2024.bngl N=73 rxn=332

The errno check, paired as you asked

Once the directory is created first, the old failure is unreachable on a normal
run. What is left is the case mkdir cannot cover — the directory removed
while a multi-hour sweep is in flight — and that is what
describe_write_failure() is for:

.../ode_fullnet/nets (the .net cache directory) no longer exists -- it is
created at the start of the pass, so it was removed mid-run; the network
itself generated fine

Three details, matching the three things you said conspired to make it hard to
read:

  1. It names the directory, not the file. And the row's status is now
    cache_write_failed rather than a bare error, so the console tag alone
    answers netgen-or-write without needing the detail at all. Nothing downstream
    enumerates statuses — run_timing.py only compares against "ok".

  2. The path comes first in the message, because the console truncates a
    row's detail at 70 characters. That truncation is why the path was cut off in
    your reproduction.

  3. One correction here, and it may save you time later: the manifest did
    carry the message, under detail rather than error.
    There is no error
    key in a manifest row, so .get("error") is None for every row including
    the healthy ones — the 585 ok rows have no error either. detail holds
    up to 400 characters and had the full path all along. Checking a
    netgen_timeout row in the current manifest:

    {"status": "netgen_timeout", "detail": "BNG2.pl netgen timed out after 180.0s"}

    So the report could have answered it; the key name was the obstacle, not the
    record. I have not renamed or aliased the key — that would rewrite the shape
    of a tracked 592-row manifest for a naming preference, which is your call to
    make, not a fix to slip into this PR.

The test

python/tests/test_gen_networks_cache_dir.py — 6 cases, no BNG2.pl needed,
which is the point: it drives main() with a model filter matching nothing, so
the run reaches its "nothing to do" exit without generating anything. That is
the ordering under test — #492 was not that the cache was never created, it was
that it was not created first.

Negative-verified precisely: removing only the ensure_cache_dir call (leaving
the helper and everything else in place) fails that one test on its assertion,
with the run still printing cache: <path> for a path that does not exist. The
other five cover idempotence on the 585-network resume, and the three
classification branches — vanished directory, a real ENOENT with the directory
present, and a non-ENOENT PermissionError (a full disk or read-only mount is
not a missing directory).

For the run machine

This is safe to take before the measurement phase, on your reasoning: it changes
what happens when the cache is missing, and nothing about what is generated
once it exists. The generated .net bytes are untouched — the only new code
paths are a mkdir and an error-message branch that a successful run never
reaches.

The mkdir -p workaround stays valid, so nothing forces the timing of this.

Phase 1 of the full-network ODE benchmark writes every generated network
into NETS = HERE / "nets" and never created that directory. On a machine
where it did not exist -- a fresh checkout, or one where the cache had
been cleaned to reclaim disk -- every model failed.
Expensively, and that is the part worth naming: the generation work
happens first and only the write fails, so six models spent 28 s to
produce nothing, and the full corpus would spend hours the same way. A
missing cache is also exactly the state in which this script is the
documented recovery path, so the one script whose job is to *build* the
cache was the one script that could not create it.
The pass now calls ensure_cache_dir(NETS) before the header, so the
"cache: ..." line it prints names a directory that exists rather than one
every write is about to fail against.
That makes the old failure unreachable on a normal run, which leaves the
case it cannot cover: the directory removed *while* a multi-hour sweep is
in flight. describe_write_failure() handles that one and says so in those
words. It matters because the old message named the wrong thing --
FileNotFoundError on a .net path reads as "this model could not be
generated", which is what a genuine BNG2.pl netgen failure looks like too
-- and because the console truncates a row's detail, so the path the
reader needed was cut off. The new detail leads with the directory, and
the row's status is cache_write_failed rather than a bare error, so the
console tag alone answers netgen-or-write without needing the detail at
all. Nothing downstream enumerates statuses; run_timing.py only compares
against "ok".
test_gen_networks_cache_dir.py drives main() with a model filter that
matches nothing, so it reaches the "nothing to do" exit without BNG2.pl
and without generating anything -- which is the ordering under test. #492
was not that the cache was never created, it was that it was not created
first. Removing only the ensure_cache_dir call fails that test on its
assertion, with the run still printing "cache: <path>" for a path that
does not exist.
One correction to the issue's third point: the manifest did carry the
message, under "detail" rather than "error" -- `.get("error")` is None
because there is no such key. The 400-char detail held the full path all
along; it was the 70-char console truncation that hid it.
@wshlavacek
wshlavacek merged commit f8be826 into mainAug 27, 2026
4 checks passed
@wshlavacek
wshlavacek deleted the bench-gen-networks-cache-dir-492 branch August 27, 2026 18:11
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

gen_networks.py does not create its own nets/ cache directory, so every model fails on a machine without it

1 participant

@wshlavacek
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); benchmarks: create the .net cache gen_networks.py exists to fill (#492) by wshlavacek · Pull Request #494 · lanl/bngsim · GitHub
Skip to content

benchmarks: create the .net cache gen_networks.py exists to fill (#492) - #494

Merged
wshlavacek merged 1 commit into
mainfrom
bench-gen-networks-cache-dir-492
Aug 27, 2026
Merged

benchmarks: create the .net cache gen_networks.py exists to fill (#492)#494
wshlavacek merged 1 commit into
mainfrom
bench-gen-networks-cache-dir-492

Conversation

@wshlavacek

Copy link
Copy Markdown
Collaborator

Closes#492.

Taken as asked — this one is on a recovery path, so the mkdir and the errno
check land together.

The fix

ensure_cache_dir(NETS) runs before the header, so the cache: ... line the
pass prints names a directory that exists rather than one every write is about
to fail against.

$ rm -rf benchmarks/suites/ode_fullnet/nets
$ python gen_networks.py --models Lang2024 --workers 1
cache: .../ode_fullnet/nets # now true when printed
[1/1] ok 0.41s Lang_2024.bngl N=73 rxn=332

The errno check, paired as you asked

Once the directory is created first, the old failure is unreachable on a normal
run. What is left is the case mkdir cannot cover — the directory removed
while a multi-hour sweep is in flight — and that is what
describe_write_failure() is for:

.../ode_fullnet/nets (the .net cache directory) no longer exists -- it is
created at the start of the pass, so it was removed mid-run; the network
itself generated fine

Three details, matching the three things you said conspired to make it hard to
read:

  1. It names the directory, not the file. And the row's status is now
    cache_write_failed rather than a bare error, so the console tag alone
    answers netgen-or-write without needing the detail at all. Nothing downstream
    enumerates statuses — run_timing.py only compares against "ok".

  2. The path comes first in the message, because the console truncates a
    row's detail at 70 characters. That truncation is why the path was cut off in
    your reproduction.

  3. One correction here, and it may save you time later: the manifest did
    carry the message, under detail rather than error.
    There is no error
    key in a manifest row, so .get("error") is None for every row including
    the healthy ones — the 585 ok rows have no error either. detail holds
    up to 400 characters and had the full path all along. Checking a
    netgen_timeout row in the current manifest:

    {"status": "netgen_timeout", "detail": "BNG2.pl netgen timed out after 180.0s"}

    So the report could have answered it; the key name was the obstacle, not the
    record. I have not renamed or aliased the key — that would rewrite the shape
    of a tracked 592-row manifest for a naming preference, which is your call to
    make, not a fix to slip into this PR.

The test

python/tests/test_gen_networks_cache_dir.py — 6 cases, no BNG2.pl needed,
which is the point: it drives main() with a model filter matching nothing, so
the run reaches its "nothing to do" exit without generating anything. That is
the ordering under test — #492 was not that the cache was never created, it was
that it was not created first.

Negative-verified precisely: removing only the ensure_cache_dir call (leaving
the helper and everything else in place) fails that one test on its assertion,
with the run still printing cache: <path> for a path that does not exist. The
other five cover idempotence on the 585-network resume, and the three
classification branches — vanished directory, a real ENOENT with the directory
present, and a non-ENOENT PermissionError (a full disk or read-only mount is
not a missing directory).

For the run machine

This is safe to take before the measurement phase, on your reasoning: it changes
what happens when the cache is missing, and nothing about what is generated
once it exists. The generated .net bytes are untouched — the only new code
paths are a mkdir and an error-message branch that a successful run never
reaches.

The mkdir -p workaround stays valid, so nothing forces the timing of this.

Phase 1 of the full-network ODE benchmark writes every generated network
into NETS = HERE / "nets" and never created that directory. On a machine
where it did not exist -- a fresh checkout, or one where the cache had
been cleaned to reclaim disk -- every model failed.
Expensively, and that is the part worth naming: the generation work
happens first and only the write fails, so six models spent 28 s to
produce nothing, and the full corpus would spend hours the same way. A
missing cache is also exactly the state in which this script is the
documented recovery path, so the one script whose job is to *build* the
cache was the one script that could not create it.
The pass now calls ensure_cache_dir(NETS) before the header, so the
"cache: ..." line it prints names a directory that exists rather than one
every write is about to fail against.
That makes the old failure unreachable on a normal run, which leaves the
case it cannot cover: the directory removed *while* a multi-hour sweep is
in flight. describe_write_failure() handles that one and says so in those
words. It matters because the old message named the wrong thing --
FileNotFoundError on a .net path reads as "this model could not be
generated", which is what a genuine BNG2.pl netgen failure looks like too
-- and because the console truncates a row's detail, so the path the
reader needed was cut off. The new detail leads with the directory, and
the row's status is cache_write_failed rather than a bare error, so the
console tag alone answers netgen-or-write without needing the detail at
all. Nothing downstream enumerates statuses; run_timing.py only compares
against "ok".
test_gen_networks_cache_dir.py drives main() with a model filter that
matches nothing, so it reaches the "nothing to do" exit without BNG2.pl
and without generating anything -- which is the ordering under test. #492
was not that the cache was never created, it was that it was not created
first. Removing only the ensure_cache_dir call fails that test on its
assertion, with the run still printing "cache: <path>" for a path that
does not exist.
One correction to the issue's third point: the manifest did carry the
message, under "detail" rather than "error" -- `.get("error")` is None
because there is no such key. The 400-char detail held the full path all
along; it was the 70-char console truncation that hid it.
@wshlavacek
wshlavacek merged commit f8be826 into mainAug 27, 2026
4 checks passed
@wshlavacek
wshlavacek deleted the bench-gen-networks-cache-dir-492 branch August 27, 2026 18:11
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

gen_networks.py does not create its own nets/ cache directory, so every model fails on a machine without it

1 participant

@wshlavacek