Calibrate the CMA-ES stagnation tolerance from the objective, not from a step length - #654

Merged
wshlavacek merged 1 commit into
mainfrom
fix/cmaes-tolfun-objective-scale
Aug 24, 2026
Merged

Calibrate the CMA-ES stagnation tolerance from the objective, not from a step length#654
wshlavacek merged 1 commit into
mainfrom
fix/cmaes-tolfun-objective-scale

Conversation

@wshlavacek

Copy link
Copy Markdown
Collaborator

Closes#653. This is the sibling of #648, in the optimizer #648's own ADR says it mirrors,
and it should have been found and fixed in the same pass.

The defect

ADR-0106 gave cmaes_tolfun its own key because it is a range in objective units while
cmaes_stop_tol is a step length in sampling space. Then it had an unset cmaes_tolfun
fall back to it anyway. The comment sitting on the line above the fallback says the two
"have no common scale and cannot share one well-set value". The line below makes them
share one. cmaes_stop_tol defaults to 1e-11, so the stagnation range was 1e-11 in
objective units.

This is #648 in the mirror.

DE (#648)CMA-ES (this)
inheritedstop_tolerance = 0.002, a dimensionless ratiocmaes_stop_tol = 1e-11, a sampling-space step
read asabsolute objective rangeabsolute objective range
effectfar too loosefar too strict
resultstops mid-descent, reports a wrong answer as convergedTolFun never fires

What the strict direction costs is the trigger the restart battery exists for. Its own
docstring calls it "the trigger the reproduction problems need", and the battery is there
because otherwise a run "polishes a local basin forever and never yields to a restart (the
IPOP/BIPOP machinery silently degenerates to one trapped run)". TolX and ConditionCov are
unaffected.

docs/config_keys.rst already told readers the default was "rarely what you want if you
rely on stagnation restarts". The defect was documented rather than fixed.

Why #648's remedy does not transfer

#648 was repaired by restoring a legacy meaning: stop_tolerance had always been a ratio,
so reading it as one again returned to known-correct behaviour. cmaes_stop_tol was never
an objective quantity, so a default has to be invented rather than restored.

Two candidates were rejected before the third, and the second is worth recording:

  • A fraction of the current objective. What ADR-0106 removed, correctly. On a
    likelihood |f| grows as the fit improves, so the threshold rises fastest where firing
    it costs most.
  • A fraction of the window being tested. Circular, and silently fatal:
    frange <= fraction * frange is never true for a small fraction, so the trigger is
    disabled rather than corrected. I tried this first and ADR-0106's own
    test_cmaes_tolfun_still_fires_on_a_genuinely_flat_history caught it immediately. That
    is the value of having kept that test.

The fix

An unset cmaes_tolfun is 1e-11 times the objective spread across the first scored
generation's population
.

  • Right units, taken from the objective rather than borrowed across a unit boundary.
  • Does not drift with |f|, because it is fixed at the first generation, before
    anything has converged. ADR-0106's objection does not reach it.
  • Not the window under test, so it is not circular.
  • Calibrated once; every IPOP/BIPOP restart reuses it. A later restart starts nearer
    the optimum and would measure a smaller spread, so recalibrating would hold the late,
    large-population restarts to the strictest bar, which is the shape of failure ADR-0106
    fixed.

The fraction is chosen so a problem whose initial population spans one objective unit gets
exactly the 1e-11 this key always defaulted to. A reference-scaled problem is unchanged
by construction; everything else scales in proportion. The run logs the value it picked.

A generation that cannot supply a spread (fewer than two finite scores, or all identical)
keeps the old fallback rather than inventing a number or setting zero. An explicit
cmaes_tolfun is never touched.

Evidence

All three of ADR-0106's regression tests pass unchanged. They construct synthetic
distribution state without scoring a generation, so they still exercise the fallback and
still assert alg.tolfun == alg.stop_tol.

Six new tests: the defect at its decision point (including that the old threshold stays
silent on a run the new one correctly stops), the anchor value, scaling across six decades,
the explicit key, three degenerate populations, and the no-recalibration-on-restart rule.

Full suite including the slow and recovery tiers CI skips: 4778 passed, 13 skipped, none
failed
.

Scope

Only reachable with cmaes_restarts > 0, which is not the default, and nothing in the
shipped corpus sets it. It cannot produce a wrong answer; it weakens the search.

The class, swept

A grep for the same pattern finds exactly two fallbacks in the codebase that cross a unit
boundary this way: de_tolfun and cmaes_tolfun. Both are now resolved. cmaes_run_maxgen
also defaults from unset, but to infinity rather than to another key, so no boundary is
crossed. ADR-0128 records the lesson: treat a defect whose ADR names a sibling as a defect
in a class, and check the sibling in the same pass.

…m a step length (#653)
This is the sibling of #648, in the optimizer #648's own ADR says it mirrors.
ADR-0106 gave cmaes_tolfun its own key because it is a range in objective units
while cmaes_stop_tol is a step length in sampling space, and then had an unset
cmaes_tolfun fall back to it anyway. The comment on the line above the fallback
says the two have no common scale and cannot share one well-set value. The line
below it makes them share one. cmaes_stop_tol defaults to 1e-11, so the
stagnation range was 1e-11 in objective units.
That is #648 in the mirror. There a dimensionless ratio read as an objective
range was far too loose, and fits stopped early reporting a wrong answer. Here a
step length read as an objective range is far too strict, so on an objective of
ordinary magnitude the trigger never fires. What that costs is the trigger the
restart battery exists for: without it a run polishes a local basin and never
yields to a restart, which is the failure the battery was built to prevent. The
documentation already told readers the default was rarely what they wanted,
which described the defect rather than fixing it.
The #648 remedy does not transfer. There, stop_tolerance had always been a ratio,
so reading it as one again returned to known-correct behaviour. cmaes_stop_tol
was never an objective quantity, so a default has to be invented rather than
restored. Two candidates were rejected first. A fraction of the current objective
is what ADR-0106 removed, because a likelihood's magnitude grows as the fit
improves. A fraction of the window being tested is circular, and silently
disables the trigger rather than fixing it; ADR-0106's own flat-history test
caught that attempt, which is the value of having kept it.
An unset cmaes_tolfun is now 1e-11 times the objective spread across the first
scored generation's population. That spread measures how much this objective
varies over the search box, it is in the units the tolerance needs, and it is
taken before anything has converged so it does not drift with the objective. It
is calibrated once and every restart reuses it, so a late restart is not held to
a stricter bar than an early one. The fraction is chosen so a problem whose
initial population spans one objective unit gets exactly the 1e-11 this key
always defaulted to, leaving a reference-scaled problem unchanged. A generation
that cannot supply a spread keeps the old fallback rather than inventing a
number, and an explicit cmaes_tolfun is never touched.
All three of ADR-0106's regression tests pass unchanged: they build synthetic
distribution state without scoring a generation, so they still exercise the
fallback. Six new tests cover the calibration, the anchor value, the scaling
across decades, the explicit key, the degenerate populations and the
no-recalibration-on-restart rule.
A sweep for the same pattern finds exactly two fallbacks in the codebase that
cross a unit boundary this way, de_tolfun and cmaes_tolfun. Both are now
resolved. cmaes_run_maxgen also defaults from unset, but to infinity rather than
to another key, so no boundary is crossed.
Only affects cmaes_restarts > 0, which is not the default. Full suite including
the slow and recovery tiers: 4778 passed, 13 skipped, none failed.
@wshlavacek
wshlavacek merged commit 6f08137 into mainAug 24, 2026
9 checks passed
@wshlavacek
wshlavacek deleted the fix/cmaes-tolfun-objective-scale branch August 24, 2026 03:05
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

An unset cmaes_tolfun inherits a sampling-space step length as an objective range, so the TolFun restart trigger never fires

1 participant

@wshlavacek
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Calibrate the CMA-ES stagnation tolerance from the objective, not from a step length - #654

Merged
wshlavacek merged 1 commit into
mainfrom
fix/cmaes-tolfun-objective-scale
Aug 24, 2026
Merged

Calibrate the CMA-ES stagnation tolerance from the objective, not from a step length#654
wshlavacek merged 1 commit into
mainfrom
fix/cmaes-tolfun-objective-scale

Conversation

@wshlavacek

Copy link
Copy Markdown
Collaborator

Closes#653. This is the sibling of #648, in the optimizer #648's own ADR says it mirrors,
and it should have been found and fixed in the same pass.

The defect

ADR-0106 gave cmaes_tolfun its own key because it is a range in objective units while
cmaes_stop_tol is a step length in sampling space. Then it had an unset cmaes_tolfun
fall back to it anyway. The comment sitting on the line above the fallback says the two
"have no common scale and cannot share one well-set value". The line below makes them
share one. cmaes_stop_tol defaults to 1e-11, so the stagnation range was 1e-11 in
objective units.

This is #648 in the mirror.

DE (#648)CMA-ES (this)
inheritedstop_tolerance = 0.002, a dimensionless ratiocmaes_stop_tol = 1e-11, a sampling-space step
read asabsolute objective rangeabsolute objective range
effectfar too loosefar too strict
resultstops mid-descent, reports a wrong answer as convergedTolFun never fires

What the strict direction costs is the trigger the restart battery exists for. Its own
docstring calls it "the trigger the reproduction problems need", and the battery is there
because otherwise a run "polishes a local basin forever and never yields to a restart (the
IPOP/BIPOP machinery silently degenerates to one trapped run)". TolX and ConditionCov are
unaffected.

docs/config_keys.rst already told readers the default was "rarely what you want if you
rely on stagnation restarts". The defect was documented rather than fixed.

Why #648's remedy does not transfer

#648 was repaired by restoring a legacy meaning: stop_tolerance had always been a ratio,
so reading it as one again returned to known-correct behaviour. cmaes_stop_tol was never
an objective quantity, so a default has to be invented rather than restored.

Two candidates were rejected before the third, and the second is worth recording:

  • A fraction of the current objective. What ADR-0106 removed, correctly. On a
    likelihood |f| grows as the fit improves, so the threshold rises fastest where firing
    it costs most.
  • A fraction of the window being tested. Circular, and silently fatal:
    frange <= fraction * frange is never true for a small fraction, so the trigger is
    disabled rather than corrected. I tried this first and ADR-0106's own
    test_cmaes_tolfun_still_fires_on_a_genuinely_flat_history caught it immediately. That
    is the value of having kept that test.

The fix

An unset cmaes_tolfun is 1e-11 times the objective spread across the first scored
generation's population
.

  • Right units, taken from the objective rather than borrowed across a unit boundary.
  • Does not drift with |f|, because it is fixed at the first generation, before
    anything has converged. ADR-0106's objection does not reach it.
  • Not the window under test, so it is not circular.
  • Calibrated once; every IPOP/BIPOP restart reuses it. A later restart starts nearer
    the optimum and would measure a smaller spread, so recalibrating would hold the late,
    large-population restarts to the strictest bar, which is the shape of failure ADR-0106
    fixed.

The fraction is chosen so a problem whose initial population spans one objective unit gets
exactly the 1e-11 this key always defaulted to. A reference-scaled problem is unchanged
by construction; everything else scales in proportion. The run logs the value it picked.

A generation that cannot supply a spread (fewer than two finite scores, or all identical)
keeps the old fallback rather than inventing a number or setting zero. An explicit
cmaes_tolfun is never touched.

Evidence

All three of ADR-0106's regression tests pass unchanged. They construct synthetic
distribution state without scoring a generation, so they still exercise the fallback and
still assert alg.tolfun == alg.stop_tol.

Six new tests: the defect at its decision point (including that the old threshold stays
silent on a run the new one correctly stops), the anchor value, scaling across six decades,
the explicit key, three degenerate populations, and the no-recalibration-on-restart rule.

Full suite including the slow and recovery tiers CI skips: 4778 passed, 13 skipped, none
failed
.

Scope

Only reachable with cmaes_restarts > 0, which is not the default, and nothing in the
shipped corpus sets it. It cannot produce a wrong answer; it weakens the search.

The class, swept

A grep for the same pattern finds exactly two fallbacks in the codebase that cross a unit
boundary this way: de_tolfun and cmaes_tolfun. Both are now resolved. cmaes_run_maxgen
also defaults from unset, but to infinity rather than to another key, so no boundary is
crossed. ADR-0128 records the lesson: treat a defect whose ADR names a sibling as a defect
in a class, and check the sibling in the same pass.

…m a step length (#653)
This is the sibling of #648, in the optimizer #648's own ADR says it mirrors.
ADR-0106 gave cmaes_tolfun its own key because it is a range in objective units
while cmaes_stop_tol is a step length in sampling space, and then had an unset
cmaes_tolfun fall back to it anyway. The comment on the line above the fallback
says the two have no common scale and cannot share one well-set value. The line
below it makes them share one. cmaes_stop_tol defaults to 1e-11, so the
stagnation range was 1e-11 in objective units.
That is #648 in the mirror. There a dimensionless ratio read as an objective
range was far too loose, and fits stopped early reporting a wrong answer. Here a
step length read as an objective range is far too strict, so on an objective of
ordinary magnitude the trigger never fires. What that costs is the trigger the
restart battery exists for: without it a run polishes a local basin and never
yields to a restart, which is the failure the battery was built to prevent. The
documentation already told readers the default was rarely what they wanted,
which described the defect rather than fixing it.
The #648 remedy does not transfer. There, stop_tolerance had always been a ratio,
so reading it as one again returned to known-correct behaviour. cmaes_stop_tol
was never an objective quantity, so a default has to be invented rather than
restored. Two candidates were rejected first. A fraction of the current objective
is what ADR-0106 removed, because a likelihood's magnitude grows as the fit
improves. A fraction of the window being tested is circular, and silently
disables the trigger rather than fixing it; ADR-0106's own flat-history test
caught that attempt, which is the value of having kept it.
An unset cmaes_tolfun is now 1e-11 times the objective spread across the first
scored generation's population. That spread measures how much this objective
varies over the search box, it is in the units the tolerance needs, and it is
taken before anything has converged so it does not drift with the objective. It
is calibrated once and every restart reuses it, so a late restart is not held to
a stricter bar than an early one. The fraction is chosen so a problem whose
initial population spans one objective unit gets exactly the 1e-11 this key
always defaulted to, leaving a reference-scaled problem unchanged. A generation
that cannot supply a spread keeps the old fallback rather than inventing a
number, and an explicit cmaes_tolfun is never touched.
All three of ADR-0106's regression tests pass unchanged: they build synthetic
distribution state without scoring a generation, so they still exercise the
fallback. Six new tests cover the calibration, the anchor value, the scaling
across decades, the explicit key, the degenerate populations and the
no-recalibration-on-restart rule.
A sweep for the same pattern finds exactly two fallbacks in the codebase that
cross a unit boundary this way, de_tolfun and cmaes_tolfun. Both are now
resolved. cmaes_run_maxgen also defaults from unset, but to infinity rather than
to another key, so no boundary is crossed.
Only affects cmaes_restarts > 0, which is not the default. Full suite including
the slow and recovery tiers: 4778 passed, 13 skipped, none failed.
@wshlavacek
wshlavacek merged commit 6f08137 into mainAug 24, 2026
9 checks passed
@wshlavacek
wshlavacek deleted the fix/cmaes-tolfun-objective-scale branch August 24, 2026 03:05
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

An unset cmaes_tolfun inherits a sampling-space step length as an objective range, so the TolFun restart trigger never fires

1 participant

@wshlavacek
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Calibrate the CMA-ES stagnation tolerance from the objective, not from a step length - #654

Merged
wshlavacek merged 1 commit into
mainfrom
fix/cmaes-tolfun-objective-scale
Aug 24, 2026
Merged

Calibrate the CMA-ES stagnation tolerance from the objective, not from a step length#654
wshlavacek merged 1 commit into
mainfrom
fix/cmaes-tolfun-objective-scale

Conversation

@wshlavacek

Copy link
Copy Markdown
Collaborator

Closes#653. This is the sibling of #648, in the optimizer #648's own ADR says it mirrors,
and it should have been found and fixed in the same pass.

The defect

ADR-0106 gave cmaes_tolfun its own key because it is a range in objective units while
cmaes_stop_tol is a step length in sampling space. Then it had an unset cmaes_tolfun
fall back to it anyway. The comment sitting on the line above the fallback says the two
"have no common scale and cannot share one well-set value". The line below makes them
share one. cmaes_stop_tol defaults to 1e-11, so the stagnation range was 1e-11 in
objective units.

This is #648 in the mirror.

DE (#648)CMA-ES (this)
inheritedstop_tolerance = 0.002, a dimensionless ratiocmaes_stop_tol = 1e-11, a sampling-space step
read asabsolute objective rangeabsolute objective range
effectfar too loosefar too strict
resultstops mid-descent, reports a wrong answer as convergedTolFun never fires

What the strict direction costs is the trigger the restart battery exists for. Its own
docstring calls it "the trigger the reproduction problems need", and the battery is there
because otherwise a run "polishes a local basin forever and never yields to a restart (the
IPOP/BIPOP machinery silently degenerates to one trapped run)". TolX and ConditionCov are
unaffected.

docs/config_keys.rst already told readers the default was "rarely what you want if you
rely on stagnation restarts". The defect was documented rather than fixed.

Why #648's remedy does not transfer

#648 was repaired by restoring a legacy meaning: stop_tolerance had always been a ratio,
so reading it as one again returned to known-correct behaviour. cmaes_stop_tol was never
an objective quantity, so a default has to be invented rather than restored.

Two candidates were rejected before the third, and the second is worth recording:

  • A fraction of the current objective. What ADR-0106 removed, correctly. On a
    likelihood |f| grows as the fit improves, so the threshold rises fastest where firing
    it costs most.
  • A fraction of the window being tested. Circular, and silently fatal:
    frange <= fraction * frange is never true for a small fraction, so the trigger is
    disabled rather than corrected. I tried this first and ADR-0106's own
    test_cmaes_tolfun_still_fires_on_a_genuinely_flat_history caught it immediately. That
    is the value of having kept that test.

The fix

An unset cmaes_tolfun is 1e-11 times the objective spread across the first scored
generation's population
.

  • Right units, taken from the objective rather than borrowed across a unit boundary.
  • Does not drift with |f|, because it is fixed at the first generation, before
    anything has converged. ADR-0106's objection does not reach it.
  • Not the window under test, so it is not circular.
  • Calibrated once; every IPOP/BIPOP restart reuses it. A later restart starts nearer
    the optimum and would measure a smaller spread, so recalibrating would hold the late,
    large-population restarts to the strictest bar, which is the shape of failure ADR-0106
    fixed.

The fraction is chosen so a problem whose initial population spans one objective unit gets
exactly the 1e-11 this key always defaulted to. A reference-scaled problem is unchanged
by construction; everything else scales in proportion. The run logs the value it picked.

A generation that cannot supply a spread (fewer than two finite scores, or all identical)
keeps the old fallback rather than inventing a number or setting zero. An explicit
cmaes_tolfun is never touched.

Evidence

All three of ADR-0106's regression tests pass unchanged. They construct synthetic
distribution state without scoring a generation, so they still exercise the fallback and
still assert alg.tolfun == alg.stop_tol.

Six new tests: the defect at its decision point (including that the old threshold stays
silent on a run the new one correctly stops), the anchor value, scaling across six decades,
the explicit key, three degenerate populations, and the no-recalibration-on-restart rule.

Full suite including the slow and recovery tiers CI skips: 4778 passed, 13 skipped, none
failed
.

Scope

Only reachable with cmaes_restarts > 0, which is not the default, and nothing in the
shipped corpus sets it. It cannot produce a wrong answer; it weakens the search.

The class, swept

A grep for the same pattern finds exactly two fallbacks in the codebase that cross a unit
boundary this way: de_tolfun and cmaes_tolfun. Both are now resolved. cmaes_run_maxgen
also defaults from unset, but to infinity rather than to another key, so no boundary is
crossed. ADR-0128 records the lesson: treat a defect whose ADR names a sibling as a defect
in a class, and check the sibling in the same pass.

…m a step length (#653)
This is the sibling of #648, in the optimizer #648's own ADR says it mirrors.
ADR-0106 gave cmaes_tolfun its own key because it is a range in objective units
while cmaes_stop_tol is a step length in sampling space, and then had an unset
cmaes_tolfun fall back to it anyway. The comment on the line above the fallback
says the two have no common scale and cannot share one well-set value. The line
below it makes them share one. cmaes_stop_tol defaults to 1e-11, so the
stagnation range was 1e-11 in objective units.
That is #648 in the mirror. There a dimensionless ratio read as an objective
range was far too loose, and fits stopped early reporting a wrong answer. Here a
step length read as an objective range is far too strict, so on an objective of
ordinary magnitude the trigger never fires. What that costs is the trigger the
restart battery exists for: without it a run polishes a local basin and never
yields to a restart, which is the failure the battery was built to prevent. The
documentation already told readers the default was rarely what they wanted,
which described the defect rather than fixing it.
The #648 remedy does not transfer. There, stop_tolerance had always been a ratio,
so reading it as one again returned to known-correct behaviour. cmaes_stop_tol
was never an objective quantity, so a default has to be invented rather than
restored. Two candidates were rejected first. A fraction of the current objective
is what ADR-0106 removed, because a likelihood's magnitude grows as the fit
improves. A fraction of the window being tested is circular, and silently
disables the trigger rather than fixing it; ADR-0106's own flat-history test
caught that attempt, which is the value of having kept it.
An unset cmaes_tolfun is now 1e-11 times the objective spread across the first
scored generation's population. That spread measures how much this objective
varies over the search box, it is in the units the tolerance needs, and it is
taken before anything has converged so it does not drift with the objective. It
is calibrated once and every restart reuses it, so a late restart is not held to
a stricter bar than an early one. The fraction is chosen so a problem whose
initial population spans one objective unit gets exactly the 1e-11 this key
always defaulted to, leaving a reference-scaled problem unchanged. A generation
that cannot supply a spread keeps the old fallback rather than inventing a
number, and an explicit cmaes_tolfun is never touched.
All three of ADR-0106's regression tests pass unchanged: they build synthetic
distribution state without scoring a generation, so they still exercise the
fallback. Six new tests cover the calibration, the anchor value, the scaling
across decades, the explicit key, the degenerate populations and the
no-recalibration-on-restart rule.
A sweep for the same pattern finds exactly two fallbacks in the codebase that
cross a unit boundary this way, de_tolfun and cmaes_tolfun. Both are now
resolved. cmaes_run_maxgen also defaults from unset, but to infinity rather than
to another key, so no boundary is crossed.
Only affects cmaes_restarts > 0, which is not the default. Full suite including
the slow and recovery tiers: 4778 passed, 13 skipped, none failed.
@wshlavacek
wshlavacek merged commit 6f08137 into mainAug 24, 2026
9 checks passed
@wshlavacek
wshlavacek deleted the fix/cmaes-tolfun-objective-scale branch August 24, 2026 03:05
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

An unset cmaes_tolfun inherits a sampling-space step length as an objective range, so the TolFun restart trigger never fires

1 participant

@wshlavacek
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Calibrate the CMA-ES stagnation tolerance from the objective, not from a step length - #654

Merged
wshlavacek merged 1 commit into
mainfrom
fix/cmaes-tolfun-objective-scale
Aug 24, 2026
Merged

Calibrate the CMA-ES stagnation tolerance from the objective, not from a step length#654
wshlavacek merged 1 commit into
mainfrom
fix/cmaes-tolfun-objective-scale

Conversation

@wshlavacek

Copy link
Copy Markdown
Collaborator

Closes#653. This is the sibling of #648, in the optimizer #648's own ADR says it mirrors,
and it should have been found and fixed in the same pass.

The defect

ADR-0106 gave cmaes_tolfun its own key because it is a range in objective units while
cmaes_stop_tol is a step length in sampling space. Then it had an unset cmaes_tolfun
fall back to it anyway. The comment sitting on the line above the fallback says the two
"have no common scale and cannot share one well-set value". The line below makes them
share one. cmaes_stop_tol defaults to 1e-11, so the stagnation range was 1e-11 in
objective units.

This is #648 in the mirror.

DE (#648)CMA-ES (this)
inheritedstop_tolerance = 0.002, a dimensionless ratiocmaes_stop_tol = 1e-11, a sampling-space step
read asabsolute objective rangeabsolute objective range
effectfar too loosefar too strict
resultstops mid-descent, reports a wrong answer as convergedTolFun never fires

What the strict direction costs is the trigger the restart battery exists for. Its own
docstring calls it "the trigger the reproduction problems need", and the battery is there
because otherwise a run "polishes a local basin forever and never yields to a restart (the
IPOP/BIPOP machinery silently degenerates to one trapped run)". TolX and ConditionCov are
unaffected.

docs/config_keys.rst already told readers the default was "rarely what you want if you
rely on stagnation restarts". The defect was documented rather than fixed.

Why #648's remedy does not transfer

#648 was repaired by restoring a legacy meaning: stop_tolerance had always been a ratio,
so reading it as one again returned to known-correct behaviour. cmaes_stop_tol was never
an objective quantity, so a default has to be invented rather than restored.

Two candidates were rejected before the third, and the second is worth recording:

  • A fraction of the current objective. What ADR-0106 removed, correctly. On a
    likelihood |f| grows as the fit improves, so the threshold rises fastest where firing
    it costs most.
  • A fraction of the window being tested. Circular, and silently fatal:
    frange <= fraction * frange is never true for a small fraction, so the trigger is
    disabled rather than corrected. I tried this first and ADR-0106's own
    test_cmaes_tolfun_still_fires_on_a_genuinely_flat_history caught it immediately. That
    is the value of having kept that test.

The fix

An unset cmaes_tolfun is 1e-11 times the objective spread across the first scored
generation's population
.

  • Right units, taken from the objective rather than borrowed across a unit boundary.
  • Does not drift with |f|, because it is fixed at the first generation, before
    anything has converged. ADR-0106's objection does not reach it.
  • Not the window under test, so it is not circular.
  • Calibrated once; every IPOP/BIPOP restart reuses it. A later restart starts nearer
    the optimum and would measure a smaller spread, so recalibrating would hold the late,
    large-population restarts to the strictest bar, which is the shape of failure ADR-0106
    fixed.

The fraction is chosen so a problem whose initial population spans one objective unit gets
exactly the 1e-11 this key always defaulted to. A reference-scaled problem is unchanged
by construction; everything else scales in proportion. The run logs the value it picked.

A generation that cannot supply a spread (fewer than two finite scores, or all identical)
keeps the old fallback rather than inventing a number or setting zero. An explicit
cmaes_tolfun is never touched.

Evidence

All three of ADR-0106's regression tests pass unchanged. They construct synthetic
distribution state without scoring a generation, so they still exercise the fallback and
still assert alg.tolfun == alg.stop_tol.

Six new tests: the defect at its decision point (including that the old threshold stays
silent on a run the new one correctly stops), the anchor value, scaling across six decades,
the explicit key, three degenerate populations, and the no-recalibration-on-restart rule.

Full suite including the slow and recovery tiers CI skips: 4778 passed, 13 skipped, none
failed
.

Scope

Only reachable with cmaes_restarts > 0, which is not the default, and nothing in the
shipped corpus sets it. It cannot produce a wrong answer; it weakens the search.

The class, swept

A grep for the same pattern finds exactly two fallbacks in the codebase that cross a unit
boundary this way: de_tolfun and cmaes_tolfun. Both are now resolved. cmaes_run_maxgen
also defaults from unset, but to infinity rather than to another key, so no boundary is
crossed. ADR-0128 records the lesson: treat a defect whose ADR names a sibling as a defect
in a class, and check the sibling in the same pass.

…m a step length (#653)
This is the sibling of #648, in the optimizer #648's own ADR says it mirrors.
ADR-0106 gave cmaes_tolfun its own key because it is a range in objective units
while cmaes_stop_tol is a step length in sampling space, and then had an unset
cmaes_tolfun fall back to it anyway. The comment on the line above the fallback
says the two have no common scale and cannot share one well-set value. The line
below it makes them share one. cmaes_stop_tol defaults to 1e-11, so the
stagnation range was 1e-11 in objective units.
That is #648 in the mirror. There a dimensionless ratio read as an objective
range was far too loose, and fits stopped early reporting a wrong answer. Here a
step length read as an objective range is far too strict, so on an objective of
ordinary magnitude the trigger never fires. What that costs is the trigger the
restart battery exists for: without it a run polishes a local basin and never
yields to a restart, which is the failure the battery was built to prevent. The
documentation already told readers the default was rarely what they wanted,
which described the defect rather than fixing it.
The #648 remedy does not transfer. There, stop_tolerance had always been a ratio,
so reading it as one again returned to known-correct behaviour. cmaes_stop_tol
was never an objective quantity, so a default has to be invented rather than
restored. Two candidates were rejected first. A fraction of the current objective
is what ADR-0106 removed, because a likelihood's magnitude grows as the fit
improves. A fraction of the window being tested is circular, and silently
disables the trigger rather than fixing it; ADR-0106's own flat-history test
caught that attempt, which is the value of having kept it.
An unset cmaes_tolfun is now 1e-11 times the objective spread across the first
scored generation's population. That spread measures how much this objective
varies over the search box, it is in the units the tolerance needs, and it is
taken before anything has converged so it does not drift with the objective. It
is calibrated once and every restart reuses it, so a late restart is not held to
a stricter bar than an early one. The fraction is chosen so a problem whose
initial population spans one objective unit gets exactly the 1e-11 this key
always defaulted to, leaving a reference-scaled problem unchanged. A generation
that cannot supply a spread keeps the old fallback rather than inventing a
number, and an explicit cmaes_tolfun is never touched.
All three of ADR-0106's regression tests pass unchanged: they build synthetic
distribution state without scoring a generation, so they still exercise the
fallback. Six new tests cover the calibration, the anchor value, the scaling
across decades, the explicit key, the degenerate populations and the
no-recalibration-on-restart rule.
A sweep for the same pattern finds exactly two fallbacks in the codebase that
cross a unit boundary this way, de_tolfun and cmaes_tolfun. Both are now
resolved. cmaes_run_maxgen also defaults from unset, but to infinity rather than
to another key, so no boundary is crossed.
Only affects cmaes_restarts > 0, which is not the default. Full suite including
the slow and recovery tiers: 4778 passed, 13 skipped, none failed.
@wshlavacek
wshlavacek merged commit 6f08137 into mainAug 24, 2026
9 checks passed
@wshlavacek
wshlavacek deleted the fix/cmaes-tolfun-objective-scale branch August 24, 2026 03:05
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

An unset cmaes_tolfun inherits a sampling-space step length as an objective range, so the TolFun restart trigger never fires

1 participant

@wshlavacek
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Calibrate the CMA-ES stagnation tolerance from the objective, not from a step length - #654

Merged
wshlavacek merged 1 commit into
mainfrom
fix/cmaes-tolfun-objective-scale
Aug 24, 2026
Merged

Calibrate the CMA-ES stagnation tolerance from the objective, not from a step length#654
wshlavacek merged 1 commit into
mainfrom
fix/cmaes-tolfun-objective-scale

Conversation

@wshlavacek

Copy link
Copy Markdown
Collaborator

Closes#653. This is the sibling of #648, in the optimizer #648's own ADR says it mirrors,
and it should have been found and fixed in the same pass.

The defect

ADR-0106 gave cmaes_tolfun its own key because it is a range in objective units while
cmaes_stop_tol is a step length in sampling space. Then it had an unset cmaes_tolfun
fall back to it anyway. The comment sitting on the line above the fallback says the two
"have no common scale and cannot share one well-set value". The line below makes them
share one. cmaes_stop_tol defaults to 1e-11, so the stagnation range was 1e-11 in
objective units.

This is #648 in the mirror.

DE (#648)CMA-ES (this)
inheritedstop_tolerance = 0.002, a dimensionless ratiocmaes_stop_tol = 1e-11, a sampling-space step
read asabsolute objective rangeabsolute objective range
effectfar too loosefar too strict
resultstops mid-descent, reports a wrong answer as convergedTolFun never fires

What the strict direction costs is the trigger the restart battery exists for. Its own
docstring calls it "the trigger the reproduction problems need", and the battery is there
because otherwise a run "polishes a local basin forever and never yields to a restart (the
IPOP/BIPOP machinery silently degenerates to one trapped run)". TolX and ConditionCov are
unaffected.

docs/config_keys.rst already told readers the default was "rarely what you want if you
rely on stagnation restarts". The defect was documented rather than fixed.

Why #648's remedy does not transfer

#648 was repaired by restoring a legacy meaning: stop_tolerance had always been a ratio,
so reading it as one again returned to known-correct behaviour. cmaes_stop_tol was never
an objective quantity, so a default has to be invented rather than restored.

Two candidates were rejected before the third, and the second is worth recording:

  • A fraction of the current objective. What ADR-0106 removed, correctly. On a
    likelihood |f| grows as the fit improves, so the threshold rises fastest where firing
    it costs most.
  • A fraction of the window being tested. Circular, and silently fatal:
    frange <= fraction * frange is never true for a small fraction, so the trigger is
    disabled rather than corrected. I tried this first and ADR-0106's own
    test_cmaes_tolfun_still_fires_on_a_genuinely_flat_history caught it immediately. That
    is the value of having kept that test.

The fix

An unset cmaes_tolfun is 1e-11 times the objective spread across the first scored
generation's population
.

  • Right units, taken from the objective rather than borrowed across a unit boundary.
  • Does not drift with |f|, because it is fixed at the first generation, before
    anything has converged. ADR-0106's objection does not reach it.
  • Not the window under test, so it is not circular.
  • Calibrated once; every IPOP/BIPOP restart reuses it. A later restart starts nearer
    the optimum and would measure a smaller spread, so recalibrating would hold the late,
    large-population restarts to the strictest bar, which is the shape of failure ADR-0106
    fixed.

The fraction is chosen so a problem whose initial population spans one objective unit gets
exactly the 1e-11 this key always defaulted to. A reference-scaled problem is unchanged
by construction; everything else scales in proportion. The run logs the value it picked.

A generation that cannot supply a spread (fewer than two finite scores, or all identical)
keeps the old fallback rather than inventing a number or setting zero. An explicit
cmaes_tolfun is never touched.

Evidence

All three of ADR-0106's regression tests pass unchanged. They construct synthetic
distribution state without scoring a generation, so they still exercise the fallback and
still assert alg.tolfun == alg.stop_tol.

Six new tests: the defect at its decision point (including that the old threshold stays
silent on a run the new one correctly stops), the anchor value, scaling across six decades,
the explicit key, three degenerate populations, and the no-recalibration-on-restart rule.

Full suite including the slow and recovery tiers CI skips: 4778 passed, 13 skipped, none
failed
.

Scope

Only reachable with cmaes_restarts > 0, which is not the default, and nothing in the
shipped corpus sets it. It cannot produce a wrong answer; it weakens the search.

The class, swept

A grep for the same pattern finds exactly two fallbacks in the codebase that cross a unit
boundary this way: de_tolfun and cmaes_tolfun. Both are now resolved. cmaes_run_maxgen
also defaults from unset, but to infinity rather than to another key, so no boundary is
crossed. ADR-0128 records the lesson: treat a defect whose ADR names a sibling as a defect
in a class, and check the sibling in the same pass.

…m a step length (#653)
This is the sibling of #648, in the optimizer #648's own ADR says it mirrors.
ADR-0106 gave cmaes_tolfun its own key because it is a range in objective units
while cmaes_stop_tol is a step length in sampling space, and then had an unset
cmaes_tolfun fall back to it anyway. The comment on the line above the fallback
says the two have no common scale and cannot share one well-set value. The line
below it makes them share one. cmaes_stop_tol defaults to 1e-11, so the
stagnation range was 1e-11 in objective units.
That is #648 in the mirror. There a dimensionless ratio read as an objective
range was far too loose, and fits stopped early reporting a wrong answer. Here a
step length read as an objective range is far too strict, so on an objective of
ordinary magnitude the trigger never fires. What that costs is the trigger the
restart battery exists for: without it a run polishes a local basin and never
yields to a restart, which is the failure the battery was built to prevent. The
documentation already told readers the default was rarely what they wanted,
which described the defect rather than fixing it.
The #648 remedy does not transfer. There, stop_tolerance had always been a ratio,
so reading it as one again returned to known-correct behaviour. cmaes_stop_tol
was never an objective quantity, so a default has to be invented rather than
restored. Two candidates were rejected first. A fraction of the current objective
is what ADR-0106 removed, because a likelihood's magnitude grows as the fit
improves. A fraction of the window being tested is circular, and silently
disables the trigger rather than fixing it; ADR-0106's own flat-history test
caught that attempt, which is the value of having kept it.
An unset cmaes_tolfun is now 1e-11 times the objective spread across the first
scored generation's population. That spread measures how much this objective
varies over the search box, it is in the units the tolerance needs, and it is
taken before anything has converged so it does not drift with the objective. It
is calibrated once and every restart reuses it, so a late restart is not held to
a stricter bar than an early one. The fraction is chosen so a problem whose
initial population spans one objective unit gets exactly the 1e-11 this key
always defaulted to, leaving a reference-scaled problem unchanged. A generation
that cannot supply a spread keeps the old fallback rather than inventing a
number, and an explicit cmaes_tolfun is never touched.
All three of ADR-0106's regression tests pass unchanged: they build synthetic
distribution state without scoring a generation, so they still exercise the
fallback. Six new tests cover the calibration, the anchor value, the scaling
across decades, the explicit key, the degenerate populations and the
no-recalibration-on-restart rule.
A sweep for the same pattern finds exactly two fallbacks in the codebase that
cross a unit boundary this way, de_tolfun and cmaes_tolfun. Both are now
resolved. cmaes_run_maxgen also defaults from unset, but to infinity rather than
to another key, so no boundary is crossed.
Only affects cmaes_restarts > 0, which is not the default. Full suite including
the slow and recovery tiers: 4778 passed, 13 skipped, none failed.
@wshlavacek
wshlavacek merged commit 6f08137 into mainAug 24, 2026
9 checks passed
@wshlavacek
wshlavacek deleted the fix/cmaes-tolfun-objective-scale branch August 24, 2026 03:05
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

An unset cmaes_tolfun inherits a sampling-space step length as an objective range, so the TolFun restart trigger never fires

1 participant

@wshlavacek
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Calibrate the CMA-ES stagnation tolerance from the objective, not from a step length - #654

Merged
wshlavacek merged 1 commit into
mainfrom
fix/cmaes-tolfun-objective-scale
Aug 24, 2026
Merged

Calibrate the CMA-ES stagnation tolerance from the objective, not from a step length#654
wshlavacek merged 1 commit into
mainfrom
fix/cmaes-tolfun-objective-scale

Conversation

@wshlavacek

Copy link
Copy Markdown
Collaborator

Closes#653. This is the sibling of #648, in the optimizer #648's own ADR says it mirrors,
and it should have been found and fixed in the same pass.

The defect

ADR-0106 gave cmaes_tolfun its own key because it is a range in objective units while
cmaes_stop_tol is a step length in sampling space. Then it had an unset cmaes_tolfun
fall back to it anyway. The comment sitting on the line above the fallback says the two
"have no common scale and cannot share one well-set value". The line below makes them
share one. cmaes_stop_tol defaults to 1e-11, so the stagnation range was 1e-11 in
objective units.

This is #648 in the mirror.

DE (#648)CMA-ES (this)
inheritedstop_tolerance = 0.002, a dimensionless ratiocmaes_stop_tol = 1e-11, a sampling-space step
read asabsolute objective rangeabsolute objective range
effectfar too loosefar too strict
resultstops mid-descent, reports a wrong answer as convergedTolFun never fires

What the strict direction costs is the trigger the restart battery exists for. Its own
docstring calls it "the trigger the reproduction problems need", and the battery is there
because otherwise a run "polishes a local basin forever and never yields to a restart (the
IPOP/BIPOP machinery silently degenerates to one trapped run)". TolX and ConditionCov are
unaffected.

docs/config_keys.rst already told readers the default was "rarely what you want if you
rely on stagnation restarts". The defect was documented rather than fixed.

Why #648's remedy does not transfer

#648 was repaired by restoring a legacy meaning: stop_tolerance had always been a ratio,
so reading it as one again returned to known-correct behaviour. cmaes_stop_tol was never
an objective quantity, so a default has to be invented rather than restored.

Two candidates were rejected before the third, and the second is worth recording:

  • A fraction of the current objective. What ADR-0106 removed, correctly. On a
    likelihood |f| grows as the fit improves, so the threshold rises fastest where firing
    it costs most.
  • A fraction of the window being tested. Circular, and silently fatal:
    frange <= fraction * frange is never true for a small fraction, so the trigger is
    disabled rather than corrected. I tried this first and ADR-0106's own
    test_cmaes_tolfun_still_fires_on_a_genuinely_flat_history caught it immediately. That
    is the value of having kept that test.

The fix

An unset cmaes_tolfun is 1e-11 times the objective spread across the first scored
generation's population
.

  • Right units, taken from the objective rather than borrowed across a unit boundary.
  • Does not drift with |f|, because it is fixed at the first generation, before
    anything has converged. ADR-0106's objection does not reach it.
  • Not the window under test, so it is not circular.
  • Calibrated once; every IPOP/BIPOP restart reuses it. A later restart starts nearer
    the optimum and would measure a smaller spread, so recalibrating would hold the late,
    large-population restarts to the strictest bar, which is the shape of failure ADR-0106
    fixed.

The fraction is chosen so a problem whose initial population spans one objective unit gets
exactly the 1e-11 this key always defaulted to. A reference-scaled problem is unchanged
by construction; everything else scales in proportion. The run logs the value it picked.

A generation that cannot supply a spread (fewer than two finite scores, or all identical)
keeps the old fallback rather than inventing a number or setting zero. An explicit
cmaes_tolfun is never touched.

Evidence

All three of ADR-0106's regression tests pass unchanged. They construct synthetic
distribution state without scoring a generation, so they still exercise the fallback and
still assert alg.tolfun == alg.stop_tol.

Six new tests: the defect at its decision point (including that the old threshold stays
silent on a run the new one correctly stops), the anchor value, scaling across six decades,
the explicit key, three degenerate populations, and the no-recalibration-on-restart rule.

Full suite including the slow and recovery tiers CI skips: 4778 passed, 13 skipped, none
failed
.

Scope

Only reachable with cmaes_restarts > 0, which is not the default, and nothing in the
shipped corpus sets it. It cannot produce a wrong answer; it weakens the search.

The class, swept

A grep for the same pattern finds exactly two fallbacks in the codebase that cross a unit
boundary this way: de_tolfun and cmaes_tolfun. Both are now resolved. cmaes_run_maxgen
also defaults from unset, but to infinity rather than to another key, so no boundary is
crossed. ADR-0128 records the lesson: treat a defect whose ADR names a sibling as a defect
in a class, and check the sibling in the same pass.

…m a step length (#653)
This is the sibling of #648, in the optimizer #648's own ADR says it mirrors.
ADR-0106 gave cmaes_tolfun its own key because it is a range in objective units
while cmaes_stop_tol is a step length in sampling space, and then had an unset
cmaes_tolfun fall back to it anyway. The comment on the line above the fallback
says the two have no common scale and cannot share one well-set value. The line
below it makes them share one. cmaes_stop_tol defaults to 1e-11, so the
stagnation range was 1e-11 in objective units.
That is #648 in the mirror. There a dimensionless ratio read as an objective
range was far too loose, and fits stopped early reporting a wrong answer. Here a
step length read as an objective range is far too strict, so on an objective of
ordinary magnitude the trigger never fires. What that costs is the trigger the
restart battery exists for: without it a run polishes a local basin and never
yields to a restart, which is the failure the battery was built to prevent. The
documentation already told readers the default was rarely what they wanted,
which described the defect rather than fixing it.
The #648 remedy does not transfer. There, stop_tolerance had always been a ratio,
so reading it as one again returned to known-correct behaviour. cmaes_stop_tol
was never an objective quantity, so a default has to be invented rather than
restored. Two candidates were rejected first. A fraction of the current objective
is what ADR-0106 removed, because a likelihood's magnitude grows as the fit
improves. A fraction of the window being tested is circular, and silently
disables the trigger rather than fixing it; ADR-0106's own flat-history test
caught that attempt, which is the value of having kept it.
An unset cmaes_tolfun is now 1e-11 times the objective spread across the first
scored generation's population. That spread measures how much this objective
varies over the search box, it is in the units the tolerance needs, and it is
taken before anything has converged so it does not drift with the objective. It
is calibrated once and every restart reuses it, so a late restart is not held to
a stricter bar than an early one. The fraction is chosen so a problem whose
initial population spans one objective unit gets exactly the 1e-11 this key
always defaulted to, leaving a reference-scaled problem unchanged. A generation
that cannot supply a spread keeps the old fallback rather than inventing a
number, and an explicit cmaes_tolfun is never touched.
All three of ADR-0106's regression tests pass unchanged: they build synthetic
distribution state without scoring a generation, so they still exercise the
fallback. Six new tests cover the calibration, the anchor value, the scaling
across decades, the explicit key, the degenerate populations and the
no-recalibration-on-restart rule.
A sweep for the same pattern finds exactly two fallbacks in the codebase that
cross a unit boundary this way, de_tolfun and cmaes_tolfun. Both are now
resolved. cmaes_run_maxgen also defaults from unset, but to infinity rather than
to another key, so no boundary is crossed.
Only affects cmaes_restarts > 0, which is not the default. Full suite including
the slow and recovery tiers: 4778 passed, 13 skipped, none failed.
@wshlavacek
wshlavacek merged commit 6f08137 into mainAug 24, 2026
9 checks passed
@wshlavacek
wshlavacek deleted the fix/cmaes-tolfun-objective-scale branch August 24, 2026 03:05
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

An unset cmaes_tolfun inherits a sampling-space step length as an objective range, so the TolFun restart trigger never fires

1 participant

@wshlavacek
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Calibrate the CMA-ES stagnation tolerance from the objective, not from a step length - #654

Merged
wshlavacek merged 1 commit into
mainfrom
fix/cmaes-tolfun-objective-scale
Aug 24, 2026
Merged

Calibrate the CMA-ES stagnation tolerance from the objective, not from a step length#654
wshlavacek merged 1 commit into
mainfrom
fix/cmaes-tolfun-objective-scale

Conversation

@wshlavacek

Copy link
Copy Markdown
Collaborator

Closes#653. This is the sibling of #648, in the optimizer #648's own ADR says it mirrors,
and it should have been found and fixed in the same pass.

The defect

ADR-0106 gave cmaes_tolfun its own key because it is a range in objective units while
cmaes_stop_tol is a step length in sampling space. Then it had an unset cmaes_tolfun
fall back to it anyway. The comment sitting on the line above the fallback says the two
"have no common scale and cannot share one well-set value". The line below makes them
share one. cmaes_stop_tol defaults to 1e-11, so the stagnation range was 1e-11 in
objective units.

This is #648 in the mirror.

DE (#648)CMA-ES (this)
inheritedstop_tolerance = 0.002, a dimensionless ratiocmaes_stop_tol = 1e-11, a sampling-space step
read asabsolute objective rangeabsolute objective range
effectfar too loosefar too strict
resultstops mid-descent, reports a wrong answer as convergedTolFun never fires

What the strict direction costs is the trigger the restart battery exists for. Its own
docstring calls it "the trigger the reproduction problems need", and the battery is there
because otherwise a run "polishes a local basin forever and never yields to a restart (the
IPOP/BIPOP machinery silently degenerates to one trapped run)". TolX and ConditionCov are
unaffected.

docs/config_keys.rst already told readers the default was "rarely what you want if you
rely on stagnation restarts". The defect was documented rather than fixed.

Why #648's remedy does not transfer

#648 was repaired by restoring a legacy meaning: stop_tolerance had always been a ratio,
so reading it as one again returned to known-correct behaviour. cmaes_stop_tol was never
an objective quantity, so a default has to be invented rather than restored.

Two candidates were rejected before the third, and the second is worth recording:

  • A fraction of the current objective. What ADR-0106 removed, correctly. On a
    likelihood |f| grows as the fit improves, so the threshold rises fastest where firing
    it costs most.
  • A fraction of the window being tested. Circular, and silently fatal:
    frange <= fraction * frange is never true for a small fraction, so the trigger is
    disabled rather than corrected. I tried this first and ADR-0106's own
    test_cmaes_tolfun_still_fires_on_a_genuinely_flat_history caught it immediately. That
    is the value of having kept that test.

The fix

An unset cmaes_tolfun is 1e-11 times the objective spread across the first scored
generation's population
.

  • Right units, taken from the objective rather than borrowed across a unit boundary.
  • Does not drift with |f|, because it is fixed at the first generation, before
    anything has converged. ADR-0106's objection does not reach it.
  • Not the window under test, so it is not circular.
  • Calibrated once; every IPOP/BIPOP restart reuses it. A later restart starts nearer
    the optimum and would measure a smaller spread, so recalibrating would hold the late,
    large-population restarts to the strictest bar, which is the shape of failure ADR-0106
    fixed.

The fraction is chosen so a problem whose initial population spans one objective unit gets
exactly the 1e-11 this key always defaulted to. A reference-scaled problem is unchanged
by construction; everything else scales in proportion. The run logs the value it picked.

A generation that cannot supply a spread (fewer than two finite scores, or all identical)
keeps the old fallback rather than inventing a number or setting zero. An explicit
cmaes_tolfun is never touched.

Evidence

All three of ADR-0106's regression tests pass unchanged. They construct synthetic
distribution state without scoring a generation, so they still exercise the fallback and
still assert alg.tolfun == alg.stop_tol.

Six new tests: the defect at its decision point (including that the old threshold stays
silent on a run the new one correctly stops), the anchor value, scaling across six decades,
the explicit key, three degenerate populations, and the no-recalibration-on-restart rule.

Full suite including the slow and recovery tiers CI skips: 4778 passed, 13 skipped, none
failed
.

Scope

Only reachable with cmaes_restarts > 0, which is not the default, and nothing in the
shipped corpus sets it. It cannot produce a wrong answer; it weakens the search.

The class, swept

A grep for the same pattern finds exactly two fallbacks in the codebase that cross a unit
boundary this way: de_tolfun and cmaes_tolfun. Both are now resolved. cmaes_run_maxgen
also defaults from unset, but to infinity rather than to another key, so no boundary is
crossed. ADR-0128 records the lesson: treat a defect whose ADR names a sibling as a defect
in a class, and check the sibling in the same pass.

…m a step length (#653)
This is the sibling of #648, in the optimizer #648's own ADR says it mirrors.
ADR-0106 gave cmaes_tolfun its own key because it is a range in objective units
while cmaes_stop_tol is a step length in sampling space, and then had an unset
cmaes_tolfun fall back to it anyway. The comment on the line above the fallback
says the two have no common scale and cannot share one well-set value. The line
below it makes them share one. cmaes_stop_tol defaults to 1e-11, so the
stagnation range was 1e-11 in objective units.
That is #648 in the mirror. There a dimensionless ratio read as an objective
range was far too loose, and fits stopped early reporting a wrong answer. Here a
step length read as an objective range is far too strict, so on an objective of
ordinary magnitude the trigger never fires. What that costs is the trigger the
restart battery exists for: without it a run polishes a local basin and never
yields to a restart, which is the failure the battery was built to prevent. The
documentation already told readers the default was rarely what they wanted,
which described the defect rather than fixing it.
The #648 remedy does not transfer. There, stop_tolerance had always been a ratio,
so reading it as one again returned to known-correct behaviour. cmaes_stop_tol
was never an objective quantity, so a default has to be invented rather than
restored. Two candidates were rejected first. A fraction of the current objective
is what ADR-0106 removed, because a likelihood's magnitude grows as the fit
improves. A fraction of the window being tested is circular, and silently
disables the trigger rather than fixing it; ADR-0106's own flat-history test
caught that attempt, which is the value of having kept it.
An unset cmaes_tolfun is now 1e-11 times the objective spread across the first
scored generation's population. That spread measures how much this objective
varies over the search box, it is in the units the tolerance needs, and it is
taken before anything has converged so it does not drift with the objective. It
is calibrated once and every restart reuses it, so a late restart is not held to
a stricter bar than an early one. The fraction is chosen so a problem whose
initial population spans one objective unit gets exactly the 1e-11 this key
always defaulted to, leaving a reference-scaled problem unchanged. A generation
that cannot supply a spread keeps the old fallback rather than inventing a
number, and an explicit cmaes_tolfun is never touched.
All three of ADR-0106's regression tests pass unchanged: they build synthetic
distribution state without scoring a generation, so they still exercise the
fallback. Six new tests cover the calibration, the anchor value, the scaling
across decades, the explicit key, the degenerate populations and the
no-recalibration-on-restart rule.
A sweep for the same pattern finds exactly two fallbacks in the codebase that
cross a unit boundary this way, de_tolfun and cmaes_tolfun. Both are now
resolved. cmaes_run_maxgen also defaults from unset, but to infinity rather than
to another key, so no boundary is crossed.
Only affects cmaes_restarts > 0, which is not the default. Full suite including
the slow and recovery tiers: 4778 passed, 13 skipped, none failed.
@wshlavacek
wshlavacek merged commit 6f08137 into mainAug 24, 2026
9 checks passed
@wshlavacek
wshlavacek deleted the fix/cmaes-tolfun-objective-scale branch August 24, 2026 03:05
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

An unset cmaes_tolfun inherits a sampling-space step length as an objective range, so the TolFun restart trigger never fires

1 participant

@wshlavacek
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Calibrate the CMA-ES stagnation tolerance from the objective, not from a step length - #654

Merged
wshlavacek merged 1 commit into
mainfrom
fix/cmaes-tolfun-objective-scale
Aug 24, 2026
Merged

Calibrate the CMA-ES stagnation tolerance from the objective, not from a step length#654
wshlavacek merged 1 commit into
mainfrom
fix/cmaes-tolfun-objective-scale

Conversation

@wshlavacek

Copy link
Copy Markdown
Collaborator

Closes#653. This is the sibling of #648, in the optimizer #648's own ADR says it mirrors,
and it should have been found and fixed in the same pass.

The defect

ADR-0106 gave cmaes_tolfun its own key because it is a range in objective units while
cmaes_stop_tol is a step length in sampling space. Then it had an unset cmaes_tolfun
fall back to it anyway. The comment sitting on the line above the fallback says the two
"have no common scale and cannot share one well-set value". The line below makes them
share one. cmaes_stop_tol defaults to 1e-11, so the stagnation range was 1e-11 in
objective units.

This is #648 in the mirror.

DE (#648)CMA-ES (this)
inheritedstop_tolerance = 0.002, a dimensionless ratiocmaes_stop_tol = 1e-11, a sampling-space step
read asabsolute objective rangeabsolute objective range
effectfar too loosefar too strict
resultstops mid-descent, reports a wrong answer as convergedTolFun never fires

What the strict direction costs is the trigger the restart battery exists for. Its own
docstring calls it "the trigger the reproduction problems need", and the battery is there
because otherwise a run "polishes a local basin forever and never yields to a restart (the
IPOP/BIPOP machinery silently degenerates to one trapped run)". TolX and ConditionCov are
unaffected.

docs/config_keys.rst already told readers the default was "rarely what you want if you
rely on stagnation restarts". The defect was documented rather than fixed.

Why #648's remedy does not transfer

#648 was repaired by restoring a legacy meaning: stop_tolerance had always been a ratio,
so reading it as one again returned to known-correct behaviour. cmaes_stop_tol was never
an objective quantity, so a default has to be invented rather than restored.

Two candidates were rejected before the third, and the second is worth recording:

  • A fraction of the current objective. What ADR-0106 removed, correctly. On a
    likelihood |f| grows as the fit improves, so the threshold rises fastest where firing
    it costs most.
  • A fraction of the window being tested. Circular, and silently fatal:
    frange <= fraction * frange is never true for a small fraction, so the trigger is
    disabled rather than corrected. I tried this first and ADR-0106's own
    test_cmaes_tolfun_still_fires_on_a_genuinely_flat_history caught it immediately. That
    is the value of having kept that test.

The fix

An unset cmaes_tolfun is 1e-11 times the objective spread across the first scored
generation's population
.

  • Right units, taken from the objective rather than borrowed across a unit boundary.
  • Does not drift with |f|, because it is fixed at the first generation, before
    anything has converged. ADR-0106's objection does not reach it.
  • Not the window under test, so it is not circular.
  • Calibrated once; every IPOP/BIPOP restart reuses it. A later restart starts nearer
    the optimum and would measure a smaller spread, so recalibrating would hold the late,
    large-population restarts to the strictest bar, which is the shape of failure ADR-0106
    fixed.

The fraction is chosen so a problem whose initial population spans one objective unit gets
exactly the 1e-11 this key always defaulted to. A reference-scaled problem is unchanged
by construction; everything else scales in proportion. The run logs the value it picked.

A generation that cannot supply a spread (fewer than two finite scores, or all identical)
keeps the old fallback rather than inventing a number or setting zero. An explicit
cmaes_tolfun is never touched.

Evidence

All three of ADR-0106's regression tests pass unchanged. They construct synthetic
distribution state without scoring a generation, so they still exercise the fallback and
still assert alg.tolfun == alg.stop_tol.

Six new tests: the defect at its decision point (including that the old threshold stays
silent on a run the new one correctly stops), the anchor value, scaling across six decades,
the explicit key, three degenerate populations, and the no-recalibration-on-restart rule.

Full suite including the slow and recovery tiers CI skips: 4778 passed, 13 skipped, none
failed
.

Scope

Only reachable with cmaes_restarts > 0, which is not the default, and nothing in the
shipped corpus sets it. It cannot produce a wrong answer; it weakens the search.

The class, swept

A grep for the same pattern finds exactly two fallbacks in the codebase that cross a unit
boundary this way: de_tolfun and cmaes_tolfun. Both are now resolved. cmaes_run_maxgen
also defaults from unset, but to infinity rather than to another key, so no boundary is
crossed. ADR-0128 records the lesson: treat a defect whose ADR names a sibling as a defect
in a class, and check the sibling in the same pass.

…m a step length (#653)
This is the sibling of #648, in the optimizer #648's own ADR says it mirrors.
ADR-0106 gave cmaes_tolfun its own key because it is a range in objective units
while cmaes_stop_tol is a step length in sampling space, and then had an unset
cmaes_tolfun fall back to it anyway. The comment on the line above the fallback
says the two have no common scale and cannot share one well-set value. The line
below it makes them share one. cmaes_stop_tol defaults to 1e-11, so the
stagnation range was 1e-11 in objective units.
That is #648 in the mirror. There a dimensionless ratio read as an objective
range was far too loose, and fits stopped early reporting a wrong answer. Here a
step length read as an objective range is far too strict, so on an objective of
ordinary magnitude the trigger never fires. What that costs is the trigger the
restart battery exists for: without it a run polishes a local basin and never
yields to a restart, which is the failure the battery was built to prevent. The
documentation already told readers the default was rarely what they wanted,
which described the defect rather than fixing it.
The #648 remedy does not transfer. There, stop_tolerance had always been a ratio,
so reading it as one again returned to known-correct behaviour. cmaes_stop_tol
was never an objective quantity, so a default has to be invented rather than
restored. Two candidates were rejected first. A fraction of the current objective
is what ADR-0106 removed, because a likelihood's magnitude grows as the fit
improves. A fraction of the window being tested is circular, and silently
disables the trigger rather than fixing it; ADR-0106's own flat-history test
caught that attempt, which is the value of having kept it.
An unset cmaes_tolfun is now 1e-11 times the objective spread across the first
scored generation's population. That spread measures how much this objective
varies over the search box, it is in the units the tolerance needs, and it is
taken before anything has converged so it does not drift with the objective. It
is calibrated once and every restart reuses it, so a late restart is not held to
a stricter bar than an early one. The fraction is chosen so a problem whose
initial population spans one objective unit gets exactly the 1e-11 this key
always defaulted to, leaving a reference-scaled problem unchanged. A generation
that cannot supply a spread keeps the old fallback rather than inventing a
number, and an explicit cmaes_tolfun is never touched.
All three of ADR-0106's regression tests pass unchanged: they build synthetic
distribution state without scoring a generation, so they still exercise the
fallback. Six new tests cover the calibration, the anchor value, the scaling
across decades, the explicit key, the degenerate populations and the
no-recalibration-on-restart rule.
A sweep for the same pattern finds exactly two fallbacks in the codebase that
cross a unit boundary this way, de_tolfun and cmaes_tolfun. Both are now
resolved. cmaes_run_maxgen also defaults from unset, but to infinity rather than
to another key, so no boundary is crossed.
Only affects cmaes_restarts > 0, which is not the default. Full suite including
the slow and recovery tiers: 4778 passed, 13 skipped, none failed.
@wshlavacek
wshlavacek merged commit 6f08137 into mainAug 24, 2026
9 checks passed
@wshlavacek
wshlavacek deleted the fix/cmaes-tolfun-objective-scale branch August 24, 2026 03:05
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

An unset cmaes_tolfun inherits a sampling-space step length as an objective range, so the TolFun restart trigger never fires

1 participant

@wshlavacek