e2e failure propagation - #1

Merged
solsson merged 4 commits into
mainfrom
e2e-failure-propagation
Jul 13, 2026
Merged

e2e failure propagation#1
solsson merged 4 commits into
mainfrom
e2e-failure-propagation

Conversation

@solsson

Copy link
Copy Markdown
Contributor

No description provided.

Yolean k8s-qa (buckety maintainer)and others added 4 commits July 10, 2026 06:24
run_scenario is invoked in an if condition, where bash ignores
set -e inside the function body. Every failure inside it, including
a non-zero assert.sh, fell through to the PASS logging and returned
0, so the harness has never been able to report FAIL. The latest
green main run (27607054780) shows six scenario executions with
explicit [assert][FAIL] or wait timeouts, all recorded PASS. Each
step now propagates failure explicitly with || return 1.
Three more defects the same run exposed:
- kafka_topic_exists/s3_bucket_exists use kubectl run -i, which
forwards and consumes the caller's stdin. The harness fed its
scenario list on stdin of a while read loop, so the first assert
that ran an -i helper swallowed the remaining list: only 4 of 11
kadm scenarios ever executed. The list is now materialised with
mapfile and assert.sh runs with stdin from /dev/null.
- The helpers piped kubectl run into grep -q under pipefail, where
an early grep exit SIGPIPEs kubectl and fails the pipeline even
on a present topic, with stderr discarded. Output is now captured
and printed on failure so assertion failures are diagnosable.
- teardown between implementations deleted the CRDs while scenario
CRs with finalizers could still be terminating, leaving the CRD
stuck and the next implementation's applies rejected with "create
not allowed while custom resource definition is terminating"
(the entire versitygw block in that run). Overlays now converge
in place; nothing deletes CRDs mid-run.
Scenario-level fixes entailed by unmasking:
- driver-version defaulted E2E_IMAGE_* to a ghcr :dev tag that is
never published, wedging the Deployment for two rollout timeouts
per run. The vars are now pass-through inputs and the scenario
logs SKIPPED when they are unset; CI provides real rotation
images (separate commit).
- oob-drift waits 90-300s for periodic drift repair, but the
controller default recheck is 5m. e2e overlays now pass
--periodic-recheck=15s.
Failed scenarios leave their namespace standing per SPEC "E2E
harness and parity" #3; KEEP_FAILED now additionally preserves
namespaces of passing scenarios.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…n CI
The reconciler gated the driver-major comparison on
status.driverMajor != 0, treating 0 as "not yet stamped". Both
shipped drivers are 0.x, so every existing resource is stamped
major=0 and the check never ran: a 0.x to 1.0 driver bump would
silently auto-apply instead of surfacing DriverVersionIncompatible,
contradicting SPEC "Driver versioning". Stampedness is now
signalled by status.backend, which is set in the same first
reconcile pass.
The driver-version e2e scenario (SPEC scenario #9) needs controller
images built with rotated driver versions. CI never built them, so
the scenario has never actually run. The workflow now builds
patch (0.1.1) and major (1.0.0) driver builds per SCAFFOLDING
"Driver version stamping" and pushes them to the run's local
registry; the harness receives them via E2E_IMAGE_*/E2E_VERSION_*.
These images are rotation fixtures only, not digest-pinned and
not published.
Release pin bumped for the reconciler change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
First honest run (29073773069): 15 PASS, 6 FAIL. The six, with
root causes:
- driver-version: status.driverMajor read back empty. The field
was int with omitempty and the stamp used a merge patch, so a
stamped 0 (any 0.x driver) neither serializes nor diffs and can
never reach the API server. omitempty dropped; the first-reconcile
stamp is a Status().Update() since a merge diff of 0 -> 0 is
still empty.
- misconfigured-startup: the scenario is config-only
(resources: []) and plain kubectl apply -k errors with "no
objects passed to apply". The harness now renders first and
skips apply when the output is empty.
- s3/happy-path on versitygw AND minio: the roundtrip Job upload
and download both succeed, then the final comparison calls diff,
which does not exist in the aws-cli image. Compare via the shell
instead.
- kadm/oob-drift: assert bug. awk /retention.ms/ also matches
retention.local.target.ms, so the check compared a two-line
string against 3600000. The controller had in fact restored the
value. Exact-match awk.
- kadm/parameter-mutation: same awk bug, plus the broker check was
a single read immediately after Ready; the alter is acked before
the config is necessarily visible to DescribeConfigs. Exact-match
awk and a 60s poll like the oob-drift scenario already uses.
Release pin bumped for the status stamping change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
kadm/retention-policy flaked in one of two runs of the same sha:
topic creation is acked before a fresh rpk metadata request
necessarily lists the topic, so single-shot existence checks race.
kafka_topic_exists now polls (6x5s), and the deleted-topic check
moves into a matching kafka_topic_absent helper.
The same-sha double run came from push + pull_request both
triggering on PR branches; push is now main-only.
No binary change; release pin untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@solsson
solsson merged commit 79c8441 into mainJul 13, 2026
2 checks passed
@solssonsolsson mentioned this pull request Jul 13, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@solsson
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

e2e failure propagation - #1

Merged
solsson merged 4 commits into
mainfrom
e2e-failure-propagation
Jul 13, 2026
Merged

e2e failure propagation#1
solsson merged 4 commits into
mainfrom
e2e-failure-propagation

Conversation

@solsson

Copy link
Copy Markdown
Contributor

No description provided.

Yolean k8s-qa (buckety maintainer)and others added 4 commits July 10, 2026 06:24
run_scenario is invoked in an if condition, where bash ignores
set -e inside the function body. Every failure inside it, including
a non-zero assert.sh, fell through to the PASS logging and returned
0, so the harness has never been able to report FAIL. The latest
green main run (27607054780) shows six scenario executions with
explicit [assert][FAIL] or wait timeouts, all recorded PASS. Each
step now propagates failure explicitly with || return 1.
Three more defects the same run exposed:
- kafka_topic_exists/s3_bucket_exists use kubectl run -i, which
forwards and consumes the caller's stdin. The harness fed its
scenario list on stdin of a while read loop, so the first assert
that ran an -i helper swallowed the remaining list: only 4 of 11
kadm scenarios ever executed. The list is now materialised with
mapfile and assert.sh runs with stdin from /dev/null.
- The helpers piped kubectl run into grep -q under pipefail, where
an early grep exit SIGPIPEs kubectl and fails the pipeline even
on a present topic, with stderr discarded. Output is now captured
and printed on failure so assertion failures are diagnosable.
- teardown between implementations deleted the CRDs while scenario
CRs with finalizers could still be terminating, leaving the CRD
stuck and the next implementation's applies rejected with "create
not allowed while custom resource definition is terminating"
(the entire versitygw block in that run). Overlays now converge
in place; nothing deletes CRDs mid-run.
Scenario-level fixes entailed by unmasking:
- driver-version defaulted E2E_IMAGE_* to a ghcr :dev tag that is
never published, wedging the Deployment for two rollout timeouts
per run. The vars are now pass-through inputs and the scenario
logs SKIPPED when they are unset; CI provides real rotation
images (separate commit).
- oob-drift waits 90-300s for periodic drift repair, but the
controller default recheck is 5m. e2e overlays now pass
--periodic-recheck=15s.
Failed scenarios leave their namespace standing per SPEC "E2E
harness and parity" #3; KEEP_FAILED now additionally preserves
namespaces of passing scenarios.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…n CI
The reconciler gated the driver-major comparison on
status.driverMajor != 0, treating 0 as "not yet stamped". Both
shipped drivers are 0.x, so every existing resource is stamped
major=0 and the check never ran: a 0.x to 1.0 driver bump would
silently auto-apply instead of surfacing DriverVersionIncompatible,
contradicting SPEC "Driver versioning". Stampedness is now
signalled by status.backend, which is set in the same first
reconcile pass.
The driver-version e2e scenario (SPEC scenario #9) needs controller
images built with rotated driver versions. CI never built them, so
the scenario has never actually run. The workflow now builds
patch (0.1.1) and major (1.0.0) driver builds per SCAFFOLDING
"Driver version stamping" and pushes them to the run's local
registry; the harness receives them via E2E_IMAGE_*/E2E_VERSION_*.
These images are rotation fixtures only, not digest-pinned and
not published.
Release pin bumped for the reconciler change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
First honest run (29073773069): 15 PASS, 6 FAIL. The six, with
root causes:
- driver-version: status.driverMajor read back empty. The field
was int with omitempty and the stamp used a merge patch, so a
stamped 0 (any 0.x driver) neither serializes nor diffs and can
never reach the API server. omitempty dropped; the first-reconcile
stamp is a Status().Update() since a merge diff of 0 -> 0 is
still empty.
- misconfigured-startup: the scenario is config-only
(resources: []) and plain kubectl apply -k errors with "no
objects passed to apply". The harness now renders first and
skips apply when the output is empty.
- s3/happy-path on versitygw AND minio: the roundtrip Job upload
and download both succeed, then the final comparison calls diff,
which does not exist in the aws-cli image. Compare via the shell
instead.
- kadm/oob-drift: assert bug. awk /retention.ms/ also matches
retention.local.target.ms, so the check compared a two-line
string against 3600000. The controller had in fact restored the
value. Exact-match awk.
- kadm/parameter-mutation: same awk bug, plus the broker check was
a single read immediately after Ready; the alter is acked before
the config is necessarily visible to DescribeConfigs. Exact-match
awk and a 60s poll like the oob-drift scenario already uses.
Release pin bumped for the status stamping change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
kadm/retention-policy flaked in one of two runs of the same sha:
topic creation is acked before a fresh rpk metadata request
necessarily lists the topic, so single-shot existence checks race.
kafka_topic_exists now polls (6x5s), and the deleted-topic check
moves into a matching kafka_topic_absent helper.
The same-sha double run came from push + pull_request both
triggering on PR branches; push is now main-only.
No binary change; release pin untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@solsson
solsson merged commit 79c8441 into mainJul 13, 2026
2 checks passed
@solssonsolsson mentioned this pull request Jul 13, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@solsson
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

e2e failure propagation - #1

Merged
solsson merged 4 commits into
mainfrom
e2e-failure-propagation
Jul 13, 2026
Merged

e2e failure propagation#1
solsson merged 4 commits into
mainfrom
e2e-failure-propagation

Conversation

@solsson

Copy link
Copy Markdown
Contributor

No description provided.

Yolean k8s-qa (buckety maintainer)and others added 4 commits July 10, 2026 06:24
run_scenario is invoked in an if condition, where bash ignores
set -e inside the function body. Every failure inside it, including
a non-zero assert.sh, fell through to the PASS logging and returned
0, so the harness has never been able to report FAIL. The latest
green main run (27607054780) shows six scenario executions with
explicit [assert][FAIL] or wait timeouts, all recorded PASS. Each
step now propagates failure explicitly with || return 1.
Three more defects the same run exposed:
- kafka_topic_exists/s3_bucket_exists use kubectl run -i, which
forwards and consumes the caller's stdin. The harness fed its
scenario list on stdin of a while read loop, so the first assert
that ran an -i helper swallowed the remaining list: only 4 of 11
kadm scenarios ever executed. The list is now materialised with
mapfile and assert.sh runs with stdin from /dev/null.
- The helpers piped kubectl run into grep -q under pipefail, where
an early grep exit SIGPIPEs kubectl and fails the pipeline even
on a present topic, with stderr discarded. Output is now captured
and printed on failure so assertion failures are diagnosable.
- teardown between implementations deleted the CRDs while scenario
CRs with finalizers could still be terminating, leaving the CRD
stuck and the next implementation's applies rejected with "create
not allowed while custom resource definition is terminating"
(the entire versitygw block in that run). Overlays now converge
in place; nothing deletes CRDs mid-run.
Scenario-level fixes entailed by unmasking:
- driver-version defaulted E2E_IMAGE_* to a ghcr :dev tag that is
never published, wedging the Deployment for two rollout timeouts
per run. The vars are now pass-through inputs and the scenario
logs SKIPPED when they are unset; CI provides real rotation
images (separate commit).
- oob-drift waits 90-300s for periodic drift repair, but the
controller default recheck is 5m. e2e overlays now pass
--periodic-recheck=15s.
Failed scenarios leave their namespace standing per SPEC "E2E
harness and parity" #3; KEEP_FAILED now additionally preserves
namespaces of passing scenarios.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…n CI
The reconciler gated the driver-major comparison on
status.driverMajor != 0, treating 0 as "not yet stamped". Both
shipped drivers are 0.x, so every existing resource is stamped
major=0 and the check never ran: a 0.x to 1.0 driver bump would
silently auto-apply instead of surfacing DriverVersionIncompatible,
contradicting SPEC "Driver versioning". Stampedness is now
signalled by status.backend, which is set in the same first
reconcile pass.
The driver-version e2e scenario (SPEC scenario #9) needs controller
images built with rotated driver versions. CI never built them, so
the scenario has never actually run. The workflow now builds
patch (0.1.1) and major (1.0.0) driver builds per SCAFFOLDING
"Driver version stamping" and pushes them to the run's local
registry; the harness receives them via E2E_IMAGE_*/E2E_VERSION_*.
These images are rotation fixtures only, not digest-pinned and
not published.
Release pin bumped for the reconciler change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
First honest run (29073773069): 15 PASS, 6 FAIL. The six, with
root causes:
- driver-version: status.driverMajor read back empty. The field
was int with omitempty and the stamp used a merge patch, so a
stamped 0 (any 0.x driver) neither serializes nor diffs and can
never reach the API server. omitempty dropped; the first-reconcile
stamp is a Status().Update() since a merge diff of 0 -> 0 is
still empty.
- misconfigured-startup: the scenario is config-only
(resources: []) and plain kubectl apply -k errors with "no
objects passed to apply". The harness now renders first and
skips apply when the output is empty.
- s3/happy-path on versitygw AND minio: the roundtrip Job upload
and download both succeed, then the final comparison calls diff,
which does not exist in the aws-cli image. Compare via the shell
instead.
- kadm/oob-drift: assert bug. awk /retention.ms/ also matches
retention.local.target.ms, so the check compared a two-line
string against 3600000. The controller had in fact restored the
value. Exact-match awk.
- kadm/parameter-mutation: same awk bug, plus the broker check was
a single read immediately after Ready; the alter is acked before
the config is necessarily visible to DescribeConfigs. Exact-match
awk and a 60s poll like the oob-drift scenario already uses.
Release pin bumped for the status stamping change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
kadm/retention-policy flaked in one of two runs of the same sha:
topic creation is acked before a fresh rpk metadata request
necessarily lists the topic, so single-shot existence checks race.
kafka_topic_exists now polls (6x5s), and the deleted-topic check
moves into a matching kafka_topic_absent helper.
The same-sha double run came from push + pull_request both
triggering on PR branches; push is now main-only.
No binary change; release pin untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@solsson
solsson merged commit 79c8441 into mainJul 13, 2026
2 checks passed
@solssonsolsson mentioned this pull request Jul 13, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@solsson
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

e2e failure propagation - #1

Merged
solsson merged 4 commits into
mainfrom
e2e-failure-propagation
Jul 13, 2026
Merged

e2e failure propagation#1
solsson merged 4 commits into
mainfrom
e2e-failure-propagation

Conversation

@solsson

Copy link
Copy Markdown
Contributor

No description provided.

Yolean k8s-qa (buckety maintainer)and others added 4 commits July 10, 2026 06:24
run_scenario is invoked in an if condition, where bash ignores
set -e inside the function body. Every failure inside it, including
a non-zero assert.sh, fell through to the PASS logging and returned
0, so the harness has never been able to report FAIL. The latest
green main run (27607054780) shows six scenario executions with
explicit [assert][FAIL] or wait timeouts, all recorded PASS. Each
step now propagates failure explicitly with || return 1.
Three more defects the same run exposed:
- kafka_topic_exists/s3_bucket_exists use kubectl run -i, which
forwards and consumes the caller's stdin. The harness fed its
scenario list on stdin of a while read loop, so the first assert
that ran an -i helper swallowed the remaining list: only 4 of 11
kadm scenarios ever executed. The list is now materialised with
mapfile and assert.sh runs with stdin from /dev/null.
- The helpers piped kubectl run into grep -q under pipefail, where
an early grep exit SIGPIPEs kubectl and fails the pipeline even
on a present topic, with stderr discarded. Output is now captured
and printed on failure so assertion failures are diagnosable.
- teardown between implementations deleted the CRDs while scenario
CRs with finalizers could still be terminating, leaving the CRD
stuck and the next implementation's applies rejected with "create
not allowed while custom resource definition is terminating"
(the entire versitygw block in that run). Overlays now converge
in place; nothing deletes CRDs mid-run.
Scenario-level fixes entailed by unmasking:
- driver-version defaulted E2E_IMAGE_* to a ghcr :dev tag that is
never published, wedging the Deployment for two rollout timeouts
per run. The vars are now pass-through inputs and the scenario
logs SKIPPED when they are unset; CI provides real rotation
images (separate commit).
- oob-drift waits 90-300s for periodic drift repair, but the
controller default recheck is 5m. e2e overlays now pass
--periodic-recheck=15s.
Failed scenarios leave their namespace standing per SPEC "E2E
harness and parity" #3; KEEP_FAILED now additionally preserves
namespaces of passing scenarios.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…n CI
The reconciler gated the driver-major comparison on
status.driverMajor != 0, treating 0 as "not yet stamped". Both
shipped drivers are 0.x, so every existing resource is stamped
major=0 and the check never ran: a 0.x to 1.0 driver bump would
silently auto-apply instead of surfacing DriverVersionIncompatible,
contradicting SPEC "Driver versioning". Stampedness is now
signalled by status.backend, which is set in the same first
reconcile pass.
The driver-version e2e scenario (SPEC scenario #9) needs controller
images built with rotated driver versions. CI never built them, so
the scenario has never actually run. The workflow now builds
patch (0.1.1) and major (1.0.0) driver builds per SCAFFOLDING
"Driver version stamping" and pushes them to the run's local
registry; the harness receives them via E2E_IMAGE_*/E2E_VERSION_*.
These images are rotation fixtures only, not digest-pinned and
not published.
Release pin bumped for the reconciler change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
First honest run (29073773069): 15 PASS, 6 FAIL. The six, with
root causes:
- driver-version: status.driverMajor read back empty. The field
was int with omitempty and the stamp used a merge patch, so a
stamped 0 (any 0.x driver) neither serializes nor diffs and can
never reach the API server. omitempty dropped; the first-reconcile
stamp is a Status().Update() since a merge diff of 0 -> 0 is
still empty.
- misconfigured-startup: the scenario is config-only
(resources: []) and plain kubectl apply -k errors with "no
objects passed to apply". The harness now renders first and
skips apply when the output is empty.
- s3/happy-path on versitygw AND minio: the roundtrip Job upload
and download both succeed, then the final comparison calls diff,
which does not exist in the aws-cli image. Compare via the shell
instead.
- kadm/oob-drift: assert bug. awk /retention.ms/ also matches
retention.local.target.ms, so the check compared a two-line
string against 3600000. The controller had in fact restored the
value. Exact-match awk.
- kadm/parameter-mutation: same awk bug, plus the broker check was
a single read immediately after Ready; the alter is acked before
the config is necessarily visible to DescribeConfigs. Exact-match
awk and a 60s poll like the oob-drift scenario already uses.
Release pin bumped for the status stamping change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
kadm/retention-policy flaked in one of two runs of the same sha:
topic creation is acked before a fresh rpk metadata request
necessarily lists the topic, so single-shot existence checks race.
kafka_topic_exists now polls (6x5s), and the deleted-topic check
moves into a matching kafka_topic_absent helper.
The same-sha double run came from push + pull_request both
triggering on PR branches; push is now main-only.
No binary change; release pin untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@solsson
solsson merged commit 79c8441 into mainJul 13, 2026
2 checks passed
@solssonsolsson mentioned this pull request Jul 13, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@solsson
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

e2e failure propagation - #1

Merged
solsson merged 4 commits into
mainfrom
e2e-failure-propagation
Jul 13, 2026
Merged

e2e failure propagation#1
solsson merged 4 commits into
mainfrom
e2e-failure-propagation

Conversation

@solsson

Copy link
Copy Markdown
Contributor

No description provided.

Yolean k8s-qa (buckety maintainer)and others added 4 commits July 10, 2026 06:24
run_scenario is invoked in an if condition, where bash ignores
set -e inside the function body. Every failure inside it, including
a non-zero assert.sh, fell through to the PASS logging and returned
0, so the harness has never been able to report FAIL. The latest
green main run (27607054780) shows six scenario executions with
explicit [assert][FAIL] or wait timeouts, all recorded PASS. Each
step now propagates failure explicitly with || return 1.
Three more defects the same run exposed:
- kafka_topic_exists/s3_bucket_exists use kubectl run -i, which
forwards and consumes the caller's stdin. The harness fed its
scenario list on stdin of a while read loop, so the first assert
that ran an -i helper swallowed the remaining list: only 4 of 11
kadm scenarios ever executed. The list is now materialised with
mapfile and assert.sh runs with stdin from /dev/null.
- The helpers piped kubectl run into grep -q under pipefail, where
an early grep exit SIGPIPEs kubectl and fails the pipeline even
on a present topic, with stderr discarded. Output is now captured
and printed on failure so assertion failures are diagnosable.
- teardown between implementations deleted the CRDs while scenario
CRs with finalizers could still be terminating, leaving the CRD
stuck and the next implementation's applies rejected with "create
not allowed while custom resource definition is terminating"
(the entire versitygw block in that run). Overlays now converge
in place; nothing deletes CRDs mid-run.
Scenario-level fixes entailed by unmasking:
- driver-version defaulted E2E_IMAGE_* to a ghcr :dev tag that is
never published, wedging the Deployment for two rollout timeouts
per run. The vars are now pass-through inputs and the scenario
logs SKIPPED when they are unset; CI provides real rotation
images (separate commit).
- oob-drift waits 90-300s for periodic drift repair, but the
controller default recheck is 5m. e2e overlays now pass
--periodic-recheck=15s.
Failed scenarios leave their namespace standing per SPEC "E2E
harness and parity" #3; KEEP_FAILED now additionally preserves
namespaces of passing scenarios.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…n CI
The reconciler gated the driver-major comparison on
status.driverMajor != 0, treating 0 as "not yet stamped". Both
shipped drivers are 0.x, so every existing resource is stamped
major=0 and the check never ran: a 0.x to 1.0 driver bump would
silently auto-apply instead of surfacing DriverVersionIncompatible,
contradicting SPEC "Driver versioning". Stampedness is now
signalled by status.backend, which is set in the same first
reconcile pass.
The driver-version e2e scenario (SPEC scenario #9) needs controller
images built with rotated driver versions. CI never built them, so
the scenario has never actually run. The workflow now builds
patch (0.1.1) and major (1.0.0) driver builds per SCAFFOLDING
"Driver version stamping" and pushes them to the run's local
registry; the harness receives them via E2E_IMAGE_*/E2E_VERSION_*.
These images are rotation fixtures only, not digest-pinned and
not published.
Release pin bumped for the reconciler change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
First honest run (29073773069): 15 PASS, 6 FAIL. The six, with
root causes:
- driver-version: status.driverMajor read back empty. The field
was int with omitempty and the stamp used a merge patch, so a
stamped 0 (any 0.x driver) neither serializes nor diffs and can
never reach the API server. omitempty dropped; the first-reconcile
stamp is a Status().Update() since a merge diff of 0 -> 0 is
still empty.
- misconfigured-startup: the scenario is config-only
(resources: []) and plain kubectl apply -k errors with "no
objects passed to apply". The harness now renders first and
skips apply when the output is empty.
- s3/happy-path on versitygw AND minio: the roundtrip Job upload
and download both succeed, then the final comparison calls diff,
which does not exist in the aws-cli image. Compare via the shell
instead.
- kadm/oob-drift: assert bug. awk /retention.ms/ also matches
retention.local.target.ms, so the check compared a two-line
string against 3600000. The controller had in fact restored the
value. Exact-match awk.
- kadm/parameter-mutation: same awk bug, plus the broker check was
a single read immediately after Ready; the alter is acked before
the config is necessarily visible to DescribeConfigs. Exact-match
awk and a 60s poll like the oob-drift scenario already uses.
Release pin bumped for the status stamping change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
kadm/retention-policy flaked in one of two runs of the same sha:
topic creation is acked before a fresh rpk metadata request
necessarily lists the topic, so single-shot existence checks race.
kafka_topic_exists now polls (6x5s), and the deleted-topic check
moves into a matching kafka_topic_absent helper.
The same-sha double run came from push + pull_request both
triggering on PR branches; push is now main-only.
No binary change; release pin untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@solsson
solsson merged commit 79c8441 into mainJul 13, 2026
2 checks passed
@solssonsolsson mentioned this pull request Jul 13, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@solsson
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

e2e failure propagation - #1

Merged
solsson merged 4 commits into
mainfrom
e2e-failure-propagation
Jul 13, 2026
Merged

e2e failure propagation#1
solsson merged 4 commits into
mainfrom
e2e-failure-propagation

Conversation

@solsson

Copy link
Copy Markdown
Contributor

No description provided.

Yolean k8s-qa (buckety maintainer)and others added 4 commits July 10, 2026 06:24
run_scenario is invoked in an if condition, where bash ignores
set -e inside the function body. Every failure inside it, including
a non-zero assert.sh, fell through to the PASS logging and returned
0, so the harness has never been able to report FAIL. The latest
green main run (27607054780) shows six scenario executions with
explicit [assert][FAIL] or wait timeouts, all recorded PASS. Each
step now propagates failure explicitly with || return 1.
Three more defects the same run exposed:
- kafka_topic_exists/s3_bucket_exists use kubectl run -i, which
forwards and consumes the caller's stdin. The harness fed its
scenario list on stdin of a while read loop, so the first assert
that ran an -i helper swallowed the remaining list: only 4 of 11
kadm scenarios ever executed. The list is now materialised with
mapfile and assert.sh runs with stdin from /dev/null.
- The helpers piped kubectl run into grep -q under pipefail, where
an early grep exit SIGPIPEs kubectl and fails the pipeline even
on a present topic, with stderr discarded. Output is now captured
and printed on failure so assertion failures are diagnosable.
- teardown between implementations deleted the CRDs while scenario
CRs with finalizers could still be terminating, leaving the CRD
stuck and the next implementation's applies rejected with "create
not allowed while custom resource definition is terminating"
(the entire versitygw block in that run). Overlays now converge
in place; nothing deletes CRDs mid-run.
Scenario-level fixes entailed by unmasking:
- driver-version defaulted E2E_IMAGE_* to a ghcr :dev tag that is
never published, wedging the Deployment for two rollout timeouts
per run. The vars are now pass-through inputs and the scenario
logs SKIPPED when they are unset; CI provides real rotation
images (separate commit).
- oob-drift waits 90-300s for periodic drift repair, but the
controller default recheck is 5m. e2e overlays now pass
--periodic-recheck=15s.
Failed scenarios leave their namespace standing per SPEC "E2E
harness and parity" #3; KEEP_FAILED now additionally preserves
namespaces of passing scenarios.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…n CI
The reconciler gated the driver-major comparison on
status.driverMajor != 0, treating 0 as "not yet stamped". Both
shipped drivers are 0.x, so every existing resource is stamped
major=0 and the check never ran: a 0.x to 1.0 driver bump would
silently auto-apply instead of surfacing DriverVersionIncompatible,
contradicting SPEC "Driver versioning". Stampedness is now
signalled by status.backend, which is set in the same first
reconcile pass.
The driver-version e2e scenario (SPEC scenario #9) needs controller
images built with rotated driver versions. CI never built them, so
the scenario has never actually run. The workflow now builds
patch (0.1.1) and major (1.0.0) driver builds per SCAFFOLDING
"Driver version stamping" and pushes them to the run's local
registry; the harness receives them via E2E_IMAGE_*/E2E_VERSION_*.
These images are rotation fixtures only, not digest-pinned and
not published.
Release pin bumped for the reconciler change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
First honest run (29073773069): 15 PASS, 6 FAIL. The six, with
root causes:
- driver-version: status.driverMajor read back empty. The field
was int with omitempty and the stamp used a merge patch, so a
stamped 0 (any 0.x driver) neither serializes nor diffs and can
never reach the API server. omitempty dropped; the first-reconcile
stamp is a Status().Update() since a merge diff of 0 -> 0 is
still empty.
- misconfigured-startup: the scenario is config-only
(resources: []) and plain kubectl apply -k errors with "no
objects passed to apply". The harness now renders first and
skips apply when the output is empty.
- s3/happy-path on versitygw AND minio: the roundtrip Job upload
and download both succeed, then the final comparison calls diff,
which does not exist in the aws-cli image. Compare via the shell
instead.
- kadm/oob-drift: assert bug. awk /retention.ms/ also matches
retention.local.target.ms, so the check compared a two-line
string against 3600000. The controller had in fact restored the
value. Exact-match awk.
- kadm/parameter-mutation: same awk bug, plus the broker check was
a single read immediately after Ready; the alter is acked before
the config is necessarily visible to DescribeConfigs. Exact-match
awk and a 60s poll like the oob-drift scenario already uses.
Release pin bumped for the status stamping change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
kadm/retention-policy flaked in one of two runs of the same sha:
topic creation is acked before a fresh rpk metadata request
necessarily lists the topic, so single-shot existence checks race.
kafka_topic_exists now polls (6x5s), and the deleted-topic check
moves into a matching kafka_topic_absent helper.
The same-sha double run came from push + pull_request both
triggering on PR branches; push is now main-only.
No binary change; release pin untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@solsson
solsson merged commit 79c8441 into mainJul 13, 2026
2 checks passed
@solssonsolsson mentioned this pull request Jul 13, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@solsson
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

e2e failure propagation - #1

Merged
solsson merged 4 commits into
mainfrom
e2e-failure-propagation
Jul 13, 2026
Merged

e2e failure propagation#1
solsson merged 4 commits into
mainfrom
e2e-failure-propagation

Conversation

@solsson

Copy link
Copy Markdown
Contributor

No description provided.

Yolean k8s-qa (buckety maintainer)and others added 4 commits July 10, 2026 06:24
run_scenario is invoked in an if condition, where bash ignores
set -e inside the function body. Every failure inside it, including
a non-zero assert.sh, fell through to the PASS logging and returned
0, so the harness has never been able to report FAIL. The latest
green main run (27607054780) shows six scenario executions with
explicit [assert][FAIL] or wait timeouts, all recorded PASS. Each
step now propagates failure explicitly with || return 1.
Three more defects the same run exposed:
- kafka_topic_exists/s3_bucket_exists use kubectl run -i, which
forwards and consumes the caller's stdin. The harness fed its
scenario list on stdin of a while read loop, so the first assert
that ran an -i helper swallowed the remaining list: only 4 of 11
kadm scenarios ever executed. The list is now materialised with
mapfile and assert.sh runs with stdin from /dev/null.
- The helpers piped kubectl run into grep -q under pipefail, where
an early grep exit SIGPIPEs kubectl and fails the pipeline even
on a present topic, with stderr discarded. Output is now captured
and printed on failure so assertion failures are diagnosable.
- teardown between implementations deleted the CRDs while scenario
CRs with finalizers could still be terminating, leaving the CRD
stuck and the next implementation's applies rejected with "create
not allowed while custom resource definition is terminating"
(the entire versitygw block in that run). Overlays now converge
in place; nothing deletes CRDs mid-run.
Scenario-level fixes entailed by unmasking:
- driver-version defaulted E2E_IMAGE_* to a ghcr :dev tag that is
never published, wedging the Deployment for two rollout timeouts
per run. The vars are now pass-through inputs and the scenario
logs SKIPPED when they are unset; CI provides real rotation
images (separate commit).
- oob-drift waits 90-300s for periodic drift repair, but the
controller default recheck is 5m. e2e overlays now pass
--periodic-recheck=15s.
Failed scenarios leave their namespace standing per SPEC "E2E
harness and parity" #3; KEEP_FAILED now additionally preserves
namespaces of passing scenarios.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…n CI
The reconciler gated the driver-major comparison on
status.driverMajor != 0, treating 0 as "not yet stamped". Both
shipped drivers are 0.x, so every existing resource is stamped
major=0 and the check never ran: a 0.x to 1.0 driver bump would
silently auto-apply instead of surfacing DriverVersionIncompatible,
contradicting SPEC "Driver versioning". Stampedness is now
signalled by status.backend, which is set in the same first
reconcile pass.
The driver-version e2e scenario (SPEC scenario #9) needs controller
images built with rotated driver versions. CI never built them, so
the scenario has never actually run. The workflow now builds
patch (0.1.1) and major (1.0.0) driver builds per SCAFFOLDING
"Driver version stamping" and pushes them to the run's local
registry; the harness receives them via E2E_IMAGE_*/E2E_VERSION_*.
These images are rotation fixtures only, not digest-pinned and
not published.
Release pin bumped for the reconciler change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
First honest run (29073773069): 15 PASS, 6 FAIL. The six, with
root causes:
- driver-version: status.driverMajor read back empty. The field
was int with omitempty and the stamp used a merge patch, so a
stamped 0 (any 0.x driver) neither serializes nor diffs and can
never reach the API server. omitempty dropped; the first-reconcile
stamp is a Status().Update() since a merge diff of 0 -> 0 is
still empty.
- misconfigured-startup: the scenario is config-only
(resources: []) and plain kubectl apply -k errors with "no
objects passed to apply". The harness now renders first and
skips apply when the output is empty.
- s3/happy-path on versitygw AND minio: the roundtrip Job upload
and download both succeed, then the final comparison calls diff,
which does not exist in the aws-cli image. Compare via the shell
instead.
- kadm/oob-drift: assert bug. awk /retention.ms/ also matches
retention.local.target.ms, so the check compared a two-line
string against 3600000. The controller had in fact restored the
value. Exact-match awk.
- kadm/parameter-mutation: same awk bug, plus the broker check was
a single read immediately after Ready; the alter is acked before
the config is necessarily visible to DescribeConfigs. Exact-match
awk and a 60s poll like the oob-drift scenario already uses.
Release pin bumped for the status stamping change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
kadm/retention-policy flaked in one of two runs of the same sha:
topic creation is acked before a fresh rpk metadata request
necessarily lists the topic, so single-shot existence checks race.
kafka_topic_exists now polls (6x5s), and the deleted-topic check
moves into a matching kafka_topic_absent helper.
The same-sha double run came from push + pull_request both
triggering on PR branches; push is now main-only.
No binary change; release pin untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@solsson
solsson merged commit 79c8441 into mainJul 13, 2026
2 checks passed
@solssonsolsson mentioned this pull request Jul 13, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@solsson
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

e2e failure propagation - #1

Merged
solsson merged 4 commits into
mainfrom
e2e-failure-propagation
Jul 13, 2026
Merged

e2e failure propagation#1
solsson merged 4 commits into
mainfrom
e2e-failure-propagation

Conversation

@solsson

Copy link
Copy Markdown
Contributor

No description provided.

Yolean k8s-qa (buckety maintainer)and others added 4 commits July 10, 2026 06:24
run_scenario is invoked in an if condition, where bash ignores
set -e inside the function body. Every failure inside it, including
a non-zero assert.sh, fell through to the PASS logging and returned
0, so the harness has never been able to report FAIL. The latest
green main run (27607054780) shows six scenario executions with
explicit [assert][FAIL] or wait timeouts, all recorded PASS. Each
step now propagates failure explicitly with || return 1.
Three more defects the same run exposed:
- kafka_topic_exists/s3_bucket_exists use kubectl run -i, which
forwards and consumes the caller's stdin. The harness fed its
scenario list on stdin of a while read loop, so the first assert
that ran an -i helper swallowed the remaining list: only 4 of 11
kadm scenarios ever executed. The list is now materialised with
mapfile and assert.sh runs with stdin from /dev/null.
- The helpers piped kubectl run into grep -q under pipefail, where
an early grep exit SIGPIPEs kubectl and fails the pipeline even
on a present topic, with stderr discarded. Output is now captured
and printed on failure so assertion failures are diagnosable.
- teardown between implementations deleted the CRDs while scenario
CRs with finalizers could still be terminating, leaving the CRD
stuck and the next implementation's applies rejected with "create
not allowed while custom resource definition is terminating"
(the entire versitygw block in that run). Overlays now converge
in place; nothing deletes CRDs mid-run.
Scenario-level fixes entailed by unmasking:
- driver-version defaulted E2E_IMAGE_* to a ghcr :dev tag that is
never published, wedging the Deployment for two rollout timeouts
per run. The vars are now pass-through inputs and the scenario
logs SKIPPED when they are unset; CI provides real rotation
images (separate commit).
- oob-drift waits 90-300s for periodic drift repair, but the
controller default recheck is 5m. e2e overlays now pass
--periodic-recheck=15s.
Failed scenarios leave their namespace standing per SPEC "E2E
harness and parity" #3; KEEP_FAILED now additionally preserves
namespaces of passing scenarios.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…n CI
The reconciler gated the driver-major comparison on
status.driverMajor != 0, treating 0 as "not yet stamped". Both
shipped drivers are 0.x, so every existing resource is stamped
major=0 and the check never ran: a 0.x to 1.0 driver bump would
silently auto-apply instead of surfacing DriverVersionIncompatible,
contradicting SPEC "Driver versioning". Stampedness is now
signalled by status.backend, which is set in the same first
reconcile pass.
The driver-version e2e scenario (SPEC scenario #9) needs controller
images built with rotated driver versions. CI never built them, so
the scenario has never actually run. The workflow now builds
patch (0.1.1) and major (1.0.0) driver builds per SCAFFOLDING
"Driver version stamping" and pushes them to the run's local
registry; the harness receives them via E2E_IMAGE_*/E2E_VERSION_*.
These images are rotation fixtures only, not digest-pinned and
not published.
Release pin bumped for the reconciler change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
First honest run (29073773069): 15 PASS, 6 FAIL. The six, with
root causes:
- driver-version: status.driverMajor read back empty. The field
was int with omitempty and the stamp used a merge patch, so a
stamped 0 (any 0.x driver) neither serializes nor diffs and can
never reach the API server. omitempty dropped; the first-reconcile
stamp is a Status().Update() since a merge diff of 0 -> 0 is
still empty.
- misconfigured-startup: the scenario is config-only
(resources: []) and plain kubectl apply -k errors with "no
objects passed to apply". The harness now renders first and
skips apply when the output is empty.
- s3/happy-path on versitygw AND minio: the roundtrip Job upload
and download both succeed, then the final comparison calls diff,
which does not exist in the aws-cli image. Compare via the shell
instead.
- kadm/oob-drift: assert bug. awk /retention.ms/ also matches
retention.local.target.ms, so the check compared a two-line
string against 3600000. The controller had in fact restored the
value. Exact-match awk.
- kadm/parameter-mutation: same awk bug, plus the broker check was
a single read immediately after Ready; the alter is acked before
the config is necessarily visible to DescribeConfigs. Exact-match
awk and a 60s poll like the oob-drift scenario already uses.
Release pin bumped for the status stamping change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
kadm/retention-policy flaked in one of two runs of the same sha:
topic creation is acked before a fresh rpk metadata request
necessarily lists the topic, so single-shot existence checks race.
kafka_topic_exists now polls (6x5s), and the deleted-topic check
moves into a matching kafka_topic_absent helper.
The same-sha double run came from push + pull_request both
triggering on PR branches; push is now main-only.
No binary change; release pin untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@solsson
solsson merged commit 79c8441 into mainJul 13, 2026
2 checks passed
@solssonsolsson mentioned this pull request Jul 13, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@solsson