Uh oh!
There was an error while loading. Please reload this page.
ci: soft-fail coverage-gate while artifact quota persists - #41
Conversation
After PR #40, the workflow conclusion is still 'failure' because the coverage-gate job fails when its first download-artifact (coverage-unit) can't find an artifact — that artifact's upload is currently soft-failed on the rust-unit job due to the GH Actions storage quota outage. The coverage-gate threshold check itself is already a soft gate (the bash exits 0 even on below-threshold; see comment on the merge step). Align the job-level setting with that intent: continue-on-error so a failure doesn't make the whole workflow red. Also mark the first download-artifact step continue-on-error so missing input doesn't prevent the threshold step (which already tolerates an empty cov/e2e directory). Revert with the other quota mitigations once GH storage usage refreshes or the quota is raised.
There was a problem hiding this comment.
Pull request overview
Adjusts CI’s advisory coverage gate to avoid failing the overall workflow while GitHub Actions artifact uploads/downloads are intermittently failing due to the org storage quota outage.
Changes:
- Marks the
coverage-gatejob ascontinue-on-error: trueso it won’t fail the workflow conclusion. - Soft-fails the
coverage-unitartifact download step so the job can proceed even when that artifact is missing.
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| # workflow when coverage-gate fails — especially while the artifact | ||
| # storage quota outage prevents the upstream coverage-unit upload. | ||
| # Revert when the gate is flipped from advisory to gating (PR #5). | ||
| continue-on-error: true |
There was a problem hiding this comment.
The new comment says to “Revert when the gate is flipped … (PR #5)”, but PR #5 (per this PR’s metadata) is about server startup/health listeners and doesn’t appear related to coverage gating. Consider removing the PR reference or linking to the correct tracking issue/PR so future readers know when to revert.
| - uses: actions/download-artifact@v4 | ||
| continue-on-error: true | ||
| with: { name: coverage-unit, path: cov/unit } |
There was a problem hiding this comment.
continue-on-error on the download-artifact step lets the job proceed, but the later merge script still assumes cov/unit/*.info exists (files=(cov/unit/*.info) + cp "${files[0]}" ...). When the unit artifact is missing, Bash will keep the literal glob, causing lcov-result-merger/cp to fail and the subsequent upload step to be skipped. Please make the merge step tolerate missing/empty artifacts (e.g., enable nullglob and build the files array only from existing matches, or explicitly handle the empty case).
Uh oh!
There was an error while loading. Please reload this page.
…42) The Packages spending limit on the moonming account was raised from $0 to $5/mo (artifact storage is billed under the Packages SKU, not Actions). Verified by re-running run 24948385951: build-ui's upload-artifact step no longer reports "Artifact storage quota has been hit", and downstream build-aisix can now download ui-dist. Reverts the four `continue-on-error` mitigations added by: - #40: rust-unit's coverage-unit upload-artifact build-ui's ui-dist upload-artifact build-bin (build-aisix) job-level - #41: coverage-gate job-level coverage-gate's coverage-unit download-artifact Preserves the two pre-existing `continue-on-error` markers that predate the quota outage: - e2e job-level (etcd service startup / undici flake) - coverage-gate's coverage-e2e download-artifact (e2e produces no coverage when its tests don't run)
Summary
main CI is currently red on commit 6bf6227 because the `coverage-gate` job fails when its first `download-artifact` (coverage-unit) can't find an artifact. That artifact's upload was already soft-failed on the rust-unit job by PR #40 due to the GitHub Actions storage quota outage. The job-level `continue-on-error` was missed on coverage-gate so the workflow conclusion still rolls up to failure.
The coverage-gate threshold check is already a soft gate (the bash explicitly `exit 0`s even when below threshold; the comment says "soft during scaffold milestones; flip to `exit 1` at PR #5"). Aligning the job-level setting with that documented intent.
Changes
Test plan