fix(collector): fail stuck recovery running jobs instead of succeeding - #1157
Open
proerror77 wants to merge 1 commit into
Open
fix(collector): fail stuck recovery running jobs instead of succeeding#1157proerror77 wants to merge 1 commit into
proerror77 wants to merge 1 commit into
Conversation
proerror77
marked this pull request as ready for review
September 12, 2026 04:47
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
cursor
Bot
force-pushed
the
cursor/collector-recovery-running-timeout-1729
branch
from
September 12, 2026 05:00
3dda17b to
28b8089
Compare
Recovery oneshot TimeoutStartSec is now 7200s so a hung drain receives TERM. An unfinished *.running job is marked failed rather than leaving systemd Result=success. Failed jobs still require an identity-bearing resume; this does not auto-retry them. Co-authored-by: Wild Card <proerror77@users.noreply.github.com>
cursor
Bot
force-pushed
the
cursor/collector-recovery-running-timeout-1729
branch
from
September 12, 2026 05:50
28b8089 to
e86f5b2
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Change
Stuck LOB recovery
*.runningjobs no longer look successful.TimeoutStartSec=0let systemd report success while an unfinished drain sat in.running. The oneshot now bounds start at 7200s and stop at 120s. Drain of an unfinished running jobmark_faileds withstep=unfinished-runningand exits 1. Stale identity stillmark_stales viafinalize_passed_running. Failed jobs still require an identity-bearingresume; this does not auto-resume them.Rebased onto current
mainafter#1154.Issue relationship
None
Validation
git diff --checkdeployment/aliyun/test-rust-lob-recovery-queue.sh— unfinished running drain fails closed into.failedwithstep=unfinished-running; unit file rejectsTimeoutStartSec=0.Runtime impact and rollback
Collector recovery only. After the next controller/runtime deploy of
host-rust-lob-recovery-queue.shandbinance-lob-archiver-recovery@.service, stuck running jobs fail instead of succeeding. Rollback is the previous recovery controller and unit. This is not a live trading change.