Uh oh!
There was an error while loading. Please reload this page.
[fix](fe) Retry meta service too busy responses - #65378
Conversation
### What problem does this PR solve? Issue Number: None Related PR: None Problem Summary: MetaServiceProxy only retried transport failures. When a meta service RPC returned successfully but carried MS_TOO_BUSY in the response status, FE returned that response immediately and let callers handle a transient overload as a normal business result. This change passes a typed status extractor into executeRequest, retries MS_TOO_BUSY responses with the existing retry counter and backoff, and throws RpcException if all attempts still return MS_TOO_BUSY. The retry path does not mark the client as failed, so it does not trigger connection shutdown. ### Release note None ### Check List (For Author) - Test: Not run (user requested not to compile or run tests) - Behavior changed: Yes (MS_TOO_BUSY meta service responses are retried by FE MetaServiceProxy and fail after retries are exhausted) - Does this need documentation: No
hello-stephen
commented
Jul 8, 2026
Thank you for your contribution to Apache Doris. Please clearly describe your PR:
|
mymeiyi
commented
Jul 8, 2026
/review |
There was a problem hiding this comment.
Pull request overview
This PR improves FE’s cloud MetaService RPC resilience by treating MS_TOO_BUSY as a transient overload signal (even when the RPC itself succeeds), retrying with the existing retry/backoff loop, and surfacing an RpcException only after retries are exhausted.
Changes:
- Extend
MetaServiceClientWrapper.executeRequest/executeWithMetricsto accept a typedstatusExtractorand retry responses whose status code isMS_TOO_BUSY. - Update all
MetaServiceProxyRPC wrappers to pass the correctgetStatusextractor. - Add FE unit tests validating: no shutdown on
MS_TOO_BUSY, retry-to-success, and failure after retries.
Review Checkpoint Conclusions (from skill checklist)
- Goal & completeness: The change matches the PR description (retry
MS_TOO_BUSYresponses) and adds targeted unit tests demonstrating the new behavior. - Scope: Focused to MetaServiceProxy retry behavior and tests; no unrelated refactors observed.
- Concurrency/lifecycle: No new concurrency primitives introduced; retry loop behavior remains synchronous and uses existing
Thread.sleepbackoff.
Reviewed changes
Copilot reviewed 2 out of 2 changed files in this pull request and generated 3 comments.
| File | Description |
|---|---|
fe/fe-core/src/main/java/org/apache/doris/cloud/rpc/MetaServiceProxy.java | Adds status-based retry for MS_TOO_BUSY by passing a response-status extractor through the existing retry/metrics wrapper. |
fe/fe-core/src/test/java/org/apache/doris/cloud/rpc/MetaServiceProxyTest.java | Updates existing tests for new API signature and adds coverage for MS_TOO_BUSY retry/failure behavior. |
Comments suppressed due to low confidence (1)
fe/fe-core/src/main/java/org/apache/doris/cloud/rpc/MetaServiceProxy.java:249
- Log message contains a typo: "meta servive" → "meta service".
requestFailed = true;
LOG.warn("failed to request meta servive trycnt {}", tried, e);
if (tried >= maxRetries) {
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
There was a problem hiding this comment.
Review complete. I found one correctness gap: the new MS_TOO_BUSY response retry is wired through the shared synchronous wrapper, but MetaServiceProxy.getInstance() still bypasses that wrapper and can return a status-bearing too-busy response directly. I left one inline comment.
Critical checkpoint conclusions:
- Goal/test: wrapper-based synchronous RPCs now retry MS_TOO_BUSY and have focused unit coverage, but getInstance is a parallel status-bearing RPC path without the new retry behavior or a test.
- Scope: the change is mostly focused and does not add config, storage format, or protocol fields.
- Concurrency/lifecycle: no new shared mutable state or lock ordering change found; the retry path intentionally avoids marking MS_TOO_BUSY as a client failure.
- Compatibility: mixed FE/meta-service compatibility is not materially changed because MS_TOO_BUSY already exists in the proto.
- Parallel paths: getInstance is the uncovered parallel path and is the blocking issue.
- Testing/style: static checks found all new status extractors refer to response types with a status field and no changed Java lines over 120 columns. I did not run FE tests because this checkout is missing .worktree_initialized, thirdparty/installed, and thirdparty/installed/bin/protoc.
User focus: no additional user-provided focus points were present.
Subagent conclusions: optimizer-rewrite proposed OPT-1, which became MAIN-1 and is submitted inline. tests-session-config proposed TSC-1 as a duplicate of OPT-1 with test coverage evidence, merged into MAIN-1. Convergence round 1 ended with both live subagents replying NO_NEW_VALUABLE_FINDINGS for the same ledger/comment set.
Uh oh!
There was an error while loading. Please reload this page.
mymeiyi
commented
Jul 8, 2026
run buildall |
hello-stephen
commented
Jul 8, 2026
FE Regression Coverage ReportIncrement line coverage |
PR approved by at least one committer and no changes requested. |
mymeiyi
commented
Jul 13, 2026
run external |
mymeiyi
commented
Jul 13, 2026
run performance |
hello-stephen
commented
Jul 13, 2026
TPC-H: Total hot run time: 29542 ms |
hello-stephen
commented
Jul 13, 2026
TPC-DS: Total hot run time: 182134 ms |
hello-stephen
commented
Jul 13, 2026
ClickBench: Total hot run time: 24.95 s |
Uh oh!
There was an error while loading. Please reload this page.
When a meta service RPC returned successfully but carried MS_TOO_BUSY in the response status, FE returned that response immediately and let callers handle a transient overload as a normal business result.
This change passes a typed status extractor into executeRequest, retries MS_TOO_BUSY responses with the existing retry counter and backoff, and throws RpcException if all attempts still return MS_TOO_BUSY.