feat(routing): fallback_on_statuses — opt selected upstream statuses into retry/failover - #752

Merged
jarvis9443 merged 2 commits into
mainfrom
feat/fallback-on-statuses
Jul 10, 2026
Merged

feat(routing): fallback_on_statuses — opt selected upstream statuses into retry/failover#752
jarvis9443 merged 2 commits into
mainfrom
feat/fallback-on-statuses

Conversation

@jarvis9443

@jarvis9443jarvis9443 commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

Problem

Some providers use non-429 4xx codes for transient conditions — model overload, queue full, quota exhaustion. The default classification (is_retryable) treats every non-429 4xx as the caller's error and relays it, so a model group fails a request that its other targets could have served. (AISIX-Cloud#1012)

Change

Routing models gain an explicit, empty-by-default opt-in list:

routing:
strategy: failovertargets: [{ model: primary }, { model: backup }]fallback_on_statuses: [408, 409, 422]

A listed status becomes retryable in is_retryable() — the single predicate every dispatch loop consults — so it participates in both same-target retries (retries) and failover (max_fallbacks), exactly like retry_on_429 does for 429 (429 in the list works without retry_on_429; the boolean stays for compatibility). Unset = behavior unchanged, bit-for-bit. 5xx are already retryable (listing them is a no-op); the list never affects non-status failures (customer-fixable config/credential errors stay terminal). Schema constrains entries to 400–599.

Threaded through all six dispatch sites: chat streaming + non-streaming + cross-target break, messages, responses, count_tokens — the whole routing family in one PR.

Observability: per-attempt telemetry already records each attempt's upstream status, error_class and attempt_kind (initial/retry/fallback), so a rule-triggered failover is fully attributable from existing fields; no new telemetry needed.

Interaction note: this is the current-request knob. The pre-existing per-direct-model cooldown.trigger_statuses benches a target for future requests and stays an independent layer.

Deliberately not included

  • Provider error-code / response-body matchers (the issue floats error.code: Throttling-style matching as a possible extension): status-code opt-in covers the reported scenarios; body matching adds a parsing surface per provider and can be layered on later if a concrete provider needs it.
  • Changing defaults (e.g. making 408/409 retryable out of the box, as some gateways do): the issue explicitly requires unchanged defaults.

Baseline: mainstream gateways retry 408/409/429/5xx by default and let users force retries per exception category; none offer per-status opt-in. We keep stricter defaults and add the more precise status-list control instead — providers in scope here signal overload with codes (e.g. 422) that category-based systems can't distinguish from validation errors.

Control-plane pairing

The CP is a closed schema — this field is unreachable by users until the paired AISIX-Cloud PR (cp-admin.yaml + validation + etcd projection + dashboard) lands. DP ships first because it rejects unknown fields (deny_unknown_fields); the CP PR follows immediately and carries the closing reference for AISIX-Cloud#1012.

Tests

  • Unit: listed 4xx retryable; unlisted 4xx/400 stay terminal; 429-via-list without retry_on_429; 5xx unaffected; non-status failures never resurrected; existing classification test updated to the new signature with &[] (pinning default-unchanged).
  • e2e (real aisix + etcd): two identical two-target groups — default group relays the first target's 422 without consulting the backup; fallback_on_statuses: [422] group fails over and succeeds on the backup; a 400 from a configured group (422-only list) still relays. Existing fallback/fallback-edges e2es pass unchanged.

Ref api7/AISIX-Cloud#1012 (closing reference rides the CP PR so the issue doesn't close on the DP half)

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added configurable HTTP status codes that can trigger upstream retries and failover.
    • Supports status codes from 400–599, with existing behavior preserved when unset.
    • Applied configurable fallback handling across chat, message, response, and token-counting requests.
  • Bug Fixes

    • Improved failover behavior for selected upstream client-error responses.
  • Tests

    • Added validation and end-to-end coverage for configured, unconfigured, and unlisted status codes.

…into retry/failover
Some providers use non-429 4xx codes for transient conditions (model
overload, queue full, quota exhaustion). The gateway's default treats
any non-429 4xx as the caller's error and relays it, so a model group
fails a request its other targets could have served
(AISIX-Cloud#1012).
Routing models gain an explicit opt-in list:
routing:
strategy: failover
targets: [...]
fallback_on_statuses: [408, 409, 422]
A listed status becomes retryable in is_retryable() — the single
predicate every dispatch loop consults — so it participates in both
same-target retries and failover, exactly like retry_on_429 does for
429. Unset means unchanged behavior. 5xx codes are already retryable;
listing them is a no-op. The list never affects non-status failures
(customer-fixable config/credentials stay terminal). Schema constrains
entries to 400-599.
Threaded through all six dispatch call sites (chat streaming +
non-streaming, messages, responses, count_tokens). Per-attempt
telemetry already records each attempt's upstream status and kind
(initial/retry/fallback), so a rule-triggered failover is attributable
from existing fields.
Control-plane exposure (cp-admin schema, validation, projection,
dashboard) ships as the paired AISIX-Cloud PR; this changes the DP
first so the CP projection lands against a DP that accepts the field
(deny_unknown_fields).
Ref api7/AISIX-Cloud#1012
@coderabbitai

coderabbitaiBot commented Jul 10, 2026

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds optional fallback_on_statuses routing configuration for HTTP statuses 400–599, threads it through proxy retry decisions, and validates default, configured, and non-listed status behavior.

Changes

Configurable fallback status routing

Layer / File(s)Summary
Routing configuration and schema
crates/aisix-core/src/models/routing.rs, crates/aisix-core/src/models/schema.rs, schemas/resources/*.json
Adds optional fallback_on_statuses, validates statuses from 400 through 599, and provides an empty default slice.
Retry classification
crates/aisix-proxy/src/routing.rs
Updates is_retryable to treat configured upstream statuses as retryable while preserving terminal behavior for unlisted statuses and non-status failures.
Proxy dispatch integration
crates/aisix-proxy/src/chat.rs, crates/aisix-proxy/src/count_tokens.rs, crates/aisix-proxy/src/messages.rs, crates/aisix-proxy/src/responses.rs
Passes model-specific fallback statuses into streaming, non-streaming, messages, responses, and token-count retry/failover paths.
Fallback behavior validation
tests/e2e/src/cases/fallback-on-statuses-e2e.test.ts
Verifies default terminal handling, configured 422 failover, and terminal handling for an unlisted 400 response.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
participant Client
participant ProxyDispatch
participant RoutingConfig
participant Upstream
Client->>ProxyDispatch: send request
ProxyDispatch->>RoutingConfig: read fallback_on_statuses
ProxyDispatch->>Upstream: try primary target
Upstream-->>ProxyDispatch: 422 upstream error
ProxyDispatch->>ProxyDispatch: classify status as retryable
ProxyDispatch->>Upstream: try fallback target
Upstream-->>ProxyDispatch: successful response
ProxyDispatch-->>Client: return response
Loading

Possibly related PRs

  • api7/aisix#472: Refactors target resolution in the same proxy routing dispatch paths extended here with fallback-status retry decisions.
🚥 Pre-merge checks | ✅ 6
✅ Passed checks (6 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title clearly summarizes the main routing change and the new opt-in fallback status behavior.
Linked Issues check✅ PassedCheck skipped because no linked issues were found for this pull request.
Out of Scope Changes check✅ PassedCheck skipped because no linked issues were found for this pull request.
E2e Test Quality Review✅ PassedReal E2E flow uses spawned app, etcd, and HTTP upstreams; scenarios cover default, opt-in, and unlisted statuses with clear, independent assertions.
Security Check✅ PassedNo issues found; the PR only adds an opt-in retry-status list and threads it through dispatch, with no secret/log/auth/TLS/db code touched.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/fallback-on-statuses

Comment @coderabbitai help to get the list of available commands.

…ange at the validator
Audit follow-ups: the six dispatch sites now borrow the configured
slice instead of cloning per request, and a schema-validation test
pins that out-of-range entries are rejected while in-range lists pass.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
crates/aisix-proxy/src/chat.rs (1)

1101-1114: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Duplicate routing-config extraction pattern; consider a Model-level helper.

fallback_statuses is derived with the same virtual_entry.value.routing.as_ref().map(|r| r.fallback_on_statuses_or_default()).unwrap_or(&[]) boilerplate twice in this file (streaming and non-streaming paths), and the identical pattern (for both retry_on_429 and now fallback_on_statuses) is repeated again in count_tokens.rs (and, per the graph context, in messages.rs/responses.rs). A small convenience method on Model (e.g. fallback_on_statuses_or_default(&self) -> &[u16], mirroring the existing Routing method) would collapse all these call sites to a single line and remove the risk of the branches drifting.

♻️ Proposed helper (illustrative)
implModel{pubfnfallback_on_statuses_or_default(&self) -> &[u16]{self.routing.as_ref().map(|r| r.fallback_on_statuses_or_default()).unwrap_or(&[])}}

Then each call site becomes let fallback_statuses = virtual_entry.value.fallback_on_statuses_or_default();.

Also applies to: 1822-1833

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@crates/aisix-proxy/src/chat.rs` around lines 1101 - 1114, Introduce a
Model-level helper, such as Model::fallback_on_statuses_or_default, that
delegates to routing when present and otherwise returns an empty slice. Replace
the repeated routing.as_ref().map(...).unwrap_or(&[]) extraction in both chat.rs
paths and the corresponding count_tokens.rs, messages.rs, and responses.rs call
sites with this helper, including the related retry configuration pattern where
applicable.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@crates/aisix-proxy/src/chat.rs`:
- Around line 1101-1114: Introduce a Model-level helper, such as
Model::fallback_on_statuses_or_default, that delegates to routing when present
and otherwise returns an empty slice. Replace the repeated
routing.as_ref().map(...).unwrap_or(&[]) extraction in both chat.rs paths and
the corresponding count_tokens.rs, messages.rs, and responses.rs call sites with
this helper, including the related retry configuration pattern where applicable.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: cfffe33d-ca44-4bce-af17-40c31da81ac5

📥 Commits

Reviewing files that changed from the base of the PR and between 31357d1 and 5c0dfcf.

📒 Files selected for processing (10)
  • crates/aisix-core/src/models/routing.rs
  • crates/aisix-core/src/models/schema.rs
  • crates/aisix-proxy/src/chat.rs
  • crates/aisix-proxy/src/count_tokens.rs
  • crates/aisix-proxy/src/messages.rs
  • crates/aisix-proxy/src/responses.rs
  • crates/aisix-proxy/src/routing.rs
  • schemas/resources/model.schema.json
  • schemas/resources/routing.schema.json
  • tests/e2e/src/cases/fallback-on-statuses-e2e.test.ts

@jarvis9443
jarvis9443 merged commit c750e4c into mainJul 10, 2026
12 checks passed
@jarvis9443
jarvis9443 deleted the feat/fallback-on-statuses branch July 10, 2026 12:57
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jarvis9443
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(routing): fallback_on_statuses — opt selected upstream statuses into retry/failover - #752

Merged
jarvis9443 merged 2 commits into
mainfrom
feat/fallback-on-statuses
Jul 10, 2026
Merged

feat(routing): fallback_on_statuses — opt selected upstream statuses into retry/failover#752
jarvis9443 merged 2 commits into
mainfrom
feat/fallback-on-statuses

Conversation

@jarvis9443

@jarvis9443jarvis9443 commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

Problem

Some providers use non-429 4xx codes for transient conditions — model overload, queue full, quota exhaustion. The default classification (is_retryable) treats every non-429 4xx as the caller's error and relays it, so a model group fails a request that its other targets could have served. (AISIX-Cloud#1012)

Change

Routing models gain an explicit, empty-by-default opt-in list:

routing:
strategy: failovertargets: [{ model: primary }, { model: backup }]fallback_on_statuses: [408, 409, 422]

A listed status becomes retryable in is_retryable() — the single predicate every dispatch loop consults — so it participates in both same-target retries (retries) and failover (max_fallbacks), exactly like retry_on_429 does for 429 (429 in the list works without retry_on_429; the boolean stays for compatibility). Unset = behavior unchanged, bit-for-bit. 5xx are already retryable (listing them is a no-op); the list never affects non-status failures (customer-fixable config/credential errors stay terminal). Schema constrains entries to 400–599.

Threaded through all six dispatch sites: chat streaming + non-streaming + cross-target break, messages, responses, count_tokens — the whole routing family in one PR.

Observability: per-attempt telemetry already records each attempt's upstream status, error_class and attempt_kind (initial/retry/fallback), so a rule-triggered failover is fully attributable from existing fields; no new telemetry needed.

Interaction note: this is the current-request knob. The pre-existing per-direct-model cooldown.trigger_statuses benches a target for future requests and stays an independent layer.

Deliberately not included

  • Provider error-code / response-body matchers (the issue floats error.code: Throttling-style matching as a possible extension): status-code opt-in covers the reported scenarios; body matching adds a parsing surface per provider and can be layered on later if a concrete provider needs it.
  • Changing defaults (e.g. making 408/409 retryable out of the box, as some gateways do): the issue explicitly requires unchanged defaults.

Baseline: mainstream gateways retry 408/409/429/5xx by default and let users force retries per exception category; none offer per-status opt-in. We keep stricter defaults and add the more precise status-list control instead — providers in scope here signal overload with codes (e.g. 422) that category-based systems can't distinguish from validation errors.

Control-plane pairing

The CP is a closed schema — this field is unreachable by users until the paired AISIX-Cloud PR (cp-admin.yaml + validation + etcd projection + dashboard) lands. DP ships first because it rejects unknown fields (deny_unknown_fields); the CP PR follows immediately and carries the closing reference for AISIX-Cloud#1012.

Tests

  • Unit: listed 4xx retryable; unlisted 4xx/400 stay terminal; 429-via-list without retry_on_429; 5xx unaffected; non-status failures never resurrected; existing classification test updated to the new signature with &[] (pinning default-unchanged).
  • e2e (real aisix + etcd): two identical two-target groups — default group relays the first target's 422 without consulting the backup; fallback_on_statuses: [422] group fails over and succeeds on the backup; a 400 from a configured group (422-only list) still relays. Existing fallback/fallback-edges e2es pass unchanged.

Ref api7/AISIX-Cloud#1012 (closing reference rides the CP PR so the issue doesn't close on the DP half)

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added configurable HTTP status codes that can trigger upstream retries and failover.
    • Supports status codes from 400–599, with existing behavior preserved when unset.
    • Applied configurable fallback handling across chat, message, response, and token-counting requests.
  • Bug Fixes

    • Improved failover behavior for selected upstream client-error responses.
  • Tests

    • Added validation and end-to-end coverage for configured, unconfigured, and unlisted status codes.

…into retry/failover
Some providers use non-429 4xx codes for transient conditions (model
overload, queue full, quota exhaustion). The gateway's default treats
any non-429 4xx as the caller's error and relays it, so a model group
fails a request its other targets could have served
(AISIX-Cloud#1012).
Routing models gain an explicit opt-in list:
routing:
strategy: failover
targets: [...]
fallback_on_statuses: [408, 409, 422]
A listed status becomes retryable in is_retryable() — the single
predicate every dispatch loop consults — so it participates in both
same-target retries and failover, exactly like retry_on_429 does for
429. Unset means unchanged behavior. 5xx codes are already retryable;
listing them is a no-op. The list never affects non-status failures
(customer-fixable config/credentials stay terminal). Schema constrains
entries to 400-599.
Threaded through all six dispatch call sites (chat streaming +
non-streaming, messages, responses, count_tokens). Per-attempt
telemetry already records each attempt's upstream status and kind
(initial/retry/fallback), so a rule-triggered failover is attributable
from existing fields.
Control-plane exposure (cp-admin schema, validation, projection,
dashboard) ships as the paired AISIX-Cloud PR; this changes the DP
first so the CP projection lands against a DP that accepts the field
(deny_unknown_fields).
Ref api7/AISIX-Cloud#1012
@coderabbitai

coderabbitaiBot commented Jul 10, 2026

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds optional fallback_on_statuses routing configuration for HTTP statuses 400–599, threads it through proxy retry decisions, and validates default, configured, and non-listed status behavior.

Changes

Configurable fallback status routing

Layer / File(s)Summary
Routing configuration and schema
crates/aisix-core/src/models/routing.rs, crates/aisix-core/src/models/schema.rs, schemas/resources/*.json
Adds optional fallback_on_statuses, validates statuses from 400 through 599, and provides an empty default slice.
Retry classification
crates/aisix-proxy/src/routing.rs
Updates is_retryable to treat configured upstream statuses as retryable while preserving terminal behavior for unlisted statuses and non-status failures.
Proxy dispatch integration
crates/aisix-proxy/src/chat.rs, crates/aisix-proxy/src/count_tokens.rs, crates/aisix-proxy/src/messages.rs, crates/aisix-proxy/src/responses.rs
Passes model-specific fallback statuses into streaming, non-streaming, messages, responses, and token-count retry/failover paths.
Fallback behavior validation
tests/e2e/src/cases/fallback-on-statuses-e2e.test.ts
Verifies default terminal handling, configured 422 failover, and terminal handling for an unlisted 400 response.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
participant Client
participant ProxyDispatch
participant RoutingConfig
participant Upstream
Client->>ProxyDispatch: send request
ProxyDispatch->>RoutingConfig: read fallback_on_statuses
ProxyDispatch->>Upstream: try primary target
Upstream-->>ProxyDispatch: 422 upstream error
ProxyDispatch->>ProxyDispatch: classify status as retryable
ProxyDispatch->>Upstream: try fallback target
Upstream-->>ProxyDispatch: successful response
ProxyDispatch-->>Client: return response
Loading

Possibly related PRs

  • api7/aisix#472: Refactors target resolution in the same proxy routing dispatch paths extended here with fallback-status retry decisions.
🚥 Pre-merge checks | ✅ 6
✅ Passed checks (6 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title clearly summarizes the main routing change and the new opt-in fallback status behavior.
Linked Issues check✅ PassedCheck skipped because no linked issues were found for this pull request.
Out of Scope Changes check✅ PassedCheck skipped because no linked issues were found for this pull request.
E2e Test Quality Review✅ PassedReal E2E flow uses spawned app, etcd, and HTTP upstreams; scenarios cover default, opt-in, and unlisted statuses with clear, independent assertions.
Security Check✅ PassedNo issues found; the PR only adds an opt-in retry-status list and threads it through dispatch, with no secret/log/auth/TLS/db code touched.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/fallback-on-statuses

Comment @coderabbitai help to get the list of available commands.

…ange at the validator
Audit follow-ups: the six dispatch sites now borrow the configured
slice instead of cloning per request, and a schema-validation test
pins that out-of-range entries are rejected while in-range lists pass.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
crates/aisix-proxy/src/chat.rs (1)

1101-1114: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Duplicate routing-config extraction pattern; consider a Model-level helper.

fallback_statuses is derived with the same virtual_entry.value.routing.as_ref().map(|r| r.fallback_on_statuses_or_default()).unwrap_or(&[]) boilerplate twice in this file (streaming and non-streaming paths), and the identical pattern (for both retry_on_429 and now fallback_on_statuses) is repeated again in count_tokens.rs (and, per the graph context, in messages.rs/responses.rs). A small convenience method on Model (e.g. fallback_on_statuses_or_default(&self) -> &[u16], mirroring the existing Routing method) would collapse all these call sites to a single line and remove the risk of the branches drifting.

♻️ Proposed helper (illustrative)
implModel{pubfnfallback_on_statuses_or_default(&self) -> &[u16]{self.routing.as_ref().map(|r| r.fallback_on_statuses_or_default()).unwrap_or(&[])}}

Then each call site becomes let fallback_statuses = virtual_entry.value.fallback_on_statuses_or_default();.

Also applies to: 1822-1833

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@crates/aisix-proxy/src/chat.rs` around lines 1101 - 1114, Introduce a
Model-level helper, such as Model::fallback_on_statuses_or_default, that
delegates to routing when present and otherwise returns an empty slice. Replace
the repeated routing.as_ref().map(...).unwrap_or(&[]) extraction in both chat.rs
paths and the corresponding count_tokens.rs, messages.rs, and responses.rs call
sites with this helper, including the related retry configuration pattern where
applicable.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@crates/aisix-proxy/src/chat.rs`:
- Around line 1101-1114: Introduce a Model-level helper, such as
Model::fallback_on_statuses_or_default, that delegates to routing when present
and otherwise returns an empty slice. Replace the repeated
routing.as_ref().map(...).unwrap_or(&[]) extraction in both chat.rs paths and
the corresponding count_tokens.rs, messages.rs, and responses.rs call sites with
this helper, including the related retry configuration pattern where applicable.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: cfffe33d-ca44-4bce-af17-40c31da81ac5

📥 Commits

Reviewing files that changed from the base of the PR and between 31357d1 and 5c0dfcf.

📒 Files selected for processing (10)
  • crates/aisix-core/src/models/routing.rs
  • crates/aisix-core/src/models/schema.rs
  • crates/aisix-proxy/src/chat.rs
  • crates/aisix-proxy/src/count_tokens.rs
  • crates/aisix-proxy/src/messages.rs
  • crates/aisix-proxy/src/responses.rs
  • crates/aisix-proxy/src/routing.rs
  • schemas/resources/model.schema.json
  • schemas/resources/routing.schema.json
  • tests/e2e/src/cases/fallback-on-statuses-e2e.test.ts

@jarvis9443
jarvis9443 merged commit c750e4c into mainJul 10, 2026
12 checks passed
@jarvis9443
jarvis9443 deleted the feat/fallback-on-statuses branch July 10, 2026 12:57
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jarvis9443
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(routing): fallback_on_statuses — opt selected upstream statuses into retry/failover - #752

Merged
jarvis9443 merged 2 commits into
mainfrom
feat/fallback-on-statuses
Jul 10, 2026
Merged

feat(routing): fallback_on_statuses — opt selected upstream statuses into retry/failover#752
jarvis9443 merged 2 commits into
mainfrom
feat/fallback-on-statuses

Conversation

@jarvis9443

@jarvis9443jarvis9443 commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

Problem

Some providers use non-429 4xx codes for transient conditions — model overload, queue full, quota exhaustion. The default classification (is_retryable) treats every non-429 4xx as the caller's error and relays it, so a model group fails a request that its other targets could have served. (AISIX-Cloud#1012)

Change

Routing models gain an explicit, empty-by-default opt-in list:

routing:
strategy: failovertargets: [{ model: primary }, { model: backup }]fallback_on_statuses: [408, 409, 422]

A listed status becomes retryable in is_retryable() — the single predicate every dispatch loop consults — so it participates in both same-target retries (retries) and failover (max_fallbacks), exactly like retry_on_429 does for 429 (429 in the list works without retry_on_429; the boolean stays for compatibility). Unset = behavior unchanged, bit-for-bit. 5xx are already retryable (listing them is a no-op); the list never affects non-status failures (customer-fixable config/credential errors stay terminal). Schema constrains entries to 400–599.

Threaded through all six dispatch sites: chat streaming + non-streaming + cross-target break, messages, responses, count_tokens — the whole routing family in one PR.

Observability: per-attempt telemetry already records each attempt's upstream status, error_class and attempt_kind (initial/retry/fallback), so a rule-triggered failover is fully attributable from existing fields; no new telemetry needed.

Interaction note: this is the current-request knob. The pre-existing per-direct-model cooldown.trigger_statuses benches a target for future requests and stays an independent layer.

Deliberately not included

  • Provider error-code / response-body matchers (the issue floats error.code: Throttling-style matching as a possible extension): status-code opt-in covers the reported scenarios; body matching adds a parsing surface per provider and can be layered on later if a concrete provider needs it.
  • Changing defaults (e.g. making 408/409 retryable out of the box, as some gateways do): the issue explicitly requires unchanged defaults.

Baseline: mainstream gateways retry 408/409/429/5xx by default and let users force retries per exception category; none offer per-status opt-in. We keep stricter defaults and add the more precise status-list control instead — providers in scope here signal overload with codes (e.g. 422) that category-based systems can't distinguish from validation errors.

Control-plane pairing

The CP is a closed schema — this field is unreachable by users until the paired AISIX-Cloud PR (cp-admin.yaml + validation + etcd projection + dashboard) lands. DP ships first because it rejects unknown fields (deny_unknown_fields); the CP PR follows immediately and carries the closing reference for AISIX-Cloud#1012.

Tests

  • Unit: listed 4xx retryable; unlisted 4xx/400 stay terminal; 429-via-list without retry_on_429; 5xx unaffected; non-status failures never resurrected; existing classification test updated to the new signature with &[] (pinning default-unchanged).
  • e2e (real aisix + etcd): two identical two-target groups — default group relays the first target's 422 without consulting the backup; fallback_on_statuses: [422] group fails over and succeeds on the backup; a 400 from a configured group (422-only list) still relays. Existing fallback/fallback-edges e2es pass unchanged.

Ref api7/AISIX-Cloud#1012 (closing reference rides the CP PR so the issue doesn't close on the DP half)

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added configurable HTTP status codes that can trigger upstream retries and failover.
    • Supports status codes from 400–599, with existing behavior preserved when unset.
    • Applied configurable fallback handling across chat, message, response, and token-counting requests.
  • Bug Fixes

    • Improved failover behavior for selected upstream client-error responses.
  • Tests

    • Added validation and end-to-end coverage for configured, unconfigured, and unlisted status codes.

…into retry/failover
Some providers use non-429 4xx codes for transient conditions (model
overload, queue full, quota exhaustion). The gateway's default treats
any non-429 4xx as the caller's error and relays it, so a model group
fails a request its other targets could have served
(AISIX-Cloud#1012).
Routing models gain an explicit opt-in list:
routing:
strategy: failover
targets: [...]
fallback_on_statuses: [408, 409, 422]
A listed status becomes retryable in is_retryable() — the single
predicate every dispatch loop consults — so it participates in both
same-target retries and failover, exactly like retry_on_429 does for
429. Unset means unchanged behavior. 5xx codes are already retryable;
listing them is a no-op. The list never affects non-status failures
(customer-fixable config/credentials stay terminal). Schema constrains
entries to 400-599.
Threaded through all six dispatch call sites (chat streaming +
non-streaming, messages, responses, count_tokens). Per-attempt
telemetry already records each attempt's upstream status and kind
(initial/retry/fallback), so a rule-triggered failover is attributable
from existing fields.
Control-plane exposure (cp-admin schema, validation, projection,
dashboard) ships as the paired AISIX-Cloud PR; this changes the DP
first so the CP projection lands against a DP that accepts the field
(deny_unknown_fields).
Ref api7/AISIX-Cloud#1012
@coderabbitai

coderabbitaiBot commented Jul 10, 2026

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds optional fallback_on_statuses routing configuration for HTTP statuses 400–599, threads it through proxy retry decisions, and validates default, configured, and non-listed status behavior.

Changes

Configurable fallback status routing

Layer / File(s)Summary
Routing configuration and schema
crates/aisix-core/src/models/routing.rs, crates/aisix-core/src/models/schema.rs, schemas/resources/*.json
Adds optional fallback_on_statuses, validates statuses from 400 through 599, and provides an empty default slice.
Retry classification
crates/aisix-proxy/src/routing.rs
Updates is_retryable to treat configured upstream statuses as retryable while preserving terminal behavior for unlisted statuses and non-status failures.
Proxy dispatch integration
crates/aisix-proxy/src/chat.rs, crates/aisix-proxy/src/count_tokens.rs, crates/aisix-proxy/src/messages.rs, crates/aisix-proxy/src/responses.rs
Passes model-specific fallback statuses into streaming, non-streaming, messages, responses, and token-count retry/failover paths.
Fallback behavior validation
tests/e2e/src/cases/fallback-on-statuses-e2e.test.ts
Verifies default terminal handling, configured 422 failover, and terminal handling for an unlisted 400 response.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
participant Client
participant ProxyDispatch
participant RoutingConfig
participant Upstream
Client->>ProxyDispatch: send request
ProxyDispatch->>RoutingConfig: read fallback_on_statuses
ProxyDispatch->>Upstream: try primary target
Upstream-->>ProxyDispatch: 422 upstream error
ProxyDispatch->>ProxyDispatch: classify status as retryable
ProxyDispatch->>Upstream: try fallback target
Upstream-->>ProxyDispatch: successful response
ProxyDispatch-->>Client: return response
Loading

Possibly related PRs

  • api7/aisix#472: Refactors target resolution in the same proxy routing dispatch paths extended here with fallback-status retry decisions.
🚥 Pre-merge checks | ✅ 6
✅ Passed checks (6 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title clearly summarizes the main routing change and the new opt-in fallback status behavior.
Linked Issues check✅ PassedCheck skipped because no linked issues were found for this pull request.
Out of Scope Changes check✅ PassedCheck skipped because no linked issues were found for this pull request.
E2e Test Quality Review✅ PassedReal E2E flow uses spawned app, etcd, and HTTP upstreams; scenarios cover default, opt-in, and unlisted statuses with clear, independent assertions.
Security Check✅ PassedNo issues found; the PR only adds an opt-in retry-status list and threads it through dispatch, with no secret/log/auth/TLS/db code touched.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/fallback-on-statuses

Comment @coderabbitai help to get the list of available commands.

…ange at the validator
Audit follow-ups: the six dispatch sites now borrow the configured
slice instead of cloning per request, and a schema-validation test
pins that out-of-range entries are rejected while in-range lists pass.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
crates/aisix-proxy/src/chat.rs (1)

1101-1114: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Duplicate routing-config extraction pattern; consider a Model-level helper.

fallback_statuses is derived with the same virtual_entry.value.routing.as_ref().map(|r| r.fallback_on_statuses_or_default()).unwrap_or(&[]) boilerplate twice in this file (streaming and non-streaming paths), and the identical pattern (for both retry_on_429 and now fallback_on_statuses) is repeated again in count_tokens.rs (and, per the graph context, in messages.rs/responses.rs). A small convenience method on Model (e.g. fallback_on_statuses_or_default(&self) -> &[u16], mirroring the existing Routing method) would collapse all these call sites to a single line and remove the risk of the branches drifting.

♻️ Proposed helper (illustrative)
implModel{pubfnfallback_on_statuses_or_default(&self) -> &[u16]{self.routing.as_ref().map(|r| r.fallback_on_statuses_or_default()).unwrap_or(&[])}}

Then each call site becomes let fallback_statuses = virtual_entry.value.fallback_on_statuses_or_default();.

Also applies to: 1822-1833

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@crates/aisix-proxy/src/chat.rs` around lines 1101 - 1114, Introduce a
Model-level helper, such as Model::fallback_on_statuses_or_default, that
delegates to routing when present and otherwise returns an empty slice. Replace
the repeated routing.as_ref().map(...).unwrap_or(&[]) extraction in both chat.rs
paths and the corresponding count_tokens.rs, messages.rs, and responses.rs call
sites with this helper, including the related retry configuration pattern where
applicable.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@crates/aisix-proxy/src/chat.rs`:
- Around line 1101-1114: Introduce a Model-level helper, such as
Model::fallback_on_statuses_or_default, that delegates to routing when present
and otherwise returns an empty slice. Replace the repeated
routing.as_ref().map(...).unwrap_or(&[]) extraction in both chat.rs paths and
the corresponding count_tokens.rs, messages.rs, and responses.rs call sites with
this helper, including the related retry configuration pattern where applicable.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: cfffe33d-ca44-4bce-af17-40c31da81ac5

📥 Commits

Reviewing files that changed from the base of the PR and between 31357d1 and 5c0dfcf.

📒 Files selected for processing (10)
  • crates/aisix-core/src/models/routing.rs
  • crates/aisix-core/src/models/schema.rs
  • crates/aisix-proxy/src/chat.rs
  • crates/aisix-proxy/src/count_tokens.rs
  • crates/aisix-proxy/src/messages.rs
  • crates/aisix-proxy/src/responses.rs
  • crates/aisix-proxy/src/routing.rs
  • schemas/resources/model.schema.json
  • schemas/resources/routing.schema.json
  • tests/e2e/src/cases/fallback-on-statuses-e2e.test.ts

@jarvis9443
jarvis9443 merged commit c750e4c into mainJul 10, 2026
12 checks passed
@jarvis9443
jarvis9443 deleted the feat/fallback-on-statuses branch July 10, 2026 12:57
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jarvis9443
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(routing): fallback_on_statuses — opt selected upstream statuses into retry/failover - #752

Merged
jarvis9443 merged 2 commits into
mainfrom
feat/fallback-on-statuses
Jul 10, 2026
Merged

feat(routing): fallback_on_statuses — opt selected upstream statuses into retry/failover#752
jarvis9443 merged 2 commits into
mainfrom
feat/fallback-on-statuses

Conversation

@jarvis9443

@jarvis9443jarvis9443 commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

Problem

Some providers use non-429 4xx codes for transient conditions — model overload, queue full, quota exhaustion. The default classification (is_retryable) treats every non-429 4xx as the caller's error and relays it, so a model group fails a request that its other targets could have served. (AISIX-Cloud#1012)

Change

Routing models gain an explicit, empty-by-default opt-in list:

routing:
strategy: failovertargets: [{ model: primary }, { model: backup }]fallback_on_statuses: [408, 409, 422]

A listed status becomes retryable in is_retryable() — the single predicate every dispatch loop consults — so it participates in both same-target retries (retries) and failover (max_fallbacks), exactly like retry_on_429 does for 429 (429 in the list works without retry_on_429; the boolean stays for compatibility). Unset = behavior unchanged, bit-for-bit. 5xx are already retryable (listing them is a no-op); the list never affects non-status failures (customer-fixable config/credential errors stay terminal). Schema constrains entries to 400–599.

Threaded through all six dispatch sites: chat streaming + non-streaming + cross-target break, messages, responses, count_tokens — the whole routing family in one PR.

Observability: per-attempt telemetry already records each attempt's upstream status, error_class and attempt_kind (initial/retry/fallback), so a rule-triggered failover is fully attributable from existing fields; no new telemetry needed.

Interaction note: this is the current-request knob. The pre-existing per-direct-model cooldown.trigger_statuses benches a target for future requests and stays an independent layer.

Deliberately not included

  • Provider error-code / response-body matchers (the issue floats error.code: Throttling-style matching as a possible extension): status-code opt-in covers the reported scenarios; body matching adds a parsing surface per provider and can be layered on later if a concrete provider needs it.
  • Changing defaults (e.g. making 408/409 retryable out of the box, as some gateways do): the issue explicitly requires unchanged defaults.

Baseline: mainstream gateways retry 408/409/429/5xx by default and let users force retries per exception category; none offer per-status opt-in. We keep stricter defaults and add the more precise status-list control instead — providers in scope here signal overload with codes (e.g. 422) that category-based systems can't distinguish from validation errors.

Control-plane pairing

The CP is a closed schema — this field is unreachable by users until the paired AISIX-Cloud PR (cp-admin.yaml + validation + etcd projection + dashboard) lands. DP ships first because it rejects unknown fields (deny_unknown_fields); the CP PR follows immediately and carries the closing reference for AISIX-Cloud#1012.

Tests

  • Unit: listed 4xx retryable; unlisted 4xx/400 stay terminal; 429-via-list without retry_on_429; 5xx unaffected; non-status failures never resurrected; existing classification test updated to the new signature with &[] (pinning default-unchanged).
  • e2e (real aisix + etcd): two identical two-target groups — default group relays the first target's 422 without consulting the backup; fallback_on_statuses: [422] group fails over and succeeds on the backup; a 400 from a configured group (422-only list) still relays. Existing fallback/fallback-edges e2es pass unchanged.

Ref api7/AISIX-Cloud#1012 (closing reference rides the CP PR so the issue doesn't close on the DP half)

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added configurable HTTP status codes that can trigger upstream retries and failover.
    • Supports status codes from 400–599, with existing behavior preserved when unset.
    • Applied configurable fallback handling across chat, message, response, and token-counting requests.
  • Bug Fixes

    • Improved failover behavior for selected upstream client-error responses.
  • Tests

    • Added validation and end-to-end coverage for configured, unconfigured, and unlisted status codes.

…into retry/failover
Some providers use non-429 4xx codes for transient conditions (model
overload, queue full, quota exhaustion). The gateway's default treats
any non-429 4xx as the caller's error and relays it, so a model group
fails a request its other targets could have served
(AISIX-Cloud#1012).
Routing models gain an explicit opt-in list:
routing:
strategy: failover
targets: [...]
fallback_on_statuses: [408, 409, 422]
A listed status becomes retryable in is_retryable() — the single
predicate every dispatch loop consults — so it participates in both
same-target retries and failover, exactly like retry_on_429 does for
429. Unset means unchanged behavior. 5xx codes are already retryable;
listing them is a no-op. The list never affects non-status failures
(customer-fixable config/credentials stay terminal). Schema constrains
entries to 400-599.
Threaded through all six dispatch call sites (chat streaming +
non-streaming, messages, responses, count_tokens). Per-attempt
telemetry already records each attempt's upstream status and kind
(initial/retry/fallback), so a rule-triggered failover is attributable
from existing fields.
Control-plane exposure (cp-admin schema, validation, projection,
dashboard) ships as the paired AISIX-Cloud PR; this changes the DP
first so the CP projection lands against a DP that accepts the field
(deny_unknown_fields).
Ref api7/AISIX-Cloud#1012
@coderabbitai

coderabbitaiBot commented Jul 10, 2026

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds optional fallback_on_statuses routing configuration for HTTP statuses 400–599, threads it through proxy retry decisions, and validates default, configured, and non-listed status behavior.

Changes

Configurable fallback status routing

Layer / File(s)Summary
Routing configuration and schema
crates/aisix-core/src/models/routing.rs, crates/aisix-core/src/models/schema.rs, schemas/resources/*.json
Adds optional fallback_on_statuses, validates statuses from 400 through 599, and provides an empty default slice.
Retry classification
crates/aisix-proxy/src/routing.rs
Updates is_retryable to treat configured upstream statuses as retryable while preserving terminal behavior for unlisted statuses and non-status failures.
Proxy dispatch integration
crates/aisix-proxy/src/chat.rs, crates/aisix-proxy/src/count_tokens.rs, crates/aisix-proxy/src/messages.rs, crates/aisix-proxy/src/responses.rs
Passes model-specific fallback statuses into streaming, non-streaming, messages, responses, and token-count retry/failover paths.
Fallback behavior validation
tests/e2e/src/cases/fallback-on-statuses-e2e.test.ts
Verifies default terminal handling, configured 422 failover, and terminal handling for an unlisted 400 response.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
participant Client
participant ProxyDispatch
participant RoutingConfig
participant Upstream
Client->>ProxyDispatch: send request
ProxyDispatch->>RoutingConfig: read fallback_on_statuses
ProxyDispatch->>Upstream: try primary target
Upstream-->>ProxyDispatch: 422 upstream error
ProxyDispatch->>ProxyDispatch: classify status as retryable
ProxyDispatch->>Upstream: try fallback target
Upstream-->>ProxyDispatch: successful response
ProxyDispatch-->>Client: return response
Loading

Possibly related PRs

  • api7/aisix#472: Refactors target resolution in the same proxy routing dispatch paths extended here with fallback-status retry decisions.
🚥 Pre-merge checks | ✅ 6
✅ Passed checks (6 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title clearly summarizes the main routing change and the new opt-in fallback status behavior.
Linked Issues check✅ PassedCheck skipped because no linked issues were found for this pull request.
Out of Scope Changes check✅ PassedCheck skipped because no linked issues were found for this pull request.
E2e Test Quality Review✅ PassedReal E2E flow uses spawned app, etcd, and HTTP upstreams; scenarios cover default, opt-in, and unlisted statuses with clear, independent assertions.
Security Check✅ PassedNo issues found; the PR only adds an opt-in retry-status list and threads it through dispatch, with no secret/log/auth/TLS/db code touched.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/fallback-on-statuses

Comment @coderabbitai help to get the list of available commands.

…ange at the validator
Audit follow-ups: the six dispatch sites now borrow the configured
slice instead of cloning per request, and a schema-validation test
pins that out-of-range entries are rejected while in-range lists pass.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
crates/aisix-proxy/src/chat.rs (1)

1101-1114: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Duplicate routing-config extraction pattern; consider a Model-level helper.

fallback_statuses is derived with the same virtual_entry.value.routing.as_ref().map(|r| r.fallback_on_statuses_or_default()).unwrap_or(&[]) boilerplate twice in this file (streaming and non-streaming paths), and the identical pattern (for both retry_on_429 and now fallback_on_statuses) is repeated again in count_tokens.rs (and, per the graph context, in messages.rs/responses.rs). A small convenience method on Model (e.g. fallback_on_statuses_or_default(&self) -> &[u16], mirroring the existing Routing method) would collapse all these call sites to a single line and remove the risk of the branches drifting.

♻️ Proposed helper (illustrative)
implModel{pubfnfallback_on_statuses_or_default(&self) -> &[u16]{self.routing.as_ref().map(|r| r.fallback_on_statuses_or_default()).unwrap_or(&[])}}

Then each call site becomes let fallback_statuses = virtual_entry.value.fallback_on_statuses_or_default();.

Also applies to: 1822-1833

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@crates/aisix-proxy/src/chat.rs` around lines 1101 - 1114, Introduce a
Model-level helper, such as Model::fallback_on_statuses_or_default, that
delegates to routing when present and otherwise returns an empty slice. Replace
the repeated routing.as_ref().map(...).unwrap_or(&[]) extraction in both chat.rs
paths and the corresponding count_tokens.rs, messages.rs, and responses.rs call
sites with this helper, including the related retry configuration pattern where
applicable.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@crates/aisix-proxy/src/chat.rs`:
- Around line 1101-1114: Introduce a Model-level helper, such as
Model::fallback_on_statuses_or_default, that delegates to routing when present
and otherwise returns an empty slice. Replace the repeated
routing.as_ref().map(...).unwrap_or(&[]) extraction in both chat.rs paths and
the corresponding count_tokens.rs, messages.rs, and responses.rs call sites with
this helper, including the related retry configuration pattern where applicable.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: cfffe33d-ca44-4bce-af17-40c31da81ac5

📥 Commits

Reviewing files that changed from the base of the PR and between 31357d1 and 5c0dfcf.

📒 Files selected for processing (10)
  • crates/aisix-core/src/models/routing.rs
  • crates/aisix-core/src/models/schema.rs
  • crates/aisix-proxy/src/chat.rs
  • crates/aisix-proxy/src/count_tokens.rs
  • crates/aisix-proxy/src/messages.rs
  • crates/aisix-proxy/src/responses.rs
  • crates/aisix-proxy/src/routing.rs
  • schemas/resources/model.schema.json
  • schemas/resources/routing.schema.json
  • tests/e2e/src/cases/fallback-on-statuses-e2e.test.ts

@jarvis9443
jarvis9443 merged commit c750e4c into mainJul 10, 2026
12 checks passed
@jarvis9443
jarvis9443 deleted the feat/fallback-on-statuses branch July 10, 2026 12:57
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jarvis9443
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(routing): fallback_on_statuses — opt selected upstream statuses into retry/failover - #752

Merged
jarvis9443 merged 2 commits into
mainfrom
feat/fallback-on-statuses
Jul 10, 2026
Merged

feat(routing): fallback_on_statuses — opt selected upstream statuses into retry/failover#752
jarvis9443 merged 2 commits into
mainfrom
feat/fallback-on-statuses

Conversation

@jarvis9443

@jarvis9443jarvis9443 commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

Problem

Some providers use non-429 4xx codes for transient conditions — model overload, queue full, quota exhaustion. The default classification (is_retryable) treats every non-429 4xx as the caller's error and relays it, so a model group fails a request that its other targets could have served. (AISIX-Cloud#1012)

Change

Routing models gain an explicit, empty-by-default opt-in list:

routing:
strategy: failovertargets: [{ model: primary }, { model: backup }]fallback_on_statuses: [408, 409, 422]

A listed status becomes retryable in is_retryable() — the single predicate every dispatch loop consults — so it participates in both same-target retries (retries) and failover (max_fallbacks), exactly like retry_on_429 does for 429 (429 in the list works without retry_on_429; the boolean stays for compatibility). Unset = behavior unchanged, bit-for-bit. 5xx are already retryable (listing them is a no-op); the list never affects non-status failures (customer-fixable config/credential errors stay terminal). Schema constrains entries to 400–599.

Threaded through all six dispatch sites: chat streaming + non-streaming + cross-target break, messages, responses, count_tokens — the whole routing family in one PR.

Observability: per-attempt telemetry already records each attempt's upstream status, error_class and attempt_kind (initial/retry/fallback), so a rule-triggered failover is fully attributable from existing fields; no new telemetry needed.

Interaction note: this is the current-request knob. The pre-existing per-direct-model cooldown.trigger_statuses benches a target for future requests and stays an independent layer.

Deliberately not included

  • Provider error-code / response-body matchers (the issue floats error.code: Throttling-style matching as a possible extension): status-code opt-in covers the reported scenarios; body matching adds a parsing surface per provider and can be layered on later if a concrete provider needs it.
  • Changing defaults (e.g. making 408/409 retryable out of the box, as some gateways do): the issue explicitly requires unchanged defaults.

Baseline: mainstream gateways retry 408/409/429/5xx by default and let users force retries per exception category; none offer per-status opt-in. We keep stricter defaults and add the more precise status-list control instead — providers in scope here signal overload with codes (e.g. 422) that category-based systems can't distinguish from validation errors.

Control-plane pairing

The CP is a closed schema — this field is unreachable by users until the paired AISIX-Cloud PR (cp-admin.yaml + validation + etcd projection + dashboard) lands. DP ships first because it rejects unknown fields (deny_unknown_fields); the CP PR follows immediately and carries the closing reference for AISIX-Cloud#1012.

Tests

  • Unit: listed 4xx retryable; unlisted 4xx/400 stay terminal; 429-via-list without retry_on_429; 5xx unaffected; non-status failures never resurrected; existing classification test updated to the new signature with &[] (pinning default-unchanged).
  • e2e (real aisix + etcd): two identical two-target groups — default group relays the first target's 422 without consulting the backup; fallback_on_statuses: [422] group fails over and succeeds on the backup; a 400 from a configured group (422-only list) still relays. Existing fallback/fallback-edges e2es pass unchanged.

Ref api7/AISIX-Cloud#1012 (closing reference rides the CP PR so the issue doesn't close on the DP half)

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added configurable HTTP status codes that can trigger upstream retries and failover.
    • Supports status codes from 400–599, with existing behavior preserved when unset.
    • Applied configurable fallback handling across chat, message, response, and token-counting requests.
  • Bug Fixes

    • Improved failover behavior for selected upstream client-error responses.
  • Tests

    • Added validation and end-to-end coverage for configured, unconfigured, and unlisted status codes.

…into retry/failover
Some providers use non-429 4xx codes for transient conditions (model
overload, queue full, quota exhaustion). The gateway's default treats
any non-429 4xx as the caller's error and relays it, so a model group
fails a request its other targets could have served
(AISIX-Cloud#1012).
Routing models gain an explicit opt-in list:
routing:
strategy: failover
targets: [...]
fallback_on_statuses: [408, 409, 422]
A listed status becomes retryable in is_retryable() — the single
predicate every dispatch loop consults — so it participates in both
same-target retries and failover, exactly like retry_on_429 does for
429. Unset means unchanged behavior. 5xx codes are already retryable;
listing them is a no-op. The list never affects non-status failures
(customer-fixable config/credentials stay terminal). Schema constrains
entries to 400-599.
Threaded through all six dispatch call sites (chat streaming +
non-streaming, messages, responses, count_tokens). Per-attempt
telemetry already records each attempt's upstream status and kind
(initial/retry/fallback), so a rule-triggered failover is attributable
from existing fields.
Control-plane exposure (cp-admin schema, validation, projection,
dashboard) ships as the paired AISIX-Cloud PR; this changes the DP
first so the CP projection lands against a DP that accepts the field
(deny_unknown_fields).
Ref api7/AISIX-Cloud#1012
@coderabbitai

coderabbitaiBot commented Jul 10, 2026

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds optional fallback_on_statuses routing configuration for HTTP statuses 400–599, threads it through proxy retry decisions, and validates default, configured, and non-listed status behavior.

Changes

Configurable fallback status routing

Layer / File(s)Summary
Routing configuration and schema
crates/aisix-core/src/models/routing.rs, crates/aisix-core/src/models/schema.rs, schemas/resources/*.json
Adds optional fallback_on_statuses, validates statuses from 400 through 599, and provides an empty default slice.
Retry classification
crates/aisix-proxy/src/routing.rs
Updates is_retryable to treat configured upstream statuses as retryable while preserving terminal behavior for unlisted statuses and non-status failures.
Proxy dispatch integration
crates/aisix-proxy/src/chat.rs, crates/aisix-proxy/src/count_tokens.rs, crates/aisix-proxy/src/messages.rs, crates/aisix-proxy/src/responses.rs
Passes model-specific fallback statuses into streaming, non-streaming, messages, responses, and token-count retry/failover paths.
Fallback behavior validation
tests/e2e/src/cases/fallback-on-statuses-e2e.test.ts
Verifies default terminal handling, configured 422 failover, and terminal handling for an unlisted 400 response.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
participant Client
participant ProxyDispatch
participant RoutingConfig
participant Upstream
Client->>ProxyDispatch: send request
ProxyDispatch->>RoutingConfig: read fallback_on_statuses
ProxyDispatch->>Upstream: try primary target
Upstream-->>ProxyDispatch: 422 upstream error
ProxyDispatch->>ProxyDispatch: classify status as retryable
ProxyDispatch->>Upstream: try fallback target
Upstream-->>ProxyDispatch: successful response
ProxyDispatch-->>Client: return response
Loading

Possibly related PRs

  • api7/aisix#472: Refactors target resolution in the same proxy routing dispatch paths extended here with fallback-status retry decisions.
🚥 Pre-merge checks | ✅ 6
✅ Passed checks (6 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title clearly summarizes the main routing change and the new opt-in fallback status behavior.
Linked Issues check✅ PassedCheck skipped because no linked issues were found for this pull request.
Out of Scope Changes check✅ PassedCheck skipped because no linked issues were found for this pull request.
E2e Test Quality Review✅ PassedReal E2E flow uses spawned app, etcd, and HTTP upstreams; scenarios cover default, opt-in, and unlisted statuses with clear, independent assertions.
Security Check✅ PassedNo issues found; the PR only adds an opt-in retry-status list and threads it through dispatch, with no secret/log/auth/TLS/db code touched.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/fallback-on-statuses

Comment @coderabbitai help to get the list of available commands.

…ange at the validator
Audit follow-ups: the six dispatch sites now borrow the configured
slice instead of cloning per request, and a schema-validation test
pins that out-of-range entries are rejected while in-range lists pass.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
crates/aisix-proxy/src/chat.rs (1)

1101-1114: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Duplicate routing-config extraction pattern; consider a Model-level helper.

fallback_statuses is derived with the same virtual_entry.value.routing.as_ref().map(|r| r.fallback_on_statuses_or_default()).unwrap_or(&[]) boilerplate twice in this file (streaming and non-streaming paths), and the identical pattern (for both retry_on_429 and now fallback_on_statuses) is repeated again in count_tokens.rs (and, per the graph context, in messages.rs/responses.rs). A small convenience method on Model (e.g. fallback_on_statuses_or_default(&self) -> &[u16], mirroring the existing Routing method) would collapse all these call sites to a single line and remove the risk of the branches drifting.

♻️ Proposed helper (illustrative)
implModel{pubfnfallback_on_statuses_or_default(&self) -> &[u16]{self.routing.as_ref().map(|r| r.fallback_on_statuses_or_default()).unwrap_or(&[])}}

Then each call site becomes let fallback_statuses = virtual_entry.value.fallback_on_statuses_or_default();.

Also applies to: 1822-1833

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@crates/aisix-proxy/src/chat.rs` around lines 1101 - 1114, Introduce a
Model-level helper, such as Model::fallback_on_statuses_or_default, that
delegates to routing when present and otherwise returns an empty slice. Replace
the repeated routing.as_ref().map(...).unwrap_or(&[]) extraction in both chat.rs
paths and the corresponding count_tokens.rs, messages.rs, and responses.rs call
sites with this helper, including the related retry configuration pattern where
applicable.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@crates/aisix-proxy/src/chat.rs`:
- Around line 1101-1114: Introduce a Model-level helper, such as
Model::fallback_on_statuses_or_default, that delegates to routing when present
and otherwise returns an empty slice. Replace the repeated
routing.as_ref().map(...).unwrap_or(&[]) extraction in both chat.rs paths and
the corresponding count_tokens.rs, messages.rs, and responses.rs call sites with
this helper, including the related retry configuration pattern where applicable.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: cfffe33d-ca44-4bce-af17-40c31da81ac5

📥 Commits

Reviewing files that changed from the base of the PR and between 31357d1 and 5c0dfcf.

📒 Files selected for processing (10)
  • crates/aisix-core/src/models/routing.rs
  • crates/aisix-core/src/models/schema.rs
  • crates/aisix-proxy/src/chat.rs
  • crates/aisix-proxy/src/count_tokens.rs
  • crates/aisix-proxy/src/messages.rs
  • crates/aisix-proxy/src/responses.rs
  • crates/aisix-proxy/src/routing.rs
  • schemas/resources/model.schema.json
  • schemas/resources/routing.schema.json
  • tests/e2e/src/cases/fallback-on-statuses-e2e.test.ts

@jarvis9443
jarvis9443 merged commit c750e4c into mainJul 10, 2026
12 checks passed
@jarvis9443
jarvis9443 deleted the feat/fallback-on-statuses branch July 10, 2026 12:57
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jarvis9443
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(routing): fallback_on_statuses — opt selected upstream statuses into retry/failover - #752

Merged
jarvis9443 merged 2 commits into
mainfrom
feat/fallback-on-statuses
Jul 10, 2026
Merged

feat(routing): fallback_on_statuses — opt selected upstream statuses into retry/failover#752
jarvis9443 merged 2 commits into
mainfrom
feat/fallback-on-statuses

Conversation

@jarvis9443

@jarvis9443jarvis9443 commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

Problem

Some providers use non-429 4xx codes for transient conditions — model overload, queue full, quota exhaustion. The default classification (is_retryable) treats every non-429 4xx as the caller's error and relays it, so a model group fails a request that its other targets could have served. (AISIX-Cloud#1012)

Change

Routing models gain an explicit, empty-by-default opt-in list:

routing:
strategy: failovertargets: [{ model: primary }, { model: backup }]fallback_on_statuses: [408, 409, 422]

A listed status becomes retryable in is_retryable() — the single predicate every dispatch loop consults — so it participates in both same-target retries (retries) and failover (max_fallbacks), exactly like retry_on_429 does for 429 (429 in the list works without retry_on_429; the boolean stays for compatibility). Unset = behavior unchanged, bit-for-bit. 5xx are already retryable (listing them is a no-op); the list never affects non-status failures (customer-fixable config/credential errors stay terminal). Schema constrains entries to 400–599.

Threaded through all six dispatch sites: chat streaming + non-streaming + cross-target break, messages, responses, count_tokens — the whole routing family in one PR.

Observability: per-attempt telemetry already records each attempt's upstream status, error_class and attempt_kind (initial/retry/fallback), so a rule-triggered failover is fully attributable from existing fields; no new telemetry needed.

Interaction note: this is the current-request knob. The pre-existing per-direct-model cooldown.trigger_statuses benches a target for future requests and stays an independent layer.

Deliberately not included

  • Provider error-code / response-body matchers (the issue floats error.code: Throttling-style matching as a possible extension): status-code opt-in covers the reported scenarios; body matching adds a parsing surface per provider and can be layered on later if a concrete provider needs it.
  • Changing defaults (e.g. making 408/409 retryable out of the box, as some gateways do): the issue explicitly requires unchanged defaults.

Baseline: mainstream gateways retry 408/409/429/5xx by default and let users force retries per exception category; none offer per-status opt-in. We keep stricter defaults and add the more precise status-list control instead — providers in scope here signal overload with codes (e.g. 422) that category-based systems can't distinguish from validation errors.

Control-plane pairing

The CP is a closed schema — this field is unreachable by users until the paired AISIX-Cloud PR (cp-admin.yaml + validation + etcd projection + dashboard) lands. DP ships first because it rejects unknown fields (deny_unknown_fields); the CP PR follows immediately and carries the closing reference for AISIX-Cloud#1012.

Tests

  • Unit: listed 4xx retryable; unlisted 4xx/400 stay terminal; 429-via-list without retry_on_429; 5xx unaffected; non-status failures never resurrected; existing classification test updated to the new signature with &[] (pinning default-unchanged).
  • e2e (real aisix + etcd): two identical two-target groups — default group relays the first target's 422 without consulting the backup; fallback_on_statuses: [422] group fails over and succeeds on the backup; a 400 from a configured group (422-only list) still relays. Existing fallback/fallback-edges e2es pass unchanged.

Ref api7/AISIX-Cloud#1012 (closing reference rides the CP PR so the issue doesn't close on the DP half)

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added configurable HTTP status codes that can trigger upstream retries and failover.
    • Supports status codes from 400–599, with existing behavior preserved when unset.
    • Applied configurable fallback handling across chat, message, response, and token-counting requests.
  • Bug Fixes

    • Improved failover behavior for selected upstream client-error responses.
  • Tests

    • Added validation and end-to-end coverage for configured, unconfigured, and unlisted status codes.

…into retry/failover
Some providers use non-429 4xx codes for transient conditions (model
overload, queue full, quota exhaustion). The gateway's default treats
any non-429 4xx as the caller's error and relays it, so a model group
fails a request its other targets could have served
(AISIX-Cloud#1012).
Routing models gain an explicit opt-in list:
routing:
strategy: failover
targets: [...]
fallback_on_statuses: [408, 409, 422]
A listed status becomes retryable in is_retryable() — the single
predicate every dispatch loop consults — so it participates in both
same-target retries and failover, exactly like retry_on_429 does for
429. Unset means unchanged behavior. 5xx codes are already retryable;
listing them is a no-op. The list never affects non-status failures
(customer-fixable config/credentials stay terminal). Schema constrains
entries to 400-599.
Threaded through all six dispatch call sites (chat streaming +
non-streaming, messages, responses, count_tokens). Per-attempt
telemetry already records each attempt's upstream status and kind
(initial/retry/fallback), so a rule-triggered failover is attributable
from existing fields.
Control-plane exposure (cp-admin schema, validation, projection,
dashboard) ships as the paired AISIX-Cloud PR; this changes the DP
first so the CP projection lands against a DP that accepts the field
(deny_unknown_fields).
Ref api7/AISIX-Cloud#1012
@coderabbitai

coderabbitaiBot commented Jul 10, 2026

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds optional fallback_on_statuses routing configuration for HTTP statuses 400–599, threads it through proxy retry decisions, and validates default, configured, and non-listed status behavior.

Changes

Configurable fallback status routing

Layer / File(s)Summary
Routing configuration and schema
crates/aisix-core/src/models/routing.rs, crates/aisix-core/src/models/schema.rs, schemas/resources/*.json
Adds optional fallback_on_statuses, validates statuses from 400 through 599, and provides an empty default slice.
Retry classification
crates/aisix-proxy/src/routing.rs
Updates is_retryable to treat configured upstream statuses as retryable while preserving terminal behavior for unlisted statuses and non-status failures.
Proxy dispatch integration
crates/aisix-proxy/src/chat.rs, crates/aisix-proxy/src/count_tokens.rs, crates/aisix-proxy/src/messages.rs, crates/aisix-proxy/src/responses.rs
Passes model-specific fallback statuses into streaming, non-streaming, messages, responses, and token-count retry/failover paths.
Fallback behavior validation
tests/e2e/src/cases/fallback-on-statuses-e2e.test.ts
Verifies default terminal handling, configured 422 failover, and terminal handling for an unlisted 400 response.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
participant Client
participant ProxyDispatch
participant RoutingConfig
participant Upstream
Client->>ProxyDispatch: send request
ProxyDispatch->>RoutingConfig: read fallback_on_statuses
ProxyDispatch->>Upstream: try primary target
Upstream-->>ProxyDispatch: 422 upstream error
ProxyDispatch->>ProxyDispatch: classify status as retryable
ProxyDispatch->>Upstream: try fallback target
Upstream-->>ProxyDispatch: successful response
ProxyDispatch-->>Client: return response
Loading

Possibly related PRs

  • api7/aisix#472: Refactors target resolution in the same proxy routing dispatch paths extended here with fallback-status retry decisions.
🚥 Pre-merge checks | ✅ 6
✅ Passed checks (6 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title clearly summarizes the main routing change and the new opt-in fallback status behavior.
Linked Issues check✅ PassedCheck skipped because no linked issues were found for this pull request.
Out of Scope Changes check✅ PassedCheck skipped because no linked issues were found for this pull request.
E2e Test Quality Review✅ PassedReal E2E flow uses spawned app, etcd, and HTTP upstreams; scenarios cover default, opt-in, and unlisted statuses with clear, independent assertions.
Security Check✅ PassedNo issues found; the PR only adds an opt-in retry-status list and threads it through dispatch, with no secret/log/auth/TLS/db code touched.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/fallback-on-statuses

Comment @coderabbitai help to get the list of available commands.

…ange at the validator
Audit follow-ups: the six dispatch sites now borrow the configured
slice instead of cloning per request, and a schema-validation test
pins that out-of-range entries are rejected while in-range lists pass.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
crates/aisix-proxy/src/chat.rs (1)

1101-1114: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Duplicate routing-config extraction pattern; consider a Model-level helper.

fallback_statuses is derived with the same virtual_entry.value.routing.as_ref().map(|r| r.fallback_on_statuses_or_default()).unwrap_or(&[]) boilerplate twice in this file (streaming and non-streaming paths), and the identical pattern (for both retry_on_429 and now fallback_on_statuses) is repeated again in count_tokens.rs (and, per the graph context, in messages.rs/responses.rs). A small convenience method on Model (e.g. fallback_on_statuses_or_default(&self) -> &[u16], mirroring the existing Routing method) would collapse all these call sites to a single line and remove the risk of the branches drifting.

♻️ Proposed helper (illustrative)
implModel{pubfnfallback_on_statuses_or_default(&self) -> &[u16]{self.routing.as_ref().map(|r| r.fallback_on_statuses_or_default()).unwrap_or(&[])}}

Then each call site becomes let fallback_statuses = virtual_entry.value.fallback_on_statuses_or_default();.

Also applies to: 1822-1833

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@crates/aisix-proxy/src/chat.rs` around lines 1101 - 1114, Introduce a
Model-level helper, such as Model::fallback_on_statuses_or_default, that
delegates to routing when present and otherwise returns an empty slice. Replace
the repeated routing.as_ref().map(...).unwrap_or(&[]) extraction in both chat.rs
paths and the corresponding count_tokens.rs, messages.rs, and responses.rs call
sites with this helper, including the related retry configuration pattern where
applicable.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@crates/aisix-proxy/src/chat.rs`:
- Around line 1101-1114: Introduce a Model-level helper, such as
Model::fallback_on_statuses_or_default, that delegates to routing when present
and otherwise returns an empty slice. Replace the repeated
routing.as_ref().map(...).unwrap_or(&[]) extraction in both chat.rs paths and
the corresponding count_tokens.rs, messages.rs, and responses.rs call sites with
this helper, including the related retry configuration pattern where applicable.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: cfffe33d-ca44-4bce-af17-40c31da81ac5

📥 Commits

Reviewing files that changed from the base of the PR and between 31357d1 and 5c0dfcf.

📒 Files selected for processing (10)
  • crates/aisix-core/src/models/routing.rs
  • crates/aisix-core/src/models/schema.rs
  • crates/aisix-proxy/src/chat.rs
  • crates/aisix-proxy/src/count_tokens.rs
  • crates/aisix-proxy/src/messages.rs
  • crates/aisix-proxy/src/responses.rs
  • crates/aisix-proxy/src/routing.rs
  • schemas/resources/model.schema.json
  • schemas/resources/routing.schema.json
  • tests/e2e/src/cases/fallback-on-statuses-e2e.test.ts

@jarvis9443
jarvis9443 merged commit c750e4c into mainJul 10, 2026
12 checks passed
@jarvis9443
jarvis9443 deleted the feat/fallback-on-statuses branch July 10, 2026 12:57
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jarvis9443
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(routing): fallback_on_statuses — opt selected upstream statuses into retry/failover - #752

Merged
jarvis9443 merged 2 commits into
mainfrom
feat/fallback-on-statuses
Jul 10, 2026
Merged

feat(routing): fallback_on_statuses — opt selected upstream statuses into retry/failover#752
jarvis9443 merged 2 commits into
mainfrom
feat/fallback-on-statuses

Conversation

@jarvis9443

@jarvis9443jarvis9443 commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

Problem

Some providers use non-429 4xx codes for transient conditions — model overload, queue full, quota exhaustion. The default classification (is_retryable) treats every non-429 4xx as the caller's error and relays it, so a model group fails a request that its other targets could have served. (AISIX-Cloud#1012)

Change

Routing models gain an explicit, empty-by-default opt-in list:

routing:
strategy: failovertargets: [{ model: primary }, { model: backup }]fallback_on_statuses: [408, 409, 422]

A listed status becomes retryable in is_retryable() — the single predicate every dispatch loop consults — so it participates in both same-target retries (retries) and failover (max_fallbacks), exactly like retry_on_429 does for 429 (429 in the list works without retry_on_429; the boolean stays for compatibility). Unset = behavior unchanged, bit-for-bit. 5xx are already retryable (listing them is a no-op); the list never affects non-status failures (customer-fixable config/credential errors stay terminal). Schema constrains entries to 400–599.

Threaded through all six dispatch sites: chat streaming + non-streaming + cross-target break, messages, responses, count_tokens — the whole routing family in one PR.

Observability: per-attempt telemetry already records each attempt's upstream status, error_class and attempt_kind (initial/retry/fallback), so a rule-triggered failover is fully attributable from existing fields; no new telemetry needed.

Interaction note: this is the current-request knob. The pre-existing per-direct-model cooldown.trigger_statuses benches a target for future requests and stays an independent layer.

Deliberately not included

  • Provider error-code / response-body matchers (the issue floats error.code: Throttling-style matching as a possible extension): status-code opt-in covers the reported scenarios; body matching adds a parsing surface per provider and can be layered on later if a concrete provider needs it.
  • Changing defaults (e.g. making 408/409 retryable out of the box, as some gateways do): the issue explicitly requires unchanged defaults.

Baseline: mainstream gateways retry 408/409/429/5xx by default and let users force retries per exception category; none offer per-status opt-in. We keep stricter defaults and add the more precise status-list control instead — providers in scope here signal overload with codes (e.g. 422) that category-based systems can't distinguish from validation errors.

Control-plane pairing

The CP is a closed schema — this field is unreachable by users until the paired AISIX-Cloud PR (cp-admin.yaml + validation + etcd projection + dashboard) lands. DP ships first because it rejects unknown fields (deny_unknown_fields); the CP PR follows immediately and carries the closing reference for AISIX-Cloud#1012.

Tests

  • Unit: listed 4xx retryable; unlisted 4xx/400 stay terminal; 429-via-list without retry_on_429; 5xx unaffected; non-status failures never resurrected; existing classification test updated to the new signature with &[] (pinning default-unchanged).
  • e2e (real aisix + etcd): two identical two-target groups — default group relays the first target's 422 without consulting the backup; fallback_on_statuses: [422] group fails over and succeeds on the backup; a 400 from a configured group (422-only list) still relays. Existing fallback/fallback-edges e2es pass unchanged.

Ref api7/AISIX-Cloud#1012 (closing reference rides the CP PR so the issue doesn't close on the DP half)

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added configurable HTTP status codes that can trigger upstream retries and failover.
    • Supports status codes from 400–599, with existing behavior preserved when unset.
    • Applied configurable fallback handling across chat, message, response, and token-counting requests.
  • Bug Fixes

    • Improved failover behavior for selected upstream client-error responses.
  • Tests

    • Added validation and end-to-end coverage for configured, unconfigured, and unlisted status codes.

…into retry/failover
Some providers use non-429 4xx codes for transient conditions (model
overload, queue full, quota exhaustion). The gateway's default treats
any non-429 4xx as the caller's error and relays it, so a model group
fails a request its other targets could have served
(AISIX-Cloud#1012).
Routing models gain an explicit opt-in list:
routing:
strategy: failover
targets: [...]
fallback_on_statuses: [408, 409, 422]
A listed status becomes retryable in is_retryable() — the single
predicate every dispatch loop consults — so it participates in both
same-target retries and failover, exactly like retry_on_429 does for
429. Unset means unchanged behavior. 5xx codes are already retryable;
listing them is a no-op. The list never affects non-status failures
(customer-fixable config/credentials stay terminal). Schema constrains
entries to 400-599.
Threaded through all six dispatch call sites (chat streaming +
non-streaming, messages, responses, count_tokens). Per-attempt
telemetry already records each attempt's upstream status and kind
(initial/retry/fallback), so a rule-triggered failover is attributable
from existing fields.
Control-plane exposure (cp-admin schema, validation, projection,
dashboard) ships as the paired AISIX-Cloud PR; this changes the DP
first so the CP projection lands against a DP that accepts the field
(deny_unknown_fields).
Ref api7/AISIX-Cloud#1012
@coderabbitai

coderabbitaiBot commented Jul 10, 2026

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds optional fallback_on_statuses routing configuration for HTTP statuses 400–599, threads it through proxy retry decisions, and validates default, configured, and non-listed status behavior.

Changes

Configurable fallback status routing

Layer / File(s)Summary
Routing configuration and schema
crates/aisix-core/src/models/routing.rs, crates/aisix-core/src/models/schema.rs, schemas/resources/*.json
Adds optional fallback_on_statuses, validates statuses from 400 through 599, and provides an empty default slice.
Retry classification
crates/aisix-proxy/src/routing.rs
Updates is_retryable to treat configured upstream statuses as retryable while preserving terminal behavior for unlisted statuses and non-status failures.
Proxy dispatch integration
crates/aisix-proxy/src/chat.rs, crates/aisix-proxy/src/count_tokens.rs, crates/aisix-proxy/src/messages.rs, crates/aisix-proxy/src/responses.rs
Passes model-specific fallback statuses into streaming, non-streaming, messages, responses, and token-count retry/failover paths.
Fallback behavior validation
tests/e2e/src/cases/fallback-on-statuses-e2e.test.ts
Verifies default terminal handling, configured 422 failover, and terminal handling for an unlisted 400 response.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
participant Client
participant ProxyDispatch
participant RoutingConfig
participant Upstream
Client->>ProxyDispatch: send request
ProxyDispatch->>RoutingConfig: read fallback_on_statuses
ProxyDispatch->>Upstream: try primary target
Upstream-->>ProxyDispatch: 422 upstream error
ProxyDispatch->>ProxyDispatch: classify status as retryable
ProxyDispatch->>Upstream: try fallback target
Upstream-->>ProxyDispatch: successful response
ProxyDispatch-->>Client: return response
Loading

Possibly related PRs

  • api7/aisix#472: Refactors target resolution in the same proxy routing dispatch paths extended here with fallback-status retry decisions.
🚥 Pre-merge checks | ✅ 6
✅ Passed checks (6 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title clearly summarizes the main routing change and the new opt-in fallback status behavior.
Linked Issues check✅ PassedCheck skipped because no linked issues were found for this pull request.
Out of Scope Changes check✅ PassedCheck skipped because no linked issues were found for this pull request.
E2e Test Quality Review✅ PassedReal E2E flow uses spawned app, etcd, and HTTP upstreams; scenarios cover default, opt-in, and unlisted statuses with clear, independent assertions.
Security Check✅ PassedNo issues found; the PR only adds an opt-in retry-status list and threads it through dispatch, with no secret/log/auth/TLS/db code touched.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/fallback-on-statuses

Comment @coderabbitai help to get the list of available commands.

…ange at the validator
Audit follow-ups: the six dispatch sites now borrow the configured
slice instead of cloning per request, and a schema-validation test
pins that out-of-range entries are rejected while in-range lists pass.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
crates/aisix-proxy/src/chat.rs (1)

1101-1114: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Duplicate routing-config extraction pattern; consider a Model-level helper.

fallback_statuses is derived with the same virtual_entry.value.routing.as_ref().map(|r| r.fallback_on_statuses_or_default()).unwrap_or(&[]) boilerplate twice in this file (streaming and non-streaming paths), and the identical pattern (for both retry_on_429 and now fallback_on_statuses) is repeated again in count_tokens.rs (and, per the graph context, in messages.rs/responses.rs). A small convenience method on Model (e.g. fallback_on_statuses_or_default(&self) -> &[u16], mirroring the existing Routing method) would collapse all these call sites to a single line and remove the risk of the branches drifting.

♻️ Proposed helper (illustrative)
implModel{pubfnfallback_on_statuses_or_default(&self) -> &[u16]{self.routing.as_ref().map(|r| r.fallback_on_statuses_or_default()).unwrap_or(&[])}}

Then each call site becomes let fallback_statuses = virtual_entry.value.fallback_on_statuses_or_default();.

Also applies to: 1822-1833

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@crates/aisix-proxy/src/chat.rs` around lines 1101 - 1114, Introduce a
Model-level helper, such as Model::fallback_on_statuses_or_default, that
delegates to routing when present and otherwise returns an empty slice. Replace
the repeated routing.as_ref().map(...).unwrap_or(&[]) extraction in both chat.rs
paths and the corresponding count_tokens.rs, messages.rs, and responses.rs call
sites with this helper, including the related retry configuration pattern where
applicable.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@crates/aisix-proxy/src/chat.rs`:
- Around line 1101-1114: Introduce a Model-level helper, such as
Model::fallback_on_statuses_or_default, that delegates to routing when present
and otherwise returns an empty slice. Replace the repeated
routing.as_ref().map(...).unwrap_or(&[]) extraction in both chat.rs paths and
the corresponding count_tokens.rs, messages.rs, and responses.rs call sites with
this helper, including the related retry configuration pattern where applicable.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: cfffe33d-ca44-4bce-af17-40c31da81ac5

📥 Commits

Reviewing files that changed from the base of the PR and between 31357d1 and 5c0dfcf.

📒 Files selected for processing (10)
  • crates/aisix-core/src/models/routing.rs
  • crates/aisix-core/src/models/schema.rs
  • crates/aisix-proxy/src/chat.rs
  • crates/aisix-proxy/src/count_tokens.rs
  • crates/aisix-proxy/src/messages.rs
  • crates/aisix-proxy/src/responses.rs
  • crates/aisix-proxy/src/routing.rs
  • schemas/resources/model.schema.json
  • schemas/resources/routing.schema.json
  • tests/e2e/src/cases/fallback-on-statuses-e2e.test.ts

@jarvis9443
jarvis9443 merged commit c750e4c into mainJul 10, 2026
12 checks passed
@jarvis9443
jarvis9443 deleted the feat/fallback-on-statuses branch July 10, 2026 12:57
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jarvis9443
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(routing): fallback_on_statuses — opt selected upstream statuses into retry/failover - #752

Merged
jarvis9443 merged 2 commits into
mainfrom
feat/fallback-on-statuses
Jul 10, 2026
Merged

feat(routing): fallback_on_statuses — opt selected upstream statuses into retry/failover#752
jarvis9443 merged 2 commits into
mainfrom
feat/fallback-on-statuses

Conversation

@jarvis9443

@jarvis9443jarvis9443 commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

Problem

Some providers use non-429 4xx codes for transient conditions — model overload, queue full, quota exhaustion. The default classification (is_retryable) treats every non-429 4xx as the caller's error and relays it, so a model group fails a request that its other targets could have served. (AISIX-Cloud#1012)

Change

Routing models gain an explicit, empty-by-default opt-in list:

routing:
strategy: failovertargets: [{ model: primary }, { model: backup }]fallback_on_statuses: [408, 409, 422]

A listed status becomes retryable in is_retryable() — the single predicate every dispatch loop consults — so it participates in both same-target retries (retries) and failover (max_fallbacks), exactly like retry_on_429 does for 429 (429 in the list works without retry_on_429; the boolean stays for compatibility). Unset = behavior unchanged, bit-for-bit. 5xx are already retryable (listing them is a no-op); the list never affects non-status failures (customer-fixable config/credential errors stay terminal). Schema constrains entries to 400–599.

Threaded through all six dispatch sites: chat streaming + non-streaming + cross-target break, messages, responses, count_tokens — the whole routing family in one PR.

Observability: per-attempt telemetry already records each attempt's upstream status, error_class and attempt_kind (initial/retry/fallback), so a rule-triggered failover is fully attributable from existing fields; no new telemetry needed.

Interaction note: this is the current-request knob. The pre-existing per-direct-model cooldown.trigger_statuses benches a target for future requests and stays an independent layer.

Deliberately not included

  • Provider error-code / response-body matchers (the issue floats error.code: Throttling-style matching as a possible extension): status-code opt-in covers the reported scenarios; body matching adds a parsing surface per provider and can be layered on later if a concrete provider needs it.
  • Changing defaults (e.g. making 408/409 retryable out of the box, as some gateways do): the issue explicitly requires unchanged defaults.

Baseline: mainstream gateways retry 408/409/429/5xx by default and let users force retries per exception category; none offer per-status opt-in. We keep stricter defaults and add the more precise status-list control instead — providers in scope here signal overload with codes (e.g. 422) that category-based systems can't distinguish from validation errors.

Control-plane pairing

The CP is a closed schema — this field is unreachable by users until the paired AISIX-Cloud PR (cp-admin.yaml + validation + etcd projection + dashboard) lands. DP ships first because it rejects unknown fields (deny_unknown_fields); the CP PR follows immediately and carries the closing reference for AISIX-Cloud#1012.

Tests

  • Unit: listed 4xx retryable; unlisted 4xx/400 stay terminal; 429-via-list without retry_on_429; 5xx unaffected; non-status failures never resurrected; existing classification test updated to the new signature with &[] (pinning default-unchanged).
  • e2e (real aisix + etcd): two identical two-target groups — default group relays the first target's 422 without consulting the backup; fallback_on_statuses: [422] group fails over and succeeds on the backup; a 400 from a configured group (422-only list) still relays. Existing fallback/fallback-edges e2es pass unchanged.

Ref api7/AISIX-Cloud#1012 (closing reference rides the CP PR so the issue doesn't close on the DP half)

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added configurable HTTP status codes that can trigger upstream retries and failover.
    • Supports status codes from 400–599, with existing behavior preserved when unset.
    • Applied configurable fallback handling across chat, message, response, and token-counting requests.
  • Bug Fixes

    • Improved failover behavior for selected upstream client-error responses.
  • Tests

    • Added validation and end-to-end coverage for configured, unconfigured, and unlisted status codes.

…into retry/failover
Some providers use non-429 4xx codes for transient conditions (model
overload, queue full, quota exhaustion). The gateway's default treats
any non-429 4xx as the caller's error and relays it, so a model group
fails a request its other targets could have served
(AISIX-Cloud#1012).
Routing models gain an explicit opt-in list:
routing:
strategy: failover
targets: [...]
fallback_on_statuses: [408, 409, 422]
A listed status becomes retryable in is_retryable() — the single
predicate every dispatch loop consults — so it participates in both
same-target retries and failover, exactly like retry_on_429 does for
429. Unset means unchanged behavior. 5xx codes are already retryable;
listing them is a no-op. The list never affects non-status failures
(customer-fixable config/credentials stay terminal). Schema constrains
entries to 400-599.
Threaded through all six dispatch call sites (chat streaming +
non-streaming, messages, responses, count_tokens). Per-attempt
telemetry already records each attempt's upstream status and kind
(initial/retry/fallback), so a rule-triggered failover is attributable
from existing fields.
Control-plane exposure (cp-admin schema, validation, projection,
dashboard) ships as the paired AISIX-Cloud PR; this changes the DP
first so the CP projection lands against a DP that accepts the field
(deny_unknown_fields).
Ref api7/AISIX-Cloud#1012
@coderabbitai

coderabbitaiBot commented Jul 10, 2026

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds optional fallback_on_statuses routing configuration for HTTP statuses 400–599, threads it through proxy retry decisions, and validates default, configured, and non-listed status behavior.

Changes

Configurable fallback status routing

Layer / File(s)Summary
Routing configuration and schema
crates/aisix-core/src/models/routing.rs, crates/aisix-core/src/models/schema.rs, schemas/resources/*.json
Adds optional fallback_on_statuses, validates statuses from 400 through 599, and provides an empty default slice.
Retry classification
crates/aisix-proxy/src/routing.rs
Updates is_retryable to treat configured upstream statuses as retryable while preserving terminal behavior for unlisted statuses and non-status failures.
Proxy dispatch integration
crates/aisix-proxy/src/chat.rs, crates/aisix-proxy/src/count_tokens.rs, crates/aisix-proxy/src/messages.rs, crates/aisix-proxy/src/responses.rs
Passes model-specific fallback statuses into streaming, non-streaming, messages, responses, and token-count retry/failover paths.
Fallback behavior validation
tests/e2e/src/cases/fallback-on-statuses-e2e.test.ts
Verifies default terminal handling, configured 422 failover, and terminal handling for an unlisted 400 response.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
participant Client
participant ProxyDispatch
participant RoutingConfig
participant Upstream
Client->>ProxyDispatch: send request
ProxyDispatch->>RoutingConfig: read fallback_on_statuses
ProxyDispatch->>Upstream: try primary target
Upstream-->>ProxyDispatch: 422 upstream error
ProxyDispatch->>ProxyDispatch: classify status as retryable
ProxyDispatch->>Upstream: try fallback target
Upstream-->>ProxyDispatch: successful response
ProxyDispatch-->>Client: return response
Loading

Possibly related PRs

  • api7/aisix#472: Refactors target resolution in the same proxy routing dispatch paths extended here with fallback-status retry decisions.
🚥 Pre-merge checks | ✅ 6
✅ Passed checks (6 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title clearly summarizes the main routing change and the new opt-in fallback status behavior.
Linked Issues check✅ PassedCheck skipped because no linked issues were found for this pull request.
Out of Scope Changes check✅ PassedCheck skipped because no linked issues were found for this pull request.
E2e Test Quality Review✅ PassedReal E2E flow uses spawned app, etcd, and HTTP upstreams; scenarios cover default, opt-in, and unlisted statuses with clear, independent assertions.
Security Check✅ PassedNo issues found; the PR only adds an opt-in retry-status list and threads it through dispatch, with no secret/log/auth/TLS/db code touched.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/fallback-on-statuses

Comment @coderabbitai help to get the list of available commands.

…ange at the validator
Audit follow-ups: the six dispatch sites now borrow the configured
slice instead of cloning per request, and a schema-validation test
pins that out-of-range entries are rejected while in-range lists pass.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
crates/aisix-proxy/src/chat.rs (1)

1101-1114: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Duplicate routing-config extraction pattern; consider a Model-level helper.

fallback_statuses is derived with the same virtual_entry.value.routing.as_ref().map(|r| r.fallback_on_statuses_or_default()).unwrap_or(&[]) boilerplate twice in this file (streaming and non-streaming paths), and the identical pattern (for both retry_on_429 and now fallback_on_statuses) is repeated again in count_tokens.rs (and, per the graph context, in messages.rs/responses.rs). A small convenience method on Model (e.g. fallback_on_statuses_or_default(&self) -> &[u16], mirroring the existing Routing method) would collapse all these call sites to a single line and remove the risk of the branches drifting.

♻️ Proposed helper (illustrative)
implModel{pubfnfallback_on_statuses_or_default(&self) -> &[u16]{self.routing.as_ref().map(|r| r.fallback_on_statuses_or_default()).unwrap_or(&[])}}

Then each call site becomes let fallback_statuses = virtual_entry.value.fallback_on_statuses_or_default();.

Also applies to: 1822-1833

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@crates/aisix-proxy/src/chat.rs` around lines 1101 - 1114, Introduce a
Model-level helper, such as Model::fallback_on_statuses_or_default, that
delegates to routing when present and otherwise returns an empty slice. Replace
the repeated routing.as_ref().map(...).unwrap_or(&[]) extraction in both chat.rs
paths and the corresponding count_tokens.rs, messages.rs, and responses.rs call
sites with this helper, including the related retry configuration pattern where
applicable.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@crates/aisix-proxy/src/chat.rs`:
- Around line 1101-1114: Introduce a Model-level helper, such as
Model::fallback_on_statuses_or_default, that delegates to routing when present
and otherwise returns an empty slice. Replace the repeated
routing.as_ref().map(...).unwrap_or(&[]) extraction in both chat.rs paths and
the corresponding count_tokens.rs, messages.rs, and responses.rs call sites with
this helper, including the related retry configuration pattern where applicable.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: cfffe33d-ca44-4bce-af17-40c31da81ac5

📥 Commits

Reviewing files that changed from the base of the PR and between 31357d1 and 5c0dfcf.

📒 Files selected for processing (10)
  • crates/aisix-core/src/models/routing.rs
  • crates/aisix-core/src/models/schema.rs
  • crates/aisix-proxy/src/chat.rs
  • crates/aisix-proxy/src/count_tokens.rs
  • crates/aisix-proxy/src/messages.rs
  • crates/aisix-proxy/src/responses.rs
  • crates/aisix-proxy/src/routing.rs
  • schemas/resources/model.schema.json
  • schemas/resources/routing.schema.json
  • tests/e2e/src/cases/fallback-on-statuses-e2e.test.ts

@jarvis9443
jarvis9443 merged commit c750e4c into mainJul 10, 2026
12 checks passed
@jarvis9443
jarvis9443 deleted the feat/fallback-on-statuses branch July 10, 2026 12:57
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jarvis9443