Skip to content

fix(dev): restore fallback capacity for canonical E2E chat verification (#345) - #468

Merged
sheepdestroyer merged 3 commits into
masterfrom
fix/dev-fallback-capacity-verification
Aug 13, 2026
Merged

fix(dev): restore fallback capacity for canonical E2E chat verification (#345)#468
sheepdestroyer merged 3 commits into
masterfrom
fix/dev-fallback-capacity-verification

Conversation

@sheepdestroyer

@sheepdestroyersheepdestroyer commented Aug 12, 2026

Copy link
Copy Markdown
Owner

Closes#345

Summary by Sourcery

Harden routing and verification infrastructure around LiteLLM/OpenRouter and Langfuse, improving auth, key handling, and canonical endpoint behavior.

Bug Fixes:

  • Restore localhost HTTP fallbacks and explicit LLAMA_* environment variable handling when resolving llama server/classifier endpoints.
  • Ensure LiteLLM master key placeholders and misconfigurations fail fast instead of silently proxying with invalid credentials.
  • Fix adaptive router roster purge assertion to accept multiple purge calls.

Enhancements:

  • Enforce bearer client authentication for the /v1/responses endpoint.
  • Introduce centralized validation of the LiteLLM master key and reuse it for proxy execution and internal calls.
  • Register OpenRouter models as LiteLLM DB models on startup, including dynamic loading from config.yaml and static fallback definitions.
  • Improve Langfuse session propagation tracing to support v4 events_only mode and longer polling, and propagate session/user IDs into Langfuse observations and updates.
  • Extend start-stack.sh to export LLAMA_* env vars and render LiteLLM config using derived classifier URLs.

Tests:

  • Update existing routing, lifespan, and responses_api tests to respect new auth and master key validation behavior.
  • Add tests covering OpenRouter model DB registration paths (no key, static fallback, and config-driven) and LiteLLM master key validation failure modes.
  • Add verification coverage for Langfuse events_only mode and non-leak behavior in session propagation checks.

@sourcery-ai

sourcery-aiBot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Reviewer's Guide

Restores and hardens environment-based configuration and verification for LiteLLM/OpenRouter routing, adds strict master-key and client auth validation, improves Langfuse tracing/session propagation, and extends tests/startup scripts to cover the new behavior and canonical endpoint derivation paths.

Sequence diagram for responses_api client auth and master key validation

sequenceDiagram
actor Client
participant Router as responses_api
participant Env as _validate_litellm_master_key
participant LiteLLM as litellm_proxy
Client->>Router: POST /responses (Authorization: Bearer client_token)
Router->>Router: validate Authorization header
Router->>Env: _validate_litellm_master_key()
Env-->>Router: master_key
Router->>LiteLLM: client.post /responses (Authorization: Bearer master_key)
LiteLLM-->>Router: response
Router-->>Client: routed response
Loading

File-Level Changes

ChangeDetailsFiles
Make LLAMA server/classifier endpoint resolution respect explicit environment variables while preserving localhost fallbacks and canonical HTTPS derivation.
  • Introduce LLAMA_SERVER_URL and LLAMA_CLASSIFIER_URL environment overrides when resolving endpoints.
  • Fallback to config values or localhost defaults when env vars are absent.
  • Prefer explicit env URLs over canonical HTTPS derivation in resolution logic.
router/main.py
start-stack.sh
Introduce centralized LiteLLM master key validation and stricter usage across responses and proxy paths.
  • Add _INVALID_MASTER_KEYS set and _validate_litellm_master_key helper that fails fast with HTTP 500 on missing/placeholder keys.
  • Use _validate_litellm_master_key in /v1/responses and execute_proxy instead of raw os.getenv lookups.
  • Update tests to patch LITELLM_MASTER_KEY and cover invalid-key behavior.
router/main.py
router/tests/test_responses_api.py
router/tests/test_routing_behavior.py
tests/conftest.py
Add OpenRouter model DB registration on startup with dynamic config loading, static fallback, and stale deployment purge.
  • Implement _register_openrouter_models_in_db that loads OpenRouter models from litellm config paths or falls back to a static openrouter-auto definition.
  • Purge stale openrouter-* deployments via _purge_stale_deployments before re-registering.
  • Wire OpenRouter registration into FastAPI lifespan and add unit tests for config-based and fallback registration paths.
router/main.py
router/tests/test_lifespan.py
router/tests/test_register_openrouter_models_in_db.py
router/tests/test_sync_adaptive_router_roster.py
Enforce client Bearer authentication and improve Langfuse tracing/session propagation behavior in responses and chat triage.
  • Require a non-empty Bearer Authorization header in responses_api, returning 401 on missing/invalid auth.
  • Pass session_id and user_id into Langfuse start_observation and update calls when available.
  • Extend Langfuse verification script to handle v4 events_only mode, longer polling, and revised leak checks.
  • Update tests to include Authorization headers and expectations for auth enforcement and tracing behavior.
router/main.py
scripts/verification/verify_canonical_endpoints.py
router/tests/test_responses_api.py
Tighten test and startup environment configuration around routing behavior and master key usage.
  • Set default LITELLM_MASTER_KEY and ROUTER_API_KEY in test conftest and specific routing tests.
  • Adjust routing behavior tests to provide Authorization headers and patched master keys.
  • Update start-stack.sh to export LLAMA_SERVER_URL/LLAMA_CLASSIFIER_URL and derive classifier URL in rendered LiteLLM config.
  • Relax sync_adaptive_router_roster purge assertion to allow multiple purge calls.
router/tests/test_routing_behavior.py
tests/conftest.py
start-stack.sh
router/tests/test_sync_adaptive_router_roster.py

Assessment against linked issues

IssueObjectiveAddressedExplanation
#345Restore functional OpenRouter-backed free-tier chat routing for dev canonical verification, including robust LiteLLM master key handling and model registration so auto/free routes (e.g. llm-routing-auto-free, agent-simple-core) can succeed.
#345Fix dev LiteLLM local safety-net routing by correctly deriving and wiring LLAMA_SERVER_URL / LLAMA_CLASSIFIER_URL from configuration/start-stack into router, so the local fallback classifier becomes reachable from the dev container without coupling to a broken production hostname.

Possibly linked issues


Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@coderabbitai

coderabbitaiBot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@sheepdestroyer, you've reached your PR review limit, so we couldn't start this review.

Next review available in:117 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 3f27a272-9764-4dc3-8e95-2f0a8a765e1e

📥 Commits

Reviewing files that changed from the base of the PR and between 830008b and 4e11105.

📒 Files selected for processing (10)
  • .env.dev
  • router/main.py
  • router/tests/test_lifespan.py
  • router/tests/test_register_openrouter_models_in_db.py
  • router/tests/test_responses_api.py
  • router/tests/test_routing_behavior.py
  • router/tests/test_sync_adaptive_router_roster.py
  • scripts/verification/verify_canonical_endpoints.py
  • start-stack.sh
  • tests/conftest.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sourcery-aisourcery-aiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 2 issues, and left some high level feedback:

  • Avoid logging the raw LITELLM_MASTER_KEY value in _validate_litellm_master_key, as the current logger.error call can expose secrets in logs; consider logging only that it is missing/invalid or a redacted version.
  • The client auth check in responses_api only validates the presence and shape of a Bearer token; if stronger authentication/authorization is expected, consider integrating actual token validation or clarifying that this endpoint is only gating on header presence.
Prompt for AI Agents
Please address the comments from this code review:
## Overall Comments- Avoid logging the raw LITELLM_MASTER_KEY value in `_validate_litellm_master_key`, as the current `logger.error` call can expose secrets in logs; consider logging only that it is missing/invalid or a redacted version.
- The client auth check in `responses_api` only validates the presence and shape of a Bearer token; if stronger authentication/authorization is expected, consider integrating actual token validation or clarifying that this endpoint is only gating on header presence.
## Individual Comments### Comment 1
<locationpath="router/main.py"line_range="797-806" />
<code_context>
+}
++
+def _validate_litellm_master_key() -> str:
+ """Validate LITELLM_MASTER_KEY environment variable.
++ Returns:
+ The valid master key string.++ Raises:
+ HTTPException(500): If master key is missing, empty, or placeholder string.+ """
+ key = (os.getenv("LITELLM_MASTER_KEY") or "").strip()
+ if not key or key in _INVALID_MASTER_KEYS or "PLACEHOLDER" in key.upper():
+ logger.error(f"Invalid or missing LITELLM_MASTER_KEY: '{key}'")
</code_context>
<issue_to_address>
**🚨 issue (security):** Avoid logging the raw master key value to prevent leaking secrets.
The `logger.error` line logs the full `LITELLM_MASTER_KEY`, which risks exposing a production secret in logs. Please avoid logging the raw key—log only that it is invalid/missing, or mask it (e.g., show a short prefix or just its length) while preserving enough context for debugging.
</issue_to_address>
### Comment 2
<locationpath="router/main.py"line_range="2306-2310" />
<code_context>
when an auto model (e.g. llm-routing-auto-free) is requested, while supporting model aliases
(such as gpt-4o-mini, local-qwen-3.6-hass) and tool/streaming executions.
"""
+# Enforce client authentication+ auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
+ if not auth_header or not auth_header.startswith("Bearer "):
+ raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")+ client_token = auth_header[7:].strip()
+ if not client_token:
+ raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
</code_context>
<issue_to_address>
**suggestion (bug_risk):** Authorization handling is case-sensitive and extracts a token that is never used.
`auth_header.startswith("Bearer ")` will reject headers with different casing or spacing (e.g. `bearer <token>`). Consider normalizing (e.g. lowercasing and/or splitting on whitespace) before checking to make the auth handling more robust. Also, `client_token` is extracted but never used; either pass it to downstream logic where needed or remove it to avoid confusion.
Suggested implementation:
```pythontry:
await _register_ollama_models_in_db(litellm_master_key)
exceptExceptionas e:
when an auto model (e.g. llm-routing-auto-free) is requested, while supporting model aliases
(such as gpt-4o-mini, local-qwen-3.6-hass) and tool/streaming executions.
""" # Enforce client authentication auth_header = request.headers.get("Authorization") or request.headers.get("authorization") if not auth_header: raise HTTPException(status_code=401, detail="Missing or invalid Authorization header") parts = auth_header.split() if len(parts) != 2 or parts[0].lower() != "bearer" or not parts[1]: raise HTTPException(status_code=401, detail="Missing or invalid Authorization header") # Store the normalized client token for downstream use request.state.client_token = parts[1] try:```1. Anywhere downstream that needs the client token should read it from `request.state.client_token` instead of re-parsing headers.2. If your router uses dependency injection (e.g. FastAPI dependencies) for auth, consider integrating this parsing into a shared dependency to avoid duplication.</issue_to_address>

Sourcery is free for open source - if you like our reviews please consider sharing them ✨
Help me be more useful! Please click 👍 or 👎 on each comment and I'll use the feedback to improve your reviews.

Comment threadrouter/main.py
Comment on lines +797 to +806
def _validate_litellm_master_key() -> str:
"""Validate LITELLM_MASTER_KEY environment variable.

Returns:
The valid master key string.

Raises:
HTTPException(500): If master key is missing, empty, or placeholder string.
"""
key = (os.getenv("LITELLM_MASTER_KEY") or "").strip()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🚨 issue (security): Avoid logging the raw master key value to prevent leaking secrets.

The logger.error line logs the full LITELLM_MASTER_KEY, which risks exposing a production secret in logs. Please avoid logging the raw key—log only that it is invalid/missing, or mask it (e.g., show a short prefix or just its length) while preserving enough context for debugging.

Comment threadrouter/main.py
Comment on lines +2306 to +2310
# Enforce client authentication
auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
if not auth_header or not auth_header.startswith("Bearer "):
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
client_token = auth_header[7:].strip()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

suggestion (bug_risk): Authorization handling is case-sensitive and extracts a token that is never used.

auth_header.startswith("Bearer ") will reject headers with different casing or spacing (e.g. bearer <token>). Consider normalizing (e.g. lowercasing and/or splitting on whitespace) before checking to make the auth handling more robust. Also, client_token is extracted but never used; either pass it to downstream logic where needed or remove it to avoid confusion.

Suggested implementation:

try:
await_register_ollama_models_in_db(litellm_master_key)
exceptExceptionase:
whenanautomodel (e.g. llm-routing-auto-free) isrequested, whilesupportingmodelaliases
(suchasgpt-4o-mini, local-qwen-3.6-hass) andtool/streamingexecutions.
"""
# Enforce client authentication
auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
if not auth_header:
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
parts = auth_header.split()
if len(parts) != 2 or parts[0].lower() != "bearer" or not parts[1]:
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
# Store the normalized client token for downstream userequest.state.client_token=parts[1]
try:
  1. Anywhere downstream that needs the client token should read it from request.state.client_token instead of re-parsing headers.
  2. If your router uses dependency injection (e.g. FastAPI dependencies) for auth, consider integrating this parsing into a shared dependency to avoid duplication.

@sheepdestroyer
sheepdestroyer merged commit 105e7b2 into masterAug 13, 2026
6 of 7 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(dev): restore fallback capacity for canonical E2E chat verification

1 participant

@sheepdestroyer
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
fix(dev): restore fallback capacity for canonical E2E chat verification (#345) by sheepdestroyer · Pull Request #468 · sheepdestroyer/LLM-Routing · GitHub
Skip to content

fix(dev): restore fallback capacity for canonical E2E chat verification (#345) - #468

Merged
sheepdestroyer merged 3 commits into
masterfrom
fix/dev-fallback-capacity-verification
Aug 13, 2026
Merged

fix(dev): restore fallback capacity for canonical E2E chat verification (#345)#468
sheepdestroyer merged 3 commits into
masterfrom
fix/dev-fallback-capacity-verification

Conversation

@sheepdestroyer

@sheepdestroyersheepdestroyer commented Aug 12, 2026

Copy link
Copy Markdown
Owner

Closes#345

Summary by Sourcery

Harden routing and verification infrastructure around LiteLLM/OpenRouter and Langfuse, improving auth, key handling, and canonical endpoint behavior.

Bug Fixes:

  • Restore localhost HTTP fallbacks and explicit LLAMA_* environment variable handling when resolving llama server/classifier endpoints.
  • Ensure LiteLLM master key placeholders and misconfigurations fail fast instead of silently proxying with invalid credentials.
  • Fix adaptive router roster purge assertion to accept multiple purge calls.

Enhancements:

  • Enforce bearer client authentication for the /v1/responses endpoint.
  • Introduce centralized validation of the LiteLLM master key and reuse it for proxy execution and internal calls.
  • Register OpenRouter models as LiteLLM DB models on startup, including dynamic loading from config.yaml and static fallback definitions.
  • Improve Langfuse session propagation tracing to support v4 events_only mode and longer polling, and propagate session/user IDs into Langfuse observations and updates.
  • Extend start-stack.sh to export LLAMA_* env vars and render LiteLLM config using derived classifier URLs.

Tests:

  • Update existing routing, lifespan, and responses_api tests to respect new auth and master key validation behavior.
  • Add tests covering OpenRouter model DB registration paths (no key, static fallback, and config-driven) and LiteLLM master key validation failure modes.
  • Add verification coverage for Langfuse events_only mode and non-leak behavior in session propagation checks.

@sourcery-ai

sourcery-aiBot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Reviewer's Guide

Restores and hardens environment-based configuration and verification for LiteLLM/OpenRouter routing, adds strict master-key and client auth validation, improves Langfuse tracing/session propagation, and extends tests/startup scripts to cover the new behavior and canonical endpoint derivation paths.

Sequence diagram for responses_api client auth and master key validation

sequenceDiagram
actor Client
participant Router as responses_api
participant Env as _validate_litellm_master_key
participant LiteLLM as litellm_proxy
Client->>Router: POST /responses (Authorization: Bearer client_token)
Router->>Router: validate Authorization header
Router->>Env: _validate_litellm_master_key()
Env-->>Router: master_key
Router->>LiteLLM: client.post /responses (Authorization: Bearer master_key)
LiteLLM-->>Router: response
Router-->>Client: routed response
Loading

File-Level Changes

ChangeDetailsFiles
Make LLAMA server/classifier endpoint resolution respect explicit environment variables while preserving localhost fallbacks and canonical HTTPS derivation.
  • Introduce LLAMA_SERVER_URL and LLAMA_CLASSIFIER_URL environment overrides when resolving endpoints.
  • Fallback to config values or localhost defaults when env vars are absent.
  • Prefer explicit env URLs over canonical HTTPS derivation in resolution logic.
router/main.py
start-stack.sh
Introduce centralized LiteLLM master key validation and stricter usage across responses and proxy paths.
  • Add _INVALID_MASTER_KEYS set and _validate_litellm_master_key helper that fails fast with HTTP 500 on missing/placeholder keys.
  • Use _validate_litellm_master_key in /v1/responses and execute_proxy instead of raw os.getenv lookups.
  • Update tests to patch LITELLM_MASTER_KEY and cover invalid-key behavior.
router/main.py
router/tests/test_responses_api.py
router/tests/test_routing_behavior.py
tests/conftest.py
Add OpenRouter model DB registration on startup with dynamic config loading, static fallback, and stale deployment purge.
  • Implement _register_openrouter_models_in_db that loads OpenRouter models from litellm config paths or falls back to a static openrouter-auto definition.
  • Purge stale openrouter-* deployments via _purge_stale_deployments before re-registering.
  • Wire OpenRouter registration into FastAPI lifespan and add unit tests for config-based and fallback registration paths.
router/main.py
router/tests/test_lifespan.py
router/tests/test_register_openrouter_models_in_db.py
router/tests/test_sync_adaptive_router_roster.py
Enforce client Bearer authentication and improve Langfuse tracing/session propagation behavior in responses and chat triage.
  • Require a non-empty Bearer Authorization header in responses_api, returning 401 on missing/invalid auth.
  • Pass session_id and user_id into Langfuse start_observation and update calls when available.
  • Extend Langfuse verification script to handle v4 events_only mode, longer polling, and revised leak checks.
  • Update tests to include Authorization headers and expectations for auth enforcement and tracing behavior.
router/main.py
scripts/verification/verify_canonical_endpoints.py
router/tests/test_responses_api.py
Tighten test and startup environment configuration around routing behavior and master key usage.
  • Set default LITELLM_MASTER_KEY and ROUTER_API_KEY in test conftest and specific routing tests.
  • Adjust routing behavior tests to provide Authorization headers and patched master keys.
  • Update start-stack.sh to export LLAMA_SERVER_URL/LLAMA_CLASSIFIER_URL and derive classifier URL in rendered LiteLLM config.
  • Relax sync_adaptive_router_roster purge assertion to allow multiple purge calls.
router/tests/test_routing_behavior.py
tests/conftest.py
start-stack.sh
router/tests/test_sync_adaptive_router_roster.py

Assessment against linked issues

IssueObjectiveAddressedExplanation
#345Restore functional OpenRouter-backed free-tier chat routing for dev canonical verification, including robust LiteLLM master key handling and model registration so auto/free routes (e.g. llm-routing-auto-free, agent-simple-core) can succeed.
#345Fix dev LiteLLM local safety-net routing by correctly deriving and wiring LLAMA_SERVER_URL / LLAMA_CLASSIFIER_URL from configuration/start-stack into router, so the local fallback classifier becomes reachable from the dev container without coupling to a broken production hostname.

Possibly linked issues


Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@coderabbitai

coderabbitaiBot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@sheepdestroyer, you've reached your PR review limit, so we couldn't start this review.

Next review available in:117 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 3f27a272-9764-4dc3-8e95-2f0a8a765e1e

📥 Commits

Reviewing files that changed from the base of the PR and between 830008b and 4e11105.

📒 Files selected for processing (10)
  • .env.dev
  • router/main.py
  • router/tests/test_lifespan.py
  • router/tests/test_register_openrouter_models_in_db.py
  • router/tests/test_responses_api.py
  • router/tests/test_routing_behavior.py
  • router/tests/test_sync_adaptive_router_roster.py
  • scripts/verification/verify_canonical_endpoints.py
  • start-stack.sh
  • tests/conftest.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sourcery-aisourcery-aiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 2 issues, and left some high level feedback:

  • Avoid logging the raw LITELLM_MASTER_KEY value in _validate_litellm_master_key, as the current logger.error call can expose secrets in logs; consider logging only that it is missing/invalid or a redacted version.
  • The client auth check in responses_api only validates the presence and shape of a Bearer token; if stronger authentication/authorization is expected, consider integrating actual token validation or clarifying that this endpoint is only gating on header presence.
Prompt for AI Agents
Please address the comments from this code review:
## Overall Comments- Avoid logging the raw LITELLM_MASTER_KEY value in `_validate_litellm_master_key`, as the current `logger.error` call can expose secrets in logs; consider logging only that it is missing/invalid or a redacted version.
- The client auth check in `responses_api` only validates the presence and shape of a Bearer token; if stronger authentication/authorization is expected, consider integrating actual token validation or clarifying that this endpoint is only gating on header presence.
## Individual Comments### Comment 1
<locationpath="router/main.py"line_range="797-806" />
<code_context>
+}
++
+def _validate_litellm_master_key() -> str:
+ """Validate LITELLM_MASTER_KEY environment variable.
++ Returns:
+ The valid master key string.++ Raises:
+ HTTPException(500): If master key is missing, empty, or placeholder string.+ """
+ key = (os.getenv("LITELLM_MASTER_KEY") or "").strip()
+ if not key or key in _INVALID_MASTER_KEYS or "PLACEHOLDER" in key.upper():
+ logger.error(f"Invalid or missing LITELLM_MASTER_KEY: '{key}'")
</code_context>
<issue_to_address>
**🚨 issue (security):** Avoid logging the raw master key value to prevent leaking secrets.
The `logger.error` line logs the full `LITELLM_MASTER_KEY`, which risks exposing a production secret in logs. Please avoid logging the raw key—log only that it is invalid/missing, or mask it (e.g., show a short prefix or just its length) while preserving enough context for debugging.
</issue_to_address>
### Comment 2
<locationpath="router/main.py"line_range="2306-2310" />
<code_context>
when an auto model (e.g. llm-routing-auto-free) is requested, while supporting model aliases
(such as gpt-4o-mini, local-qwen-3.6-hass) and tool/streaming executions.
"""
+# Enforce client authentication+ auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
+ if not auth_header or not auth_header.startswith("Bearer "):
+ raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")+ client_token = auth_header[7:].strip()
+ if not client_token:
+ raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
</code_context>
<issue_to_address>
**suggestion (bug_risk):** Authorization handling is case-sensitive and extracts a token that is never used.
`auth_header.startswith("Bearer ")` will reject headers with different casing or spacing (e.g. `bearer <token>`). Consider normalizing (e.g. lowercasing and/or splitting on whitespace) before checking to make the auth handling more robust. Also, `client_token` is extracted but never used; either pass it to downstream logic where needed or remove it to avoid confusion.
Suggested implementation:
```pythontry:
await _register_ollama_models_in_db(litellm_master_key)
exceptExceptionas e:
when an auto model (e.g. llm-routing-auto-free) is requested, while supporting model aliases
(such as gpt-4o-mini, local-qwen-3.6-hass) and tool/streaming executions.
""" # Enforce client authentication auth_header = request.headers.get("Authorization") or request.headers.get("authorization") if not auth_header: raise HTTPException(status_code=401, detail="Missing or invalid Authorization header") parts = auth_header.split() if len(parts) != 2 or parts[0].lower() != "bearer" or not parts[1]: raise HTTPException(status_code=401, detail="Missing or invalid Authorization header") # Store the normalized client token for downstream use request.state.client_token = parts[1] try:```1. Anywhere downstream that needs the client token should read it from `request.state.client_token` instead of re-parsing headers.2. If your router uses dependency injection (e.g. FastAPI dependencies) for auth, consider integrating this parsing into a shared dependency to avoid duplication.</issue_to_address>

Sourcery is free for open source - if you like our reviews please consider sharing them ✨
Help me be more useful! Please click 👍 or 👎 on each comment and I'll use the feedback to improve your reviews.

Comment threadrouter/main.py
Comment on lines +797 to +806
def _validate_litellm_master_key() -> str:
"""Validate LITELLM_MASTER_KEY environment variable.

Returns:
The valid master key string.

Raises:
HTTPException(500): If master key is missing, empty, or placeholder string.
"""
key = (os.getenv("LITELLM_MASTER_KEY") or "").strip()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🚨 issue (security): Avoid logging the raw master key value to prevent leaking secrets.

The logger.error line logs the full LITELLM_MASTER_KEY, which risks exposing a production secret in logs. Please avoid logging the raw key—log only that it is invalid/missing, or mask it (e.g., show a short prefix or just its length) while preserving enough context for debugging.

Comment threadrouter/main.py
Comment on lines +2306 to +2310
# Enforce client authentication
auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
if not auth_header or not auth_header.startswith("Bearer "):
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
client_token = auth_header[7:].strip()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

suggestion (bug_risk): Authorization handling is case-sensitive and extracts a token that is never used.

auth_header.startswith("Bearer ") will reject headers with different casing or spacing (e.g. bearer <token>). Consider normalizing (e.g. lowercasing and/or splitting on whitespace) before checking to make the auth handling more robust. Also, client_token is extracted but never used; either pass it to downstream logic where needed or remove it to avoid confusion.

Suggested implementation:

try:
await_register_ollama_models_in_db(litellm_master_key)
exceptExceptionase:
whenanautomodel (e.g. llm-routing-auto-free) isrequested, whilesupportingmodelaliases
(suchasgpt-4o-mini, local-qwen-3.6-hass) andtool/streamingexecutions.
"""
# Enforce client authentication
auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
if not auth_header:
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
parts = auth_header.split()
if len(parts) != 2 or parts[0].lower() != "bearer" or not parts[1]:
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
# Store the normalized client token for downstream userequest.state.client_token=parts[1]
try:
  1. Anywhere downstream that needs the client token should read it from request.state.client_token instead of re-parsing headers.
  2. If your router uses dependency injection (e.g. FastAPI dependencies) for auth, consider integrating this parsing into a shared dependency to avoid duplication.

@sheepdestroyer
sheepdestroyer merged commit 105e7b2 into masterAug 13, 2026
6 of 7 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(dev): restore fallback capacity for canonical E2E chat verification

1 participant

@sheepdestroyer
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(dev): restore fallback capacity for canonical E2E chat verification (#345) by sheepdestroyer · Pull Request #468 · sheepdestroyer/LLM-Routing · GitHub
Skip to content

fix(dev): restore fallback capacity for canonical E2E chat verification (#345) - #468

Merged
sheepdestroyer merged 3 commits into
masterfrom
fix/dev-fallback-capacity-verification
Aug 13, 2026
Merged

fix(dev): restore fallback capacity for canonical E2E chat verification (#345)#468
sheepdestroyer merged 3 commits into
masterfrom
fix/dev-fallback-capacity-verification

Conversation

@sheepdestroyer

@sheepdestroyersheepdestroyer commented Aug 12, 2026

Copy link
Copy Markdown
Owner

Closes#345

Summary by Sourcery

Harden routing and verification infrastructure around LiteLLM/OpenRouter and Langfuse, improving auth, key handling, and canonical endpoint behavior.

Bug Fixes:

  • Restore localhost HTTP fallbacks and explicit LLAMA_* environment variable handling when resolving llama server/classifier endpoints.
  • Ensure LiteLLM master key placeholders and misconfigurations fail fast instead of silently proxying with invalid credentials.
  • Fix adaptive router roster purge assertion to accept multiple purge calls.

Enhancements:

  • Enforce bearer client authentication for the /v1/responses endpoint.
  • Introduce centralized validation of the LiteLLM master key and reuse it for proxy execution and internal calls.
  • Register OpenRouter models as LiteLLM DB models on startup, including dynamic loading from config.yaml and static fallback definitions.
  • Improve Langfuse session propagation tracing to support v4 events_only mode and longer polling, and propagate session/user IDs into Langfuse observations and updates.
  • Extend start-stack.sh to export LLAMA_* env vars and render LiteLLM config using derived classifier URLs.

Tests:

  • Update existing routing, lifespan, and responses_api tests to respect new auth and master key validation behavior.
  • Add tests covering OpenRouter model DB registration paths (no key, static fallback, and config-driven) and LiteLLM master key validation failure modes.
  • Add verification coverage for Langfuse events_only mode and non-leak behavior in session propagation checks.

@sourcery-ai

sourcery-aiBot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Reviewer's Guide

Restores and hardens environment-based configuration and verification for LiteLLM/OpenRouter routing, adds strict master-key and client auth validation, improves Langfuse tracing/session propagation, and extends tests/startup scripts to cover the new behavior and canonical endpoint derivation paths.

Sequence diagram for responses_api client auth and master key validation

sequenceDiagram
actor Client
participant Router as responses_api
participant Env as _validate_litellm_master_key
participant LiteLLM as litellm_proxy
Client->>Router: POST /responses (Authorization: Bearer client_token)
Router->>Router: validate Authorization header
Router->>Env: _validate_litellm_master_key()
Env-->>Router: master_key
Router->>LiteLLM: client.post /responses (Authorization: Bearer master_key)
LiteLLM-->>Router: response
Router-->>Client: routed response
Loading

File-Level Changes

ChangeDetailsFiles
Make LLAMA server/classifier endpoint resolution respect explicit environment variables while preserving localhost fallbacks and canonical HTTPS derivation.
  • Introduce LLAMA_SERVER_URL and LLAMA_CLASSIFIER_URL environment overrides when resolving endpoints.
  • Fallback to config values or localhost defaults when env vars are absent.
  • Prefer explicit env URLs over canonical HTTPS derivation in resolution logic.
router/main.py
start-stack.sh
Introduce centralized LiteLLM master key validation and stricter usage across responses and proxy paths.
  • Add _INVALID_MASTER_KEYS set and _validate_litellm_master_key helper that fails fast with HTTP 500 on missing/placeholder keys.
  • Use _validate_litellm_master_key in /v1/responses and execute_proxy instead of raw os.getenv lookups.
  • Update tests to patch LITELLM_MASTER_KEY and cover invalid-key behavior.
router/main.py
router/tests/test_responses_api.py
router/tests/test_routing_behavior.py
tests/conftest.py
Add OpenRouter model DB registration on startup with dynamic config loading, static fallback, and stale deployment purge.
  • Implement _register_openrouter_models_in_db that loads OpenRouter models from litellm config paths or falls back to a static openrouter-auto definition.
  • Purge stale openrouter-* deployments via _purge_stale_deployments before re-registering.
  • Wire OpenRouter registration into FastAPI lifespan and add unit tests for config-based and fallback registration paths.
router/main.py
router/tests/test_lifespan.py
router/tests/test_register_openrouter_models_in_db.py
router/tests/test_sync_adaptive_router_roster.py
Enforce client Bearer authentication and improve Langfuse tracing/session propagation behavior in responses and chat triage.
  • Require a non-empty Bearer Authorization header in responses_api, returning 401 on missing/invalid auth.
  • Pass session_id and user_id into Langfuse start_observation and update calls when available.
  • Extend Langfuse verification script to handle v4 events_only mode, longer polling, and revised leak checks.
  • Update tests to include Authorization headers and expectations for auth enforcement and tracing behavior.
router/main.py
scripts/verification/verify_canonical_endpoints.py
router/tests/test_responses_api.py
Tighten test and startup environment configuration around routing behavior and master key usage.
  • Set default LITELLM_MASTER_KEY and ROUTER_API_KEY in test conftest and specific routing tests.
  • Adjust routing behavior tests to provide Authorization headers and patched master keys.
  • Update start-stack.sh to export LLAMA_SERVER_URL/LLAMA_CLASSIFIER_URL and derive classifier URL in rendered LiteLLM config.
  • Relax sync_adaptive_router_roster purge assertion to allow multiple purge calls.
router/tests/test_routing_behavior.py
tests/conftest.py
start-stack.sh
router/tests/test_sync_adaptive_router_roster.py

Assessment against linked issues

IssueObjectiveAddressedExplanation
#345Restore functional OpenRouter-backed free-tier chat routing for dev canonical verification, including robust LiteLLM master key handling and model registration so auto/free routes (e.g. llm-routing-auto-free, agent-simple-core) can succeed.
#345Fix dev LiteLLM local safety-net routing by correctly deriving and wiring LLAMA_SERVER_URL / LLAMA_CLASSIFIER_URL from configuration/start-stack into router, so the local fallback classifier becomes reachable from the dev container without coupling to a broken production hostname.

Possibly linked issues


Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@coderabbitai

coderabbitaiBot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@sheepdestroyer, you've reached your PR review limit, so we couldn't start this review.

Next review available in:117 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 3f27a272-9764-4dc3-8e95-2f0a8a765e1e

📥 Commits

Reviewing files that changed from the base of the PR and between 830008b and 4e11105.

📒 Files selected for processing (10)
  • .env.dev
  • router/main.py
  • router/tests/test_lifespan.py
  • router/tests/test_register_openrouter_models_in_db.py
  • router/tests/test_responses_api.py
  • router/tests/test_routing_behavior.py
  • router/tests/test_sync_adaptive_router_roster.py
  • scripts/verification/verify_canonical_endpoints.py
  • start-stack.sh
  • tests/conftest.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sourcery-aisourcery-aiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 2 issues, and left some high level feedback:

  • Avoid logging the raw LITELLM_MASTER_KEY value in _validate_litellm_master_key, as the current logger.error call can expose secrets in logs; consider logging only that it is missing/invalid or a redacted version.
  • The client auth check in responses_api only validates the presence and shape of a Bearer token; if stronger authentication/authorization is expected, consider integrating actual token validation or clarifying that this endpoint is only gating on header presence.
Prompt for AI Agents
Please address the comments from this code review:
## Overall Comments- Avoid logging the raw LITELLM_MASTER_KEY value in `_validate_litellm_master_key`, as the current `logger.error` call can expose secrets in logs; consider logging only that it is missing/invalid or a redacted version.
- The client auth check in `responses_api` only validates the presence and shape of a Bearer token; if stronger authentication/authorization is expected, consider integrating actual token validation or clarifying that this endpoint is only gating on header presence.
## Individual Comments### Comment 1
<locationpath="router/main.py"line_range="797-806" />
<code_context>
+}
++
+def _validate_litellm_master_key() -> str:
+ """Validate LITELLM_MASTER_KEY environment variable.
++ Returns:
+ The valid master key string.++ Raises:
+ HTTPException(500): If master key is missing, empty, or placeholder string.+ """
+ key = (os.getenv("LITELLM_MASTER_KEY") or "").strip()
+ if not key or key in _INVALID_MASTER_KEYS or "PLACEHOLDER" in key.upper():
+ logger.error(f"Invalid or missing LITELLM_MASTER_KEY: '{key}'")
</code_context>
<issue_to_address>
**🚨 issue (security):** Avoid logging the raw master key value to prevent leaking secrets.
The `logger.error` line logs the full `LITELLM_MASTER_KEY`, which risks exposing a production secret in logs. Please avoid logging the raw key—log only that it is invalid/missing, or mask it (e.g., show a short prefix or just its length) while preserving enough context for debugging.
</issue_to_address>
### Comment 2
<locationpath="router/main.py"line_range="2306-2310" />
<code_context>
when an auto model (e.g. llm-routing-auto-free) is requested, while supporting model aliases
(such as gpt-4o-mini, local-qwen-3.6-hass) and tool/streaming executions.
"""
+# Enforce client authentication+ auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
+ if not auth_header or not auth_header.startswith("Bearer "):
+ raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")+ client_token = auth_header[7:].strip()
+ if not client_token:
+ raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
</code_context>
<issue_to_address>
**suggestion (bug_risk):** Authorization handling is case-sensitive and extracts a token that is never used.
`auth_header.startswith("Bearer ")` will reject headers with different casing or spacing (e.g. `bearer <token>`). Consider normalizing (e.g. lowercasing and/or splitting on whitespace) before checking to make the auth handling more robust. Also, `client_token` is extracted but never used; either pass it to downstream logic where needed or remove it to avoid confusion.
Suggested implementation:
```pythontry:
await _register_ollama_models_in_db(litellm_master_key)
exceptExceptionas e:
when an auto model (e.g. llm-routing-auto-free) is requested, while supporting model aliases
(such as gpt-4o-mini, local-qwen-3.6-hass) and tool/streaming executions.
""" # Enforce client authentication auth_header = request.headers.get("Authorization") or request.headers.get("authorization") if not auth_header: raise HTTPException(status_code=401, detail="Missing or invalid Authorization header") parts = auth_header.split() if len(parts) != 2 or parts[0].lower() != "bearer" or not parts[1]: raise HTTPException(status_code=401, detail="Missing or invalid Authorization header") # Store the normalized client token for downstream use request.state.client_token = parts[1] try:```1. Anywhere downstream that needs the client token should read it from `request.state.client_token` instead of re-parsing headers.2. If your router uses dependency injection (e.g. FastAPI dependencies) for auth, consider integrating this parsing into a shared dependency to avoid duplication.</issue_to_address>

Sourcery is free for open source - if you like our reviews please consider sharing them ✨
Help me be more useful! Please click 👍 or 👎 on each comment and I'll use the feedback to improve your reviews.

Comment threadrouter/main.py
Comment on lines +797 to +806
def _validate_litellm_master_key() -> str:
"""Validate LITELLM_MASTER_KEY environment variable.

Returns:
The valid master key string.

Raises:
HTTPException(500): If master key is missing, empty, or placeholder string.
"""
key = (os.getenv("LITELLM_MASTER_KEY") or "").strip()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🚨 issue (security): Avoid logging the raw master key value to prevent leaking secrets.

The logger.error line logs the full LITELLM_MASTER_KEY, which risks exposing a production secret in logs. Please avoid logging the raw key—log only that it is invalid/missing, or mask it (e.g., show a short prefix or just its length) while preserving enough context for debugging.

Comment threadrouter/main.py
Comment on lines +2306 to +2310
# Enforce client authentication
auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
if not auth_header or not auth_header.startswith("Bearer "):
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
client_token = auth_header[7:].strip()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

suggestion (bug_risk): Authorization handling is case-sensitive and extracts a token that is never used.

auth_header.startswith("Bearer ") will reject headers with different casing or spacing (e.g. bearer <token>). Consider normalizing (e.g. lowercasing and/or splitting on whitespace) before checking to make the auth handling more robust. Also, client_token is extracted but never used; either pass it to downstream logic where needed or remove it to avoid confusion.

Suggested implementation:

try:
await_register_ollama_models_in_db(litellm_master_key)
exceptExceptionase:
whenanautomodel (e.g. llm-routing-auto-free) isrequested, whilesupportingmodelaliases
(suchasgpt-4o-mini, local-qwen-3.6-hass) andtool/streamingexecutions.
"""
# Enforce client authentication
auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
if not auth_header:
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
parts = auth_header.split()
if len(parts) != 2 or parts[0].lower() != "bearer" or not parts[1]:
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
# Store the normalized client token for downstream userequest.state.client_token=parts[1]
try:
  1. Anywhere downstream that needs the client token should read it from request.state.client_token instead of re-parsing headers.
  2. If your router uses dependency injection (e.g. FastAPI dependencies) for auth, consider integrating this parsing into a shared dependency to avoid duplication.

@sheepdestroyer
sheepdestroyer merged commit 105e7b2 into masterAug 13, 2026
6 of 7 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(dev): restore fallback capacity for canonical E2E chat verification

1 participant

@sheepdestroyer
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(dev): restore fallback capacity for canonical E2E chat verification (#345) by sheepdestroyer · Pull Request #468 · sheepdestroyer/LLM-Routing · GitHub
Skip to content

fix(dev): restore fallback capacity for canonical E2E chat verification (#345) - #468

Merged
sheepdestroyer merged 3 commits into
masterfrom
fix/dev-fallback-capacity-verification
Aug 13, 2026
Merged

fix(dev): restore fallback capacity for canonical E2E chat verification (#345)#468
sheepdestroyer merged 3 commits into
masterfrom
fix/dev-fallback-capacity-verification

Conversation

@sheepdestroyer

@sheepdestroyersheepdestroyer commented Aug 12, 2026

Copy link
Copy Markdown
Owner

Closes#345

Summary by Sourcery

Harden routing and verification infrastructure around LiteLLM/OpenRouter and Langfuse, improving auth, key handling, and canonical endpoint behavior.

Bug Fixes:

  • Restore localhost HTTP fallbacks and explicit LLAMA_* environment variable handling when resolving llama server/classifier endpoints.
  • Ensure LiteLLM master key placeholders and misconfigurations fail fast instead of silently proxying with invalid credentials.
  • Fix adaptive router roster purge assertion to accept multiple purge calls.

Enhancements:

  • Enforce bearer client authentication for the /v1/responses endpoint.
  • Introduce centralized validation of the LiteLLM master key and reuse it for proxy execution and internal calls.
  • Register OpenRouter models as LiteLLM DB models on startup, including dynamic loading from config.yaml and static fallback definitions.
  • Improve Langfuse session propagation tracing to support v4 events_only mode and longer polling, and propagate session/user IDs into Langfuse observations and updates.
  • Extend start-stack.sh to export LLAMA_* env vars and render LiteLLM config using derived classifier URLs.

Tests:

  • Update existing routing, lifespan, and responses_api tests to respect new auth and master key validation behavior.
  • Add tests covering OpenRouter model DB registration paths (no key, static fallback, and config-driven) and LiteLLM master key validation failure modes.
  • Add verification coverage for Langfuse events_only mode and non-leak behavior in session propagation checks.

@sourcery-ai

sourcery-aiBot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Reviewer's Guide

Restores and hardens environment-based configuration and verification for LiteLLM/OpenRouter routing, adds strict master-key and client auth validation, improves Langfuse tracing/session propagation, and extends tests/startup scripts to cover the new behavior and canonical endpoint derivation paths.

Sequence diagram for responses_api client auth and master key validation

sequenceDiagram
actor Client
participant Router as responses_api
participant Env as _validate_litellm_master_key
participant LiteLLM as litellm_proxy
Client->>Router: POST /responses (Authorization: Bearer client_token)
Router->>Router: validate Authorization header
Router->>Env: _validate_litellm_master_key()
Env-->>Router: master_key
Router->>LiteLLM: client.post /responses (Authorization: Bearer master_key)
LiteLLM-->>Router: response
Router-->>Client: routed response
Loading

File-Level Changes

ChangeDetailsFiles
Make LLAMA server/classifier endpoint resolution respect explicit environment variables while preserving localhost fallbacks and canonical HTTPS derivation.
  • Introduce LLAMA_SERVER_URL and LLAMA_CLASSIFIER_URL environment overrides when resolving endpoints.
  • Fallback to config values or localhost defaults when env vars are absent.
  • Prefer explicit env URLs over canonical HTTPS derivation in resolution logic.
router/main.py
start-stack.sh
Introduce centralized LiteLLM master key validation and stricter usage across responses and proxy paths.
  • Add _INVALID_MASTER_KEYS set and _validate_litellm_master_key helper that fails fast with HTTP 500 on missing/placeholder keys.
  • Use _validate_litellm_master_key in /v1/responses and execute_proxy instead of raw os.getenv lookups.
  • Update tests to patch LITELLM_MASTER_KEY and cover invalid-key behavior.
router/main.py
router/tests/test_responses_api.py
router/tests/test_routing_behavior.py
tests/conftest.py
Add OpenRouter model DB registration on startup with dynamic config loading, static fallback, and stale deployment purge.
  • Implement _register_openrouter_models_in_db that loads OpenRouter models from litellm config paths or falls back to a static openrouter-auto definition.
  • Purge stale openrouter-* deployments via _purge_stale_deployments before re-registering.
  • Wire OpenRouter registration into FastAPI lifespan and add unit tests for config-based and fallback registration paths.
router/main.py
router/tests/test_lifespan.py
router/tests/test_register_openrouter_models_in_db.py
router/tests/test_sync_adaptive_router_roster.py
Enforce client Bearer authentication and improve Langfuse tracing/session propagation behavior in responses and chat triage.
  • Require a non-empty Bearer Authorization header in responses_api, returning 401 on missing/invalid auth.
  • Pass session_id and user_id into Langfuse start_observation and update calls when available.
  • Extend Langfuse verification script to handle v4 events_only mode, longer polling, and revised leak checks.
  • Update tests to include Authorization headers and expectations for auth enforcement and tracing behavior.
router/main.py
scripts/verification/verify_canonical_endpoints.py
router/tests/test_responses_api.py
Tighten test and startup environment configuration around routing behavior and master key usage.
  • Set default LITELLM_MASTER_KEY and ROUTER_API_KEY in test conftest and specific routing tests.
  • Adjust routing behavior tests to provide Authorization headers and patched master keys.
  • Update start-stack.sh to export LLAMA_SERVER_URL/LLAMA_CLASSIFIER_URL and derive classifier URL in rendered LiteLLM config.
  • Relax sync_adaptive_router_roster purge assertion to allow multiple purge calls.
router/tests/test_routing_behavior.py
tests/conftest.py
start-stack.sh
router/tests/test_sync_adaptive_router_roster.py

Assessment against linked issues

IssueObjectiveAddressedExplanation
#345Restore functional OpenRouter-backed free-tier chat routing for dev canonical verification, including robust LiteLLM master key handling and model registration so auto/free routes (e.g. llm-routing-auto-free, agent-simple-core) can succeed.
#345Fix dev LiteLLM local safety-net routing by correctly deriving and wiring LLAMA_SERVER_URL / LLAMA_CLASSIFIER_URL from configuration/start-stack into router, so the local fallback classifier becomes reachable from the dev container without coupling to a broken production hostname.

Possibly linked issues


Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@coderabbitai

coderabbitaiBot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@sheepdestroyer, you've reached your PR review limit, so we couldn't start this review.

Next review available in:117 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 3f27a272-9764-4dc3-8e95-2f0a8a765e1e

📥 Commits

Reviewing files that changed from the base of the PR and between 830008b and 4e11105.

📒 Files selected for processing (10)
  • .env.dev
  • router/main.py
  • router/tests/test_lifespan.py
  • router/tests/test_register_openrouter_models_in_db.py
  • router/tests/test_responses_api.py
  • router/tests/test_routing_behavior.py
  • router/tests/test_sync_adaptive_router_roster.py
  • scripts/verification/verify_canonical_endpoints.py
  • start-stack.sh
  • tests/conftest.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sourcery-aisourcery-aiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 2 issues, and left some high level feedback:

  • Avoid logging the raw LITELLM_MASTER_KEY value in _validate_litellm_master_key, as the current logger.error call can expose secrets in logs; consider logging only that it is missing/invalid or a redacted version.
  • The client auth check in responses_api only validates the presence and shape of a Bearer token; if stronger authentication/authorization is expected, consider integrating actual token validation or clarifying that this endpoint is only gating on header presence.
Prompt for AI Agents
Please address the comments from this code review:
## Overall Comments- Avoid logging the raw LITELLM_MASTER_KEY value in `_validate_litellm_master_key`, as the current `logger.error` call can expose secrets in logs; consider logging only that it is missing/invalid or a redacted version.
- The client auth check in `responses_api` only validates the presence and shape of a Bearer token; if stronger authentication/authorization is expected, consider integrating actual token validation or clarifying that this endpoint is only gating on header presence.
## Individual Comments### Comment 1
<locationpath="router/main.py"line_range="797-806" />
<code_context>
+}
++
+def _validate_litellm_master_key() -> str:
+ """Validate LITELLM_MASTER_KEY environment variable.
++ Returns:
+ The valid master key string.++ Raises:
+ HTTPException(500): If master key is missing, empty, or placeholder string.+ """
+ key = (os.getenv("LITELLM_MASTER_KEY") or "").strip()
+ if not key or key in _INVALID_MASTER_KEYS or "PLACEHOLDER" in key.upper():
+ logger.error(f"Invalid or missing LITELLM_MASTER_KEY: '{key}'")
</code_context>
<issue_to_address>
**🚨 issue (security):** Avoid logging the raw master key value to prevent leaking secrets.
The `logger.error` line logs the full `LITELLM_MASTER_KEY`, which risks exposing a production secret in logs. Please avoid logging the raw key—log only that it is invalid/missing, or mask it (e.g., show a short prefix or just its length) while preserving enough context for debugging.
</issue_to_address>
### Comment 2
<locationpath="router/main.py"line_range="2306-2310" />
<code_context>
when an auto model (e.g. llm-routing-auto-free) is requested, while supporting model aliases
(such as gpt-4o-mini, local-qwen-3.6-hass) and tool/streaming executions.
"""
+# Enforce client authentication+ auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
+ if not auth_header or not auth_header.startswith("Bearer "):
+ raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")+ client_token = auth_header[7:].strip()
+ if not client_token:
+ raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
</code_context>
<issue_to_address>
**suggestion (bug_risk):** Authorization handling is case-sensitive and extracts a token that is never used.
`auth_header.startswith("Bearer ")` will reject headers with different casing or spacing (e.g. `bearer <token>`). Consider normalizing (e.g. lowercasing and/or splitting on whitespace) before checking to make the auth handling more robust. Also, `client_token` is extracted but never used; either pass it to downstream logic where needed or remove it to avoid confusion.
Suggested implementation:
```pythontry:
await _register_ollama_models_in_db(litellm_master_key)
exceptExceptionas e:
when an auto model (e.g. llm-routing-auto-free) is requested, while supporting model aliases
(such as gpt-4o-mini, local-qwen-3.6-hass) and tool/streaming executions.
""" # Enforce client authentication auth_header = request.headers.get("Authorization") or request.headers.get("authorization") if not auth_header: raise HTTPException(status_code=401, detail="Missing or invalid Authorization header") parts = auth_header.split() if len(parts) != 2 or parts[0].lower() != "bearer" or not parts[1]: raise HTTPException(status_code=401, detail="Missing or invalid Authorization header") # Store the normalized client token for downstream use request.state.client_token = parts[1] try:```1. Anywhere downstream that needs the client token should read it from `request.state.client_token` instead of re-parsing headers.2. If your router uses dependency injection (e.g. FastAPI dependencies) for auth, consider integrating this parsing into a shared dependency to avoid duplication.</issue_to_address>

Sourcery is free for open source - if you like our reviews please consider sharing them ✨
Help me be more useful! Please click 👍 or 👎 on each comment and I'll use the feedback to improve your reviews.

Comment threadrouter/main.py
Comment on lines +797 to +806
def _validate_litellm_master_key() -> str:
"""Validate LITELLM_MASTER_KEY environment variable.

Returns:
The valid master key string.

Raises:
HTTPException(500): If master key is missing, empty, or placeholder string.
"""
key = (os.getenv("LITELLM_MASTER_KEY") or "").strip()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🚨 issue (security): Avoid logging the raw master key value to prevent leaking secrets.

The logger.error line logs the full LITELLM_MASTER_KEY, which risks exposing a production secret in logs. Please avoid logging the raw key—log only that it is invalid/missing, or mask it (e.g., show a short prefix or just its length) while preserving enough context for debugging.

Comment threadrouter/main.py
Comment on lines +2306 to +2310
# Enforce client authentication
auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
if not auth_header or not auth_header.startswith("Bearer "):
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
client_token = auth_header[7:].strip()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

suggestion (bug_risk): Authorization handling is case-sensitive and extracts a token that is never used.

auth_header.startswith("Bearer ") will reject headers with different casing or spacing (e.g. bearer <token>). Consider normalizing (e.g. lowercasing and/or splitting on whitespace) before checking to make the auth handling more robust. Also, client_token is extracted but never used; either pass it to downstream logic where needed or remove it to avoid confusion.

Suggested implementation:

try:
await_register_ollama_models_in_db(litellm_master_key)
exceptExceptionase:
whenanautomodel (e.g. llm-routing-auto-free) isrequested, whilesupportingmodelaliases
(suchasgpt-4o-mini, local-qwen-3.6-hass) andtool/streamingexecutions.
"""
# Enforce client authentication
auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
if not auth_header:
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
parts = auth_header.split()
if len(parts) != 2 or parts[0].lower() != "bearer" or not parts[1]:
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
# Store the normalized client token for downstream userequest.state.client_token=parts[1]
try:
  1. Anywhere downstream that needs the client token should read it from request.state.client_token instead of re-parsing headers.
  2. If your router uses dependency injection (e.g. FastAPI dependencies) for auth, consider integrating this parsing into a shared dependency to avoid duplication.

@sheepdestroyer
sheepdestroyer merged commit 105e7b2 into masterAug 13, 2026
6 of 7 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(dev): restore fallback capacity for canonical E2E chat verification

1 participant

@sheepdestroyer
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' fix(dev): restore fallback capacity for canonical E2E chat verification (#345) by sheepdestroyer · Pull Request #468 · sheepdestroyer/LLM-Routing · GitHub
Skip to content

fix(dev): restore fallback capacity for canonical E2E chat verification (#345) - #468

Merged
sheepdestroyer merged 3 commits into
masterfrom
fix/dev-fallback-capacity-verification
Aug 13, 2026
Merged

fix(dev): restore fallback capacity for canonical E2E chat verification (#345)#468
sheepdestroyer merged 3 commits into
masterfrom
fix/dev-fallback-capacity-verification

Conversation

@sheepdestroyer

@sheepdestroyersheepdestroyer commented Aug 12, 2026

Copy link
Copy Markdown
Owner

Closes#345

Summary by Sourcery

Harden routing and verification infrastructure around LiteLLM/OpenRouter and Langfuse, improving auth, key handling, and canonical endpoint behavior.

Bug Fixes:

  • Restore localhost HTTP fallbacks and explicit LLAMA_* environment variable handling when resolving llama server/classifier endpoints.
  • Ensure LiteLLM master key placeholders and misconfigurations fail fast instead of silently proxying with invalid credentials.
  • Fix adaptive router roster purge assertion to accept multiple purge calls.

Enhancements:

  • Enforce bearer client authentication for the /v1/responses endpoint.
  • Introduce centralized validation of the LiteLLM master key and reuse it for proxy execution and internal calls.
  • Register OpenRouter models as LiteLLM DB models on startup, including dynamic loading from config.yaml and static fallback definitions.
  • Improve Langfuse session propagation tracing to support v4 events_only mode and longer polling, and propagate session/user IDs into Langfuse observations and updates.
  • Extend start-stack.sh to export LLAMA_* env vars and render LiteLLM config using derived classifier URLs.

Tests:

  • Update existing routing, lifespan, and responses_api tests to respect new auth and master key validation behavior.
  • Add tests covering OpenRouter model DB registration paths (no key, static fallback, and config-driven) and LiteLLM master key validation failure modes.
  • Add verification coverage for Langfuse events_only mode and non-leak behavior in session propagation checks.

@sourcery-ai

sourcery-aiBot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Reviewer's Guide

Restores and hardens environment-based configuration and verification for LiteLLM/OpenRouter routing, adds strict master-key and client auth validation, improves Langfuse tracing/session propagation, and extends tests/startup scripts to cover the new behavior and canonical endpoint derivation paths.

Sequence diagram for responses_api client auth and master key validation

sequenceDiagram
actor Client
participant Router as responses_api
participant Env as _validate_litellm_master_key
participant LiteLLM as litellm_proxy
Client->>Router: POST /responses (Authorization: Bearer client_token)
Router->>Router: validate Authorization header
Router->>Env: _validate_litellm_master_key()
Env-->>Router: master_key
Router->>LiteLLM: client.post /responses (Authorization: Bearer master_key)
LiteLLM-->>Router: response
Router-->>Client: routed response
Loading

File-Level Changes

ChangeDetailsFiles
Make LLAMA server/classifier endpoint resolution respect explicit environment variables while preserving localhost fallbacks and canonical HTTPS derivation.
  • Introduce LLAMA_SERVER_URL and LLAMA_CLASSIFIER_URL environment overrides when resolving endpoints.
  • Fallback to config values or localhost defaults when env vars are absent.
  • Prefer explicit env URLs over canonical HTTPS derivation in resolution logic.
router/main.py
start-stack.sh
Introduce centralized LiteLLM master key validation and stricter usage across responses and proxy paths.
  • Add _INVALID_MASTER_KEYS set and _validate_litellm_master_key helper that fails fast with HTTP 500 on missing/placeholder keys.
  • Use _validate_litellm_master_key in /v1/responses and execute_proxy instead of raw os.getenv lookups.
  • Update tests to patch LITELLM_MASTER_KEY and cover invalid-key behavior.
router/main.py
router/tests/test_responses_api.py
router/tests/test_routing_behavior.py
tests/conftest.py
Add OpenRouter model DB registration on startup with dynamic config loading, static fallback, and stale deployment purge.
  • Implement _register_openrouter_models_in_db that loads OpenRouter models from litellm config paths or falls back to a static openrouter-auto definition.
  • Purge stale openrouter-* deployments via _purge_stale_deployments before re-registering.
  • Wire OpenRouter registration into FastAPI lifespan and add unit tests for config-based and fallback registration paths.
router/main.py
router/tests/test_lifespan.py
router/tests/test_register_openrouter_models_in_db.py
router/tests/test_sync_adaptive_router_roster.py
Enforce client Bearer authentication and improve Langfuse tracing/session propagation behavior in responses and chat triage.
  • Require a non-empty Bearer Authorization header in responses_api, returning 401 on missing/invalid auth.
  • Pass session_id and user_id into Langfuse start_observation and update calls when available.
  • Extend Langfuse verification script to handle v4 events_only mode, longer polling, and revised leak checks.
  • Update tests to include Authorization headers and expectations for auth enforcement and tracing behavior.
router/main.py
scripts/verification/verify_canonical_endpoints.py
router/tests/test_responses_api.py
Tighten test and startup environment configuration around routing behavior and master key usage.
  • Set default LITELLM_MASTER_KEY and ROUTER_API_KEY in test conftest and specific routing tests.
  • Adjust routing behavior tests to provide Authorization headers and patched master keys.
  • Update start-stack.sh to export LLAMA_SERVER_URL/LLAMA_CLASSIFIER_URL and derive classifier URL in rendered LiteLLM config.
  • Relax sync_adaptive_router_roster purge assertion to allow multiple purge calls.
router/tests/test_routing_behavior.py
tests/conftest.py
start-stack.sh
router/tests/test_sync_adaptive_router_roster.py

Assessment against linked issues

IssueObjectiveAddressedExplanation
#345Restore functional OpenRouter-backed free-tier chat routing for dev canonical verification, including robust LiteLLM master key handling and model registration so auto/free routes (e.g. llm-routing-auto-free, agent-simple-core) can succeed.
#345Fix dev LiteLLM local safety-net routing by correctly deriving and wiring LLAMA_SERVER_URL / LLAMA_CLASSIFIER_URL from configuration/start-stack into router, so the local fallback classifier becomes reachable from the dev container without coupling to a broken production hostname.

Possibly linked issues


Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@coderabbitai

coderabbitaiBot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@sheepdestroyer, you've reached your PR review limit, so we couldn't start this review.

Next review available in:117 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 3f27a272-9764-4dc3-8e95-2f0a8a765e1e

📥 Commits

Reviewing files that changed from the base of the PR and between 830008b and 4e11105.

📒 Files selected for processing (10)
  • .env.dev
  • router/main.py
  • router/tests/test_lifespan.py
  • router/tests/test_register_openrouter_models_in_db.py
  • router/tests/test_responses_api.py
  • router/tests/test_routing_behavior.py
  • router/tests/test_sync_adaptive_router_roster.py
  • scripts/verification/verify_canonical_endpoints.py
  • start-stack.sh
  • tests/conftest.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sourcery-aisourcery-aiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 2 issues, and left some high level feedback:

  • Avoid logging the raw LITELLM_MASTER_KEY value in _validate_litellm_master_key, as the current logger.error call can expose secrets in logs; consider logging only that it is missing/invalid or a redacted version.
  • The client auth check in responses_api only validates the presence and shape of a Bearer token; if stronger authentication/authorization is expected, consider integrating actual token validation or clarifying that this endpoint is only gating on header presence.
Prompt for AI Agents
Please address the comments from this code review:
## Overall Comments- Avoid logging the raw LITELLM_MASTER_KEY value in `_validate_litellm_master_key`, as the current `logger.error` call can expose secrets in logs; consider logging only that it is missing/invalid or a redacted version.
- The client auth check in `responses_api` only validates the presence and shape of a Bearer token; if stronger authentication/authorization is expected, consider integrating actual token validation or clarifying that this endpoint is only gating on header presence.
## Individual Comments### Comment 1
<locationpath="router/main.py"line_range="797-806" />
<code_context>
+}
++
+def _validate_litellm_master_key() -> str:
+ """Validate LITELLM_MASTER_KEY environment variable.
++ Returns:
+ The valid master key string.++ Raises:
+ HTTPException(500): If master key is missing, empty, or placeholder string.+ """
+ key = (os.getenv("LITELLM_MASTER_KEY") or "").strip()
+ if not key or key in _INVALID_MASTER_KEYS or "PLACEHOLDER" in key.upper():
+ logger.error(f"Invalid or missing LITELLM_MASTER_KEY: '{key}'")
</code_context>
<issue_to_address>
**🚨 issue (security):** Avoid logging the raw master key value to prevent leaking secrets.
The `logger.error` line logs the full `LITELLM_MASTER_KEY`, which risks exposing a production secret in logs. Please avoid logging the raw key—log only that it is invalid/missing, or mask it (e.g., show a short prefix or just its length) while preserving enough context for debugging.
</issue_to_address>
### Comment 2
<locationpath="router/main.py"line_range="2306-2310" />
<code_context>
when an auto model (e.g. llm-routing-auto-free) is requested, while supporting model aliases
(such as gpt-4o-mini, local-qwen-3.6-hass) and tool/streaming executions.
"""
+# Enforce client authentication+ auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
+ if not auth_header or not auth_header.startswith("Bearer "):
+ raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")+ client_token = auth_header[7:].strip()
+ if not client_token:
+ raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
</code_context>
<issue_to_address>
**suggestion (bug_risk):** Authorization handling is case-sensitive and extracts a token that is never used.
`auth_header.startswith("Bearer ")` will reject headers with different casing or spacing (e.g. `bearer <token>`). Consider normalizing (e.g. lowercasing and/or splitting on whitespace) before checking to make the auth handling more robust. Also, `client_token` is extracted but never used; either pass it to downstream logic where needed or remove it to avoid confusion.
Suggested implementation:
```pythontry:
await _register_ollama_models_in_db(litellm_master_key)
exceptExceptionas e:
when an auto model (e.g. llm-routing-auto-free) is requested, while supporting model aliases
(such as gpt-4o-mini, local-qwen-3.6-hass) and tool/streaming executions.
""" # Enforce client authentication auth_header = request.headers.get("Authorization") or request.headers.get("authorization") if not auth_header: raise HTTPException(status_code=401, detail="Missing or invalid Authorization header") parts = auth_header.split() if len(parts) != 2 or parts[0].lower() != "bearer" or not parts[1]: raise HTTPException(status_code=401, detail="Missing or invalid Authorization header") # Store the normalized client token for downstream use request.state.client_token = parts[1] try:```1. Anywhere downstream that needs the client token should read it from `request.state.client_token` instead of re-parsing headers.2. If your router uses dependency injection (e.g. FastAPI dependencies) for auth, consider integrating this parsing into a shared dependency to avoid duplication.</issue_to_address>

Sourcery is free for open source - if you like our reviews please consider sharing them ✨
Help me be more useful! Please click 👍 or 👎 on each comment and I'll use the feedback to improve your reviews.

Comment threadrouter/main.py
Comment on lines +797 to +806
def _validate_litellm_master_key() -> str:
"""Validate LITELLM_MASTER_KEY environment variable.

Returns:
The valid master key string.

Raises:
HTTPException(500): If master key is missing, empty, or placeholder string.
"""
key = (os.getenv("LITELLM_MASTER_KEY") or "").strip()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🚨 issue (security): Avoid logging the raw master key value to prevent leaking secrets.

The logger.error line logs the full LITELLM_MASTER_KEY, which risks exposing a production secret in logs. Please avoid logging the raw key—log only that it is invalid/missing, or mask it (e.g., show a short prefix or just its length) while preserving enough context for debugging.

Comment threadrouter/main.py
Comment on lines +2306 to +2310
# Enforce client authentication
auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
if not auth_header or not auth_header.startswith("Bearer "):
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
client_token = auth_header[7:].strip()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

suggestion (bug_risk): Authorization handling is case-sensitive and extracts a token that is never used.

auth_header.startswith("Bearer ") will reject headers with different casing or spacing (e.g. bearer <token>). Consider normalizing (e.g. lowercasing and/or splitting on whitespace) before checking to make the auth handling more robust. Also, client_token is extracted but never used; either pass it to downstream logic where needed or remove it to avoid confusion.

Suggested implementation:

try:
await_register_ollama_models_in_db(litellm_master_key)
exceptExceptionase:
whenanautomodel (e.g. llm-routing-auto-free) isrequested, whilesupportingmodelaliases
(suchasgpt-4o-mini, local-qwen-3.6-hass) andtool/streamingexecutions.
"""
# Enforce client authentication
auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
if not auth_header:
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
parts = auth_header.split()
if len(parts) != 2 or parts[0].lower() != "bearer" or not parts[1]:
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
# Store the normalized client token for downstream userequest.state.client_token=parts[1]
try:
  1. Anywhere downstream that needs the client token should read it from request.state.client_token instead of re-parsing headers.
  2. If your router uses dependency injection (e.g. FastAPI dependencies) for auth, consider integrating this parsing into a shared dependency to avoid duplication.

@sheepdestroyer
sheepdestroyer merged commit 105e7b2 into masterAug 13, 2026
6 of 7 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(dev): restore fallback capacity for canonical E2E chat verification

1 participant

@sheepdestroyer
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(dev): restore fallback capacity for canonical E2E chat verification (#345) by sheepdestroyer · Pull Request #468 · sheepdestroyer/LLM-Routing · GitHub
Skip to content

fix(dev): restore fallback capacity for canonical E2E chat verification (#345) - #468

Merged
sheepdestroyer merged 3 commits into
masterfrom
fix/dev-fallback-capacity-verification
Aug 13, 2026
Merged

fix(dev): restore fallback capacity for canonical E2E chat verification (#345)#468
sheepdestroyer merged 3 commits into
masterfrom
fix/dev-fallback-capacity-verification

Conversation

@sheepdestroyer

@sheepdestroyersheepdestroyer commented Aug 12, 2026

Copy link
Copy Markdown
Owner

Closes#345

Summary by Sourcery

Harden routing and verification infrastructure around LiteLLM/OpenRouter and Langfuse, improving auth, key handling, and canonical endpoint behavior.

Bug Fixes:

  • Restore localhost HTTP fallbacks and explicit LLAMA_* environment variable handling when resolving llama server/classifier endpoints.
  • Ensure LiteLLM master key placeholders and misconfigurations fail fast instead of silently proxying with invalid credentials.
  • Fix adaptive router roster purge assertion to accept multiple purge calls.

Enhancements:

  • Enforce bearer client authentication for the /v1/responses endpoint.
  • Introduce centralized validation of the LiteLLM master key and reuse it for proxy execution and internal calls.
  • Register OpenRouter models as LiteLLM DB models on startup, including dynamic loading from config.yaml and static fallback definitions.
  • Improve Langfuse session propagation tracing to support v4 events_only mode and longer polling, and propagate session/user IDs into Langfuse observations and updates.
  • Extend start-stack.sh to export LLAMA_* env vars and render LiteLLM config using derived classifier URLs.

Tests:

  • Update existing routing, lifespan, and responses_api tests to respect new auth and master key validation behavior.
  • Add tests covering OpenRouter model DB registration paths (no key, static fallback, and config-driven) and LiteLLM master key validation failure modes.
  • Add verification coverage for Langfuse events_only mode and non-leak behavior in session propagation checks.

@sourcery-ai

sourcery-aiBot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Reviewer's Guide

Restores and hardens environment-based configuration and verification for LiteLLM/OpenRouter routing, adds strict master-key and client auth validation, improves Langfuse tracing/session propagation, and extends tests/startup scripts to cover the new behavior and canonical endpoint derivation paths.

Sequence diagram for responses_api client auth and master key validation

sequenceDiagram
actor Client
participant Router as responses_api
participant Env as _validate_litellm_master_key
participant LiteLLM as litellm_proxy
Client->>Router: POST /responses (Authorization: Bearer client_token)
Router->>Router: validate Authorization header
Router->>Env: _validate_litellm_master_key()
Env-->>Router: master_key
Router->>LiteLLM: client.post /responses (Authorization: Bearer master_key)
LiteLLM-->>Router: response
Router-->>Client: routed response
Loading

File-Level Changes

ChangeDetailsFiles
Make LLAMA server/classifier endpoint resolution respect explicit environment variables while preserving localhost fallbacks and canonical HTTPS derivation.
  • Introduce LLAMA_SERVER_URL and LLAMA_CLASSIFIER_URL environment overrides when resolving endpoints.
  • Fallback to config values or localhost defaults when env vars are absent.
  • Prefer explicit env URLs over canonical HTTPS derivation in resolution logic.
router/main.py
start-stack.sh
Introduce centralized LiteLLM master key validation and stricter usage across responses and proxy paths.
  • Add _INVALID_MASTER_KEYS set and _validate_litellm_master_key helper that fails fast with HTTP 500 on missing/placeholder keys.
  • Use _validate_litellm_master_key in /v1/responses and execute_proxy instead of raw os.getenv lookups.
  • Update tests to patch LITELLM_MASTER_KEY and cover invalid-key behavior.
router/main.py
router/tests/test_responses_api.py
router/tests/test_routing_behavior.py
tests/conftest.py
Add OpenRouter model DB registration on startup with dynamic config loading, static fallback, and stale deployment purge.
  • Implement _register_openrouter_models_in_db that loads OpenRouter models from litellm config paths or falls back to a static openrouter-auto definition.
  • Purge stale openrouter-* deployments via _purge_stale_deployments before re-registering.
  • Wire OpenRouter registration into FastAPI lifespan and add unit tests for config-based and fallback registration paths.
router/main.py
router/tests/test_lifespan.py
router/tests/test_register_openrouter_models_in_db.py
router/tests/test_sync_adaptive_router_roster.py
Enforce client Bearer authentication and improve Langfuse tracing/session propagation behavior in responses and chat triage.
  • Require a non-empty Bearer Authorization header in responses_api, returning 401 on missing/invalid auth.
  • Pass session_id and user_id into Langfuse start_observation and update calls when available.
  • Extend Langfuse verification script to handle v4 events_only mode, longer polling, and revised leak checks.
  • Update tests to include Authorization headers and expectations for auth enforcement and tracing behavior.
router/main.py
scripts/verification/verify_canonical_endpoints.py
router/tests/test_responses_api.py
Tighten test and startup environment configuration around routing behavior and master key usage.
  • Set default LITELLM_MASTER_KEY and ROUTER_API_KEY in test conftest and specific routing tests.
  • Adjust routing behavior tests to provide Authorization headers and patched master keys.
  • Update start-stack.sh to export LLAMA_SERVER_URL/LLAMA_CLASSIFIER_URL and derive classifier URL in rendered LiteLLM config.
  • Relax sync_adaptive_router_roster purge assertion to allow multiple purge calls.
router/tests/test_routing_behavior.py
tests/conftest.py
start-stack.sh
router/tests/test_sync_adaptive_router_roster.py

Assessment against linked issues

IssueObjectiveAddressedExplanation
#345Restore functional OpenRouter-backed free-tier chat routing for dev canonical verification, including robust LiteLLM master key handling and model registration so auto/free routes (e.g. llm-routing-auto-free, agent-simple-core) can succeed.
#345Fix dev LiteLLM local safety-net routing by correctly deriving and wiring LLAMA_SERVER_URL / LLAMA_CLASSIFIER_URL from configuration/start-stack into router, so the local fallback classifier becomes reachable from the dev container without coupling to a broken production hostname.

Possibly linked issues


Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@coderabbitai

coderabbitaiBot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@sheepdestroyer, you've reached your PR review limit, so we couldn't start this review.

Next review available in:117 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 3f27a272-9764-4dc3-8e95-2f0a8a765e1e

📥 Commits

Reviewing files that changed from the base of the PR and between 830008b and 4e11105.

📒 Files selected for processing (10)
  • .env.dev
  • router/main.py
  • router/tests/test_lifespan.py
  • router/tests/test_register_openrouter_models_in_db.py
  • router/tests/test_responses_api.py
  • router/tests/test_routing_behavior.py
  • router/tests/test_sync_adaptive_router_roster.py
  • scripts/verification/verify_canonical_endpoints.py
  • start-stack.sh
  • tests/conftest.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sourcery-aisourcery-aiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 2 issues, and left some high level feedback:

  • Avoid logging the raw LITELLM_MASTER_KEY value in _validate_litellm_master_key, as the current logger.error call can expose secrets in logs; consider logging only that it is missing/invalid or a redacted version.
  • The client auth check in responses_api only validates the presence and shape of a Bearer token; if stronger authentication/authorization is expected, consider integrating actual token validation or clarifying that this endpoint is only gating on header presence.
Prompt for AI Agents
Please address the comments from this code review:
## Overall Comments- Avoid logging the raw LITELLM_MASTER_KEY value in `_validate_litellm_master_key`, as the current `logger.error` call can expose secrets in logs; consider logging only that it is missing/invalid or a redacted version.
- The client auth check in `responses_api` only validates the presence and shape of a Bearer token; if stronger authentication/authorization is expected, consider integrating actual token validation or clarifying that this endpoint is only gating on header presence.
## Individual Comments### Comment 1
<locationpath="router/main.py"line_range="797-806" />
<code_context>
+}
++
+def _validate_litellm_master_key() -> str:
+ """Validate LITELLM_MASTER_KEY environment variable.
++ Returns:
+ The valid master key string.++ Raises:
+ HTTPException(500): If master key is missing, empty, or placeholder string.+ """
+ key = (os.getenv("LITELLM_MASTER_KEY") or "").strip()
+ if not key or key in _INVALID_MASTER_KEYS or "PLACEHOLDER" in key.upper():
+ logger.error(f"Invalid or missing LITELLM_MASTER_KEY: '{key}'")
</code_context>
<issue_to_address>
**🚨 issue (security):** Avoid logging the raw master key value to prevent leaking secrets.
The `logger.error` line logs the full `LITELLM_MASTER_KEY`, which risks exposing a production secret in logs. Please avoid logging the raw key—log only that it is invalid/missing, or mask it (e.g., show a short prefix or just its length) while preserving enough context for debugging.
</issue_to_address>
### Comment 2
<locationpath="router/main.py"line_range="2306-2310" />
<code_context>
when an auto model (e.g. llm-routing-auto-free) is requested, while supporting model aliases
(such as gpt-4o-mini, local-qwen-3.6-hass) and tool/streaming executions.
"""
+# Enforce client authentication+ auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
+ if not auth_header or not auth_header.startswith("Bearer "):
+ raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")+ client_token = auth_header[7:].strip()
+ if not client_token:
+ raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
</code_context>
<issue_to_address>
**suggestion (bug_risk):** Authorization handling is case-sensitive and extracts a token that is never used.
`auth_header.startswith("Bearer ")` will reject headers with different casing or spacing (e.g. `bearer <token>`). Consider normalizing (e.g. lowercasing and/or splitting on whitespace) before checking to make the auth handling more robust. Also, `client_token` is extracted but never used; either pass it to downstream logic where needed or remove it to avoid confusion.
Suggested implementation:
```pythontry:
await _register_ollama_models_in_db(litellm_master_key)
exceptExceptionas e:
when an auto model (e.g. llm-routing-auto-free) is requested, while supporting model aliases
(such as gpt-4o-mini, local-qwen-3.6-hass) and tool/streaming executions.
""" # Enforce client authentication auth_header = request.headers.get("Authorization") or request.headers.get("authorization") if not auth_header: raise HTTPException(status_code=401, detail="Missing or invalid Authorization header") parts = auth_header.split() if len(parts) != 2 or parts[0].lower() != "bearer" or not parts[1]: raise HTTPException(status_code=401, detail="Missing or invalid Authorization header") # Store the normalized client token for downstream use request.state.client_token = parts[1] try:```1. Anywhere downstream that needs the client token should read it from `request.state.client_token` instead of re-parsing headers.2. If your router uses dependency injection (e.g. FastAPI dependencies) for auth, consider integrating this parsing into a shared dependency to avoid duplication.</issue_to_address>

Sourcery is free for open source - if you like our reviews please consider sharing them ✨
Help me be more useful! Please click 👍 or 👎 on each comment and I'll use the feedback to improve your reviews.

Comment threadrouter/main.py
Comment on lines +797 to +806
def _validate_litellm_master_key() -> str:
"""Validate LITELLM_MASTER_KEY environment variable.

Returns:
The valid master key string.

Raises:
HTTPException(500): If master key is missing, empty, or placeholder string.
"""
key = (os.getenv("LITELLM_MASTER_KEY") or "").strip()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🚨 issue (security): Avoid logging the raw master key value to prevent leaking secrets.

The logger.error line logs the full LITELLM_MASTER_KEY, which risks exposing a production secret in logs. Please avoid logging the raw key—log only that it is invalid/missing, or mask it (e.g., show a short prefix or just its length) while preserving enough context for debugging.

Comment threadrouter/main.py
Comment on lines +2306 to +2310
# Enforce client authentication
auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
if not auth_header or not auth_header.startswith("Bearer "):
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
client_token = auth_header[7:].strip()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

suggestion (bug_risk): Authorization handling is case-sensitive and extracts a token that is never used.

auth_header.startswith("Bearer ") will reject headers with different casing or spacing (e.g. bearer <token>). Consider normalizing (e.g. lowercasing and/or splitting on whitespace) before checking to make the auth handling more robust. Also, client_token is extracted but never used; either pass it to downstream logic where needed or remove it to avoid confusion.

Suggested implementation:

try:
await_register_ollama_models_in_db(litellm_master_key)
exceptExceptionase:
whenanautomodel (e.g. llm-routing-auto-free) isrequested, whilesupportingmodelaliases
(suchasgpt-4o-mini, local-qwen-3.6-hass) andtool/streamingexecutions.
"""
# Enforce client authentication
auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
if not auth_header:
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
parts = auth_header.split()
if len(parts) != 2 or parts[0].lower() != "bearer" or not parts[1]:
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
# Store the normalized client token for downstream userequest.state.client_token=parts[1]
try:
  1. Anywhere downstream that needs the client token should read it from request.state.client_token instead of re-parsing headers.
  2. If your router uses dependency injection (e.g. FastAPI dependencies) for auth, consider integrating this parsing into a shared dependency to avoid duplication.

@sheepdestroyer
sheepdestroyer merged commit 105e7b2 into masterAug 13, 2026
6 of 7 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(dev): restore fallback capacity for canonical E2E chat verification

1 participant

@sheepdestroyer
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(dev): restore fallback capacity for canonical E2E chat verification (#345) by sheepdestroyer · Pull Request #468 · sheepdestroyer/LLM-Routing · GitHub
Skip to content

fix(dev): restore fallback capacity for canonical E2E chat verification (#345) - #468

Merged
sheepdestroyer merged 3 commits into
masterfrom
fix/dev-fallback-capacity-verification
Aug 13, 2026
Merged

fix(dev): restore fallback capacity for canonical E2E chat verification (#345)#468
sheepdestroyer merged 3 commits into
masterfrom
fix/dev-fallback-capacity-verification

Conversation

@sheepdestroyer

@sheepdestroyersheepdestroyer commented Aug 12, 2026

Copy link
Copy Markdown
Owner

Closes#345

Summary by Sourcery

Harden routing and verification infrastructure around LiteLLM/OpenRouter and Langfuse, improving auth, key handling, and canonical endpoint behavior.

Bug Fixes:

  • Restore localhost HTTP fallbacks and explicit LLAMA_* environment variable handling when resolving llama server/classifier endpoints.
  • Ensure LiteLLM master key placeholders and misconfigurations fail fast instead of silently proxying with invalid credentials.
  • Fix adaptive router roster purge assertion to accept multiple purge calls.

Enhancements:

  • Enforce bearer client authentication for the /v1/responses endpoint.
  • Introduce centralized validation of the LiteLLM master key and reuse it for proxy execution and internal calls.
  • Register OpenRouter models as LiteLLM DB models on startup, including dynamic loading from config.yaml and static fallback definitions.
  • Improve Langfuse session propagation tracing to support v4 events_only mode and longer polling, and propagate session/user IDs into Langfuse observations and updates.
  • Extend start-stack.sh to export LLAMA_* env vars and render LiteLLM config using derived classifier URLs.

Tests:

  • Update existing routing, lifespan, and responses_api tests to respect new auth and master key validation behavior.
  • Add tests covering OpenRouter model DB registration paths (no key, static fallback, and config-driven) and LiteLLM master key validation failure modes.
  • Add verification coverage for Langfuse events_only mode and non-leak behavior in session propagation checks.

@sourcery-ai

sourcery-aiBot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Reviewer's Guide

Restores and hardens environment-based configuration and verification for LiteLLM/OpenRouter routing, adds strict master-key and client auth validation, improves Langfuse tracing/session propagation, and extends tests/startup scripts to cover the new behavior and canonical endpoint derivation paths.

Sequence diagram for responses_api client auth and master key validation

sequenceDiagram
actor Client
participant Router as responses_api
participant Env as _validate_litellm_master_key
participant LiteLLM as litellm_proxy
Client->>Router: POST /responses (Authorization: Bearer client_token)
Router->>Router: validate Authorization header
Router->>Env: _validate_litellm_master_key()
Env-->>Router: master_key
Router->>LiteLLM: client.post /responses (Authorization: Bearer master_key)
LiteLLM-->>Router: response
Router-->>Client: routed response
Loading

File-Level Changes

ChangeDetailsFiles
Make LLAMA server/classifier endpoint resolution respect explicit environment variables while preserving localhost fallbacks and canonical HTTPS derivation.
  • Introduce LLAMA_SERVER_URL and LLAMA_CLASSIFIER_URL environment overrides when resolving endpoints.
  • Fallback to config values or localhost defaults when env vars are absent.
  • Prefer explicit env URLs over canonical HTTPS derivation in resolution logic.
router/main.py
start-stack.sh
Introduce centralized LiteLLM master key validation and stricter usage across responses and proxy paths.
  • Add _INVALID_MASTER_KEYS set and _validate_litellm_master_key helper that fails fast with HTTP 500 on missing/placeholder keys.
  • Use _validate_litellm_master_key in /v1/responses and execute_proxy instead of raw os.getenv lookups.
  • Update tests to patch LITELLM_MASTER_KEY and cover invalid-key behavior.
router/main.py
router/tests/test_responses_api.py
router/tests/test_routing_behavior.py
tests/conftest.py
Add OpenRouter model DB registration on startup with dynamic config loading, static fallback, and stale deployment purge.
  • Implement _register_openrouter_models_in_db that loads OpenRouter models from litellm config paths or falls back to a static openrouter-auto definition.
  • Purge stale openrouter-* deployments via _purge_stale_deployments before re-registering.
  • Wire OpenRouter registration into FastAPI lifespan and add unit tests for config-based and fallback registration paths.
router/main.py
router/tests/test_lifespan.py
router/tests/test_register_openrouter_models_in_db.py
router/tests/test_sync_adaptive_router_roster.py
Enforce client Bearer authentication and improve Langfuse tracing/session propagation behavior in responses and chat triage.
  • Require a non-empty Bearer Authorization header in responses_api, returning 401 on missing/invalid auth.
  • Pass session_id and user_id into Langfuse start_observation and update calls when available.
  • Extend Langfuse verification script to handle v4 events_only mode, longer polling, and revised leak checks.
  • Update tests to include Authorization headers and expectations for auth enforcement and tracing behavior.
router/main.py
scripts/verification/verify_canonical_endpoints.py
router/tests/test_responses_api.py
Tighten test and startup environment configuration around routing behavior and master key usage.
  • Set default LITELLM_MASTER_KEY and ROUTER_API_KEY in test conftest and specific routing tests.
  • Adjust routing behavior tests to provide Authorization headers and patched master keys.
  • Update start-stack.sh to export LLAMA_SERVER_URL/LLAMA_CLASSIFIER_URL and derive classifier URL in rendered LiteLLM config.
  • Relax sync_adaptive_router_roster purge assertion to allow multiple purge calls.
router/tests/test_routing_behavior.py
tests/conftest.py
start-stack.sh
router/tests/test_sync_adaptive_router_roster.py

Assessment against linked issues

IssueObjectiveAddressedExplanation
#345Restore functional OpenRouter-backed free-tier chat routing for dev canonical verification, including robust LiteLLM master key handling and model registration so auto/free routes (e.g. llm-routing-auto-free, agent-simple-core) can succeed.
#345Fix dev LiteLLM local safety-net routing by correctly deriving and wiring LLAMA_SERVER_URL / LLAMA_CLASSIFIER_URL from configuration/start-stack into router, so the local fallback classifier becomes reachable from the dev container without coupling to a broken production hostname.

Possibly linked issues


Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@coderabbitai

coderabbitaiBot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@sheepdestroyer, you've reached your PR review limit, so we couldn't start this review.

Next review available in:117 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 3f27a272-9764-4dc3-8e95-2f0a8a765e1e

📥 Commits

Reviewing files that changed from the base of the PR and between 830008b and 4e11105.

📒 Files selected for processing (10)
  • .env.dev
  • router/main.py
  • router/tests/test_lifespan.py
  • router/tests/test_register_openrouter_models_in_db.py
  • router/tests/test_responses_api.py
  • router/tests/test_routing_behavior.py
  • router/tests/test_sync_adaptive_router_roster.py
  • scripts/verification/verify_canonical_endpoints.py
  • start-stack.sh
  • tests/conftest.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sourcery-aisourcery-aiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 2 issues, and left some high level feedback:

  • Avoid logging the raw LITELLM_MASTER_KEY value in _validate_litellm_master_key, as the current logger.error call can expose secrets in logs; consider logging only that it is missing/invalid or a redacted version.
  • The client auth check in responses_api only validates the presence and shape of a Bearer token; if stronger authentication/authorization is expected, consider integrating actual token validation or clarifying that this endpoint is only gating on header presence.
Prompt for AI Agents
Please address the comments from this code review:
## Overall Comments- Avoid logging the raw LITELLM_MASTER_KEY value in `_validate_litellm_master_key`, as the current `logger.error` call can expose secrets in logs; consider logging only that it is missing/invalid or a redacted version.
- The client auth check in `responses_api` only validates the presence and shape of a Bearer token; if stronger authentication/authorization is expected, consider integrating actual token validation or clarifying that this endpoint is only gating on header presence.
## Individual Comments### Comment 1
<locationpath="router/main.py"line_range="797-806" />
<code_context>
+}
++
+def _validate_litellm_master_key() -> str:
+ """Validate LITELLM_MASTER_KEY environment variable.
++ Returns:
+ The valid master key string.++ Raises:
+ HTTPException(500): If master key is missing, empty, or placeholder string.+ """
+ key = (os.getenv("LITELLM_MASTER_KEY") or "").strip()
+ if not key or key in _INVALID_MASTER_KEYS or "PLACEHOLDER" in key.upper():
+ logger.error(f"Invalid or missing LITELLM_MASTER_KEY: '{key}'")
</code_context>
<issue_to_address>
**🚨 issue (security):** Avoid logging the raw master key value to prevent leaking secrets.
The `logger.error` line logs the full `LITELLM_MASTER_KEY`, which risks exposing a production secret in logs. Please avoid logging the raw key—log only that it is invalid/missing, or mask it (e.g., show a short prefix or just its length) while preserving enough context for debugging.
</issue_to_address>
### Comment 2
<locationpath="router/main.py"line_range="2306-2310" />
<code_context>
when an auto model (e.g. llm-routing-auto-free) is requested, while supporting model aliases
(such as gpt-4o-mini, local-qwen-3.6-hass) and tool/streaming executions.
"""
+# Enforce client authentication+ auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
+ if not auth_header or not auth_header.startswith("Bearer "):
+ raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")+ client_token = auth_header[7:].strip()
+ if not client_token:
+ raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
</code_context>
<issue_to_address>
**suggestion (bug_risk):** Authorization handling is case-sensitive and extracts a token that is never used.
`auth_header.startswith("Bearer ")` will reject headers with different casing or spacing (e.g. `bearer <token>`). Consider normalizing (e.g. lowercasing and/or splitting on whitespace) before checking to make the auth handling more robust. Also, `client_token` is extracted but never used; either pass it to downstream logic where needed or remove it to avoid confusion.
Suggested implementation:
```pythontry:
await _register_ollama_models_in_db(litellm_master_key)
exceptExceptionas e:
when an auto model (e.g. llm-routing-auto-free) is requested, while supporting model aliases
(such as gpt-4o-mini, local-qwen-3.6-hass) and tool/streaming executions.
""" # Enforce client authentication auth_header = request.headers.get("Authorization") or request.headers.get("authorization") if not auth_header: raise HTTPException(status_code=401, detail="Missing or invalid Authorization header") parts = auth_header.split() if len(parts) != 2 or parts[0].lower() != "bearer" or not parts[1]: raise HTTPException(status_code=401, detail="Missing or invalid Authorization header") # Store the normalized client token for downstream use request.state.client_token = parts[1] try:```1. Anywhere downstream that needs the client token should read it from `request.state.client_token` instead of re-parsing headers.2. If your router uses dependency injection (e.g. FastAPI dependencies) for auth, consider integrating this parsing into a shared dependency to avoid duplication.</issue_to_address>

Sourcery is free for open source - if you like our reviews please consider sharing them ✨
Help me be more useful! Please click 👍 or 👎 on each comment and I'll use the feedback to improve your reviews.

Comment threadrouter/main.py
Comment on lines +797 to +806
def _validate_litellm_master_key() -> str:
"""Validate LITELLM_MASTER_KEY environment variable.

Returns:
The valid master key string.

Raises:
HTTPException(500): If master key is missing, empty, or placeholder string.
"""
key = (os.getenv("LITELLM_MASTER_KEY") or "").strip()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🚨 issue (security): Avoid logging the raw master key value to prevent leaking secrets.

The logger.error line logs the full LITELLM_MASTER_KEY, which risks exposing a production secret in logs. Please avoid logging the raw key—log only that it is invalid/missing, or mask it (e.g., show a short prefix or just its length) while preserving enough context for debugging.

Comment threadrouter/main.py
Comment on lines +2306 to +2310
# Enforce client authentication
auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
if not auth_header or not auth_header.startswith("Bearer "):
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
client_token = auth_header[7:].strip()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

suggestion (bug_risk): Authorization handling is case-sensitive and extracts a token that is never used.

auth_header.startswith("Bearer ") will reject headers with different casing or spacing (e.g. bearer <token>). Consider normalizing (e.g. lowercasing and/or splitting on whitespace) before checking to make the auth handling more robust. Also, client_token is extracted but never used; either pass it to downstream logic where needed or remove it to avoid confusion.

Suggested implementation:

try:
await_register_ollama_models_in_db(litellm_master_key)
exceptExceptionase:
whenanautomodel (e.g. llm-routing-auto-free) isrequested, whilesupportingmodelaliases
(suchasgpt-4o-mini, local-qwen-3.6-hass) andtool/streamingexecutions.
"""
# Enforce client authentication
auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
if not auth_header:
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
parts = auth_header.split()
if len(parts) != 2 or parts[0].lower() != "bearer" or not parts[1]:
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
# Store the normalized client token for downstream userequest.state.client_token=parts[1]
try:
  1. Anywhere downstream that needs the client token should read it from request.state.client_token instead of re-parsing headers.
  2. If your router uses dependency injection (e.g. FastAPI dependencies) for auth, consider integrating this parsing into a shared dependency to avoid duplication.

@sheepdestroyer
sheepdestroyer merged commit 105e7b2 into masterAug 13, 2026
6 of 7 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(dev): restore fallback capacity for canonical E2E chat verification

1 participant

@sheepdestroyer
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); fix(dev): restore fallback capacity for canonical E2E chat verification (#345) by sheepdestroyer · Pull Request #468 · sheepdestroyer/LLM-Routing · GitHub
Skip to content

fix(dev): restore fallback capacity for canonical E2E chat verification (#345) - #468

Merged
sheepdestroyer merged 3 commits into
masterfrom
fix/dev-fallback-capacity-verification
Aug 13, 2026
Merged

fix(dev): restore fallback capacity for canonical E2E chat verification (#345)#468
sheepdestroyer merged 3 commits into
masterfrom
fix/dev-fallback-capacity-verification

Conversation

@sheepdestroyer

@sheepdestroyersheepdestroyer commented Aug 12, 2026

Copy link
Copy Markdown
Owner

Closes#345

Summary by Sourcery

Harden routing and verification infrastructure around LiteLLM/OpenRouter and Langfuse, improving auth, key handling, and canonical endpoint behavior.

Bug Fixes:

  • Restore localhost HTTP fallbacks and explicit LLAMA_* environment variable handling when resolving llama server/classifier endpoints.
  • Ensure LiteLLM master key placeholders and misconfigurations fail fast instead of silently proxying with invalid credentials.
  • Fix adaptive router roster purge assertion to accept multiple purge calls.

Enhancements:

  • Enforce bearer client authentication for the /v1/responses endpoint.
  • Introduce centralized validation of the LiteLLM master key and reuse it for proxy execution and internal calls.
  • Register OpenRouter models as LiteLLM DB models on startup, including dynamic loading from config.yaml and static fallback definitions.
  • Improve Langfuse session propagation tracing to support v4 events_only mode and longer polling, and propagate session/user IDs into Langfuse observations and updates.
  • Extend start-stack.sh to export LLAMA_* env vars and render LiteLLM config using derived classifier URLs.

Tests:

  • Update existing routing, lifespan, and responses_api tests to respect new auth and master key validation behavior.
  • Add tests covering OpenRouter model DB registration paths (no key, static fallback, and config-driven) and LiteLLM master key validation failure modes.
  • Add verification coverage for Langfuse events_only mode and non-leak behavior in session propagation checks.

@sourcery-ai

sourcery-aiBot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Reviewer's Guide

Restores and hardens environment-based configuration and verification for LiteLLM/OpenRouter routing, adds strict master-key and client auth validation, improves Langfuse tracing/session propagation, and extends tests/startup scripts to cover the new behavior and canonical endpoint derivation paths.

Sequence diagram for responses_api client auth and master key validation

sequenceDiagram
actor Client
participant Router as responses_api
participant Env as _validate_litellm_master_key
participant LiteLLM as litellm_proxy
Client->>Router: POST /responses (Authorization: Bearer client_token)
Router->>Router: validate Authorization header
Router->>Env: _validate_litellm_master_key()
Env-->>Router: master_key
Router->>LiteLLM: client.post /responses (Authorization: Bearer master_key)
LiteLLM-->>Router: response
Router-->>Client: routed response
Loading

File-Level Changes

ChangeDetailsFiles
Make LLAMA server/classifier endpoint resolution respect explicit environment variables while preserving localhost fallbacks and canonical HTTPS derivation.
  • Introduce LLAMA_SERVER_URL and LLAMA_CLASSIFIER_URL environment overrides when resolving endpoints.
  • Fallback to config values or localhost defaults when env vars are absent.
  • Prefer explicit env URLs over canonical HTTPS derivation in resolution logic.
router/main.py
start-stack.sh
Introduce centralized LiteLLM master key validation and stricter usage across responses and proxy paths.
  • Add _INVALID_MASTER_KEYS set and _validate_litellm_master_key helper that fails fast with HTTP 500 on missing/placeholder keys.
  • Use _validate_litellm_master_key in /v1/responses and execute_proxy instead of raw os.getenv lookups.
  • Update tests to patch LITELLM_MASTER_KEY and cover invalid-key behavior.
router/main.py
router/tests/test_responses_api.py
router/tests/test_routing_behavior.py
tests/conftest.py
Add OpenRouter model DB registration on startup with dynamic config loading, static fallback, and stale deployment purge.
  • Implement _register_openrouter_models_in_db that loads OpenRouter models from litellm config paths or falls back to a static openrouter-auto definition.
  • Purge stale openrouter-* deployments via _purge_stale_deployments before re-registering.
  • Wire OpenRouter registration into FastAPI lifespan and add unit tests for config-based and fallback registration paths.
router/main.py
router/tests/test_lifespan.py
router/tests/test_register_openrouter_models_in_db.py
router/tests/test_sync_adaptive_router_roster.py
Enforce client Bearer authentication and improve Langfuse tracing/session propagation behavior in responses and chat triage.
  • Require a non-empty Bearer Authorization header in responses_api, returning 401 on missing/invalid auth.
  • Pass session_id and user_id into Langfuse start_observation and update calls when available.
  • Extend Langfuse verification script to handle v4 events_only mode, longer polling, and revised leak checks.
  • Update tests to include Authorization headers and expectations for auth enforcement and tracing behavior.
router/main.py
scripts/verification/verify_canonical_endpoints.py
router/tests/test_responses_api.py
Tighten test and startup environment configuration around routing behavior and master key usage.
  • Set default LITELLM_MASTER_KEY and ROUTER_API_KEY in test conftest and specific routing tests.
  • Adjust routing behavior tests to provide Authorization headers and patched master keys.
  • Update start-stack.sh to export LLAMA_SERVER_URL/LLAMA_CLASSIFIER_URL and derive classifier URL in rendered LiteLLM config.
  • Relax sync_adaptive_router_roster purge assertion to allow multiple purge calls.
router/tests/test_routing_behavior.py
tests/conftest.py
start-stack.sh
router/tests/test_sync_adaptive_router_roster.py

Assessment against linked issues

IssueObjectiveAddressedExplanation
#345Restore functional OpenRouter-backed free-tier chat routing for dev canonical verification, including robust LiteLLM master key handling and model registration so auto/free routes (e.g. llm-routing-auto-free, agent-simple-core) can succeed.
#345Fix dev LiteLLM local safety-net routing by correctly deriving and wiring LLAMA_SERVER_URL / LLAMA_CLASSIFIER_URL from configuration/start-stack into router, so the local fallback classifier becomes reachable from the dev container without coupling to a broken production hostname.

Possibly linked issues


Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@coderabbitai

coderabbitaiBot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@sheepdestroyer, you've reached your PR review limit, so we couldn't start this review.

Next review available in:117 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 3f27a272-9764-4dc3-8e95-2f0a8a765e1e

📥 Commits

Reviewing files that changed from the base of the PR and between 830008b and 4e11105.

📒 Files selected for processing (10)
  • .env.dev
  • router/main.py
  • router/tests/test_lifespan.py
  • router/tests/test_register_openrouter_models_in_db.py
  • router/tests/test_responses_api.py
  • router/tests/test_routing_behavior.py
  • router/tests/test_sync_adaptive_router_roster.py
  • scripts/verification/verify_canonical_endpoints.py
  • start-stack.sh
  • tests/conftest.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sourcery-aisourcery-aiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 2 issues, and left some high level feedback:

  • Avoid logging the raw LITELLM_MASTER_KEY value in _validate_litellm_master_key, as the current logger.error call can expose secrets in logs; consider logging only that it is missing/invalid or a redacted version.
  • The client auth check in responses_api only validates the presence and shape of a Bearer token; if stronger authentication/authorization is expected, consider integrating actual token validation or clarifying that this endpoint is only gating on header presence.
Prompt for AI Agents
Please address the comments from this code review:
## Overall Comments- Avoid logging the raw LITELLM_MASTER_KEY value in `_validate_litellm_master_key`, as the current `logger.error` call can expose secrets in logs; consider logging only that it is missing/invalid or a redacted version.
- The client auth check in `responses_api` only validates the presence and shape of a Bearer token; if stronger authentication/authorization is expected, consider integrating actual token validation or clarifying that this endpoint is only gating on header presence.
## Individual Comments### Comment 1
<locationpath="router/main.py"line_range="797-806" />
<code_context>
+}
++
+def _validate_litellm_master_key() -> str:
+ """Validate LITELLM_MASTER_KEY environment variable.
++ Returns:
+ The valid master key string.++ Raises:
+ HTTPException(500): If master key is missing, empty, or placeholder string.+ """
+ key = (os.getenv("LITELLM_MASTER_KEY") or "").strip()
+ if not key or key in _INVALID_MASTER_KEYS or "PLACEHOLDER" in key.upper():
+ logger.error(f"Invalid or missing LITELLM_MASTER_KEY: '{key}'")
</code_context>
<issue_to_address>
**🚨 issue (security):** Avoid logging the raw master key value to prevent leaking secrets.
The `logger.error` line logs the full `LITELLM_MASTER_KEY`, which risks exposing a production secret in logs. Please avoid logging the raw key—log only that it is invalid/missing, or mask it (e.g., show a short prefix or just its length) while preserving enough context for debugging.
</issue_to_address>
### Comment 2
<locationpath="router/main.py"line_range="2306-2310" />
<code_context>
when an auto model (e.g. llm-routing-auto-free) is requested, while supporting model aliases
(such as gpt-4o-mini, local-qwen-3.6-hass) and tool/streaming executions.
"""
+# Enforce client authentication+ auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
+ if not auth_header or not auth_header.startswith("Bearer "):
+ raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")+ client_token = auth_header[7:].strip()
+ if not client_token:
+ raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
</code_context>
<issue_to_address>
**suggestion (bug_risk):** Authorization handling is case-sensitive and extracts a token that is never used.
`auth_header.startswith("Bearer ")` will reject headers with different casing or spacing (e.g. `bearer <token>`). Consider normalizing (e.g. lowercasing and/or splitting on whitespace) before checking to make the auth handling more robust. Also, `client_token` is extracted but never used; either pass it to downstream logic where needed or remove it to avoid confusion.
Suggested implementation:
```pythontry:
await _register_ollama_models_in_db(litellm_master_key)
exceptExceptionas e:
when an auto model (e.g. llm-routing-auto-free) is requested, while supporting model aliases
(such as gpt-4o-mini, local-qwen-3.6-hass) and tool/streaming executions.
""" # Enforce client authentication auth_header = request.headers.get("Authorization") or request.headers.get("authorization") if not auth_header: raise HTTPException(status_code=401, detail="Missing or invalid Authorization header") parts = auth_header.split() if len(parts) != 2 or parts[0].lower() != "bearer" or not parts[1]: raise HTTPException(status_code=401, detail="Missing or invalid Authorization header") # Store the normalized client token for downstream use request.state.client_token = parts[1] try:```1. Anywhere downstream that needs the client token should read it from `request.state.client_token` instead of re-parsing headers.2. If your router uses dependency injection (e.g. FastAPI dependencies) for auth, consider integrating this parsing into a shared dependency to avoid duplication.</issue_to_address>

Sourcery is free for open source - if you like our reviews please consider sharing them ✨
Help me be more useful! Please click 👍 or 👎 on each comment and I'll use the feedback to improve your reviews.

Comment threadrouter/main.py
Comment on lines +797 to +806
def _validate_litellm_master_key() -> str:
"""Validate LITELLM_MASTER_KEY environment variable.

Returns:
The valid master key string.

Raises:
HTTPException(500): If master key is missing, empty, or placeholder string.
"""
key = (os.getenv("LITELLM_MASTER_KEY") or "").strip()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🚨 issue (security): Avoid logging the raw master key value to prevent leaking secrets.

The logger.error line logs the full LITELLM_MASTER_KEY, which risks exposing a production secret in logs. Please avoid logging the raw key—log only that it is invalid/missing, or mask it (e.g., show a short prefix or just its length) while preserving enough context for debugging.

Comment threadrouter/main.py
Comment on lines +2306 to +2310
# Enforce client authentication
auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
if not auth_header or not auth_header.startswith("Bearer "):
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
client_token = auth_header[7:].strip()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

suggestion (bug_risk): Authorization handling is case-sensitive and extracts a token that is never used.

auth_header.startswith("Bearer ") will reject headers with different casing or spacing (e.g. bearer <token>). Consider normalizing (e.g. lowercasing and/or splitting on whitespace) before checking to make the auth handling more robust. Also, client_token is extracted but never used; either pass it to downstream logic where needed or remove it to avoid confusion.

Suggested implementation:

try:
await_register_ollama_models_in_db(litellm_master_key)
exceptExceptionase:
whenanautomodel (e.g. llm-routing-auto-free) isrequested, whilesupportingmodelaliases
(suchasgpt-4o-mini, local-qwen-3.6-hass) andtool/streamingexecutions.
"""
# Enforce client authentication
auth_header = request.headers.get("Authorization") or request.headers.get("authorization")
if not auth_header:
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
parts = auth_header.split()
if len(parts) != 2 or parts[0].lower() != "bearer" or not parts[1]:
raise HTTPException(status_code=401, detail="Missing or invalid Authorization header")
# Store the normalized client token for downstream userequest.state.client_token=parts[1]
try:
  1. Anywhere downstream that needs the client token should read it from request.state.client_token instead of re-parsing headers.
  2. If your router uses dependency injection (e.g. FastAPI dependencies) for auth, consider integrating this parsing into a shared dependency to avoid duplication.

@sheepdestroyer
sheepdestroyer merged commit 105e7b2 into masterAug 13, 2026
6 of 7 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(dev): restore fallback capacity for canonical E2E chat verification

1 participant

@sheepdestroyer