test(e2e): C8 structured outputs / JSON mode passthrough - #172

Merged
moonming merged 2 commits into
mainfrom
test/e2e-c8-json-mode
May 9, 2026
Merged

test(e2e): C8 structured outputs / JSON mode passthrough#172
moonming merged 2 commits into
mainfrom
test/e2e-c8-json-mode

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Third sub-area of #151's C8 row (multimodal/tools). The first two surfaced product gaps:

JSON mode is the simpler request-passthrough field — no new content-block parsing, no cross-provider translation needed, just propagation of the `response_format` field.

What's pinned

CaseUser journeyAsserts
`json_object``response_format: { type: "json_object" }`Upstream's body has the exact `{type: "json_object"}` field; caller-side content (a JSON string) reaches the SDK byte-for-byte and parses as valid JSON
`json_schema``response_format: { type: "json_schema", json_schema: { name, schema, strict: true } }`Upstream-side body preserves `type === "json_schema"`, `json_schema.name`, full schema (every property + `required` array + `additionalProperties`), and `strict: true`

Why this matters

Every modern agent / RAG framework that depends on parseable structured output sets `response_format` on every chat completion. A regression that dropped this field would leave the model in default text mode and silently break every JSON-output caller — the symptom is "I expected JSON but got prose" with no error envelope.

Source-blind discipline

Every assertion derives from external contracts:

No internal Rust paths or struct field names referenced.

Test plan

  • `npm test` (full e2e suite) — 37/37 passing locally (was 35)
  • No mock-data-only paths: each case exercises the real `aisix` binary, real etcd config propagation, real OpenAI Node SDK reverse-call
  • CI green

Refs #151.

CopilotAI review requested due to automatic review settings May 9, 2026 14:44
@coderabbitai

coderabbitaiBot commented May 9, 2026

Copy link
Copy Markdown

Warning

Rate limit exceeded

@moonming has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 10 minutes and 24 seconds before requesting another review.

You’ve run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: b3857ceb-1c54-4744-849f-a1b8d2f3b47d

📥 Commits

Reviewing files that changed from the base of the PR and between 43d5180 and 518d5fe.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/json-mode-e2e.test.ts

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds end-to-end coverage for OpenAI Chat Completions structured outputs (“JSON mode”) to ensure the gateway forwards response_format through to the upstream unchanged, using the real openai Node SDK + the existing OpenAI mock upstream harness.

Changes:

  • Introduces a new e2e suite that exercises /v1/chat/completions with response_format: { type: "json_object" }.
  • Adds a second e2e case for response_format: { type: "json_schema", json_schema: { name, schema, strict: true } } and asserts upstream body preservation.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +178 to +190
test("response_format: { type: 'json_schema', json_schema: {...} } forwarded verbatim", async (ctx) => {
if (!etcdReachable || !app || !upstream) {
ctx.skip();
return;
}

const client = new OpenAI({
apiKey: CALLER_PLAINTEXT,
baseURL: `${app.proxyUrl}/v1`,
maxRetries: 0,
});

// OpenAI's structured-outputs json_schema descriptor per
Comment on lines +211 to +218
const baseline = upstream.receivedRequests.length;
await client.chat.completions.create({
model: "jsonmode-e2e",
messages: [
{ role: "user", content: "Return facts about SF as JSON." },
],
response_format: responseFormat,
});
moonming added a commit that referenced this pull request May 9, 2026
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
moonming added a commit that referenced this pull request May 9, 2026
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
CopilotAI review requested due to automatic review settings May 9, 2026 14:50
@moonming
moonmingforce-pushed the test/e2e-c8-json-mode branch from 263aec4 to fa0c4d0CompareMay 9, 2026 14:50

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment on lines +211 to +218
const baseline = upstream.receivedRequests.length;
await client.chat.completions.create({
model: "jsonmode-e2e",
messages: [
{ role: "user", content: "Return facts about SF as JSON." },
],
response_format: responseFormat,
});
moonming added 2 commits May 9, 2026 23:24
Third C8 sub-area covered. The first two (#170 tools-cross-provider,
#171 vision input parser) surfaced product gaps and are held back
pending fixes. JSON mode is the simpler request-passthrough field —
no new content-block parsing, no cross-provider translation — and
it works.
Two cases pinned:
1. `response_format: { type: "json_object" }` — caller asserts
upstream-side body has `response_format: {type: "json_object"}`
verbatim, and the JSON-shape content reaches the SDK
byte-for-byte and parses as valid JSON.
2. `response_format: { type: "json_schema", json_schema: {name,
schema, strict} }` — caller asserts the full descriptor
(name, full JSON schema, `strict: true`) reaches upstream
verbatim. A regression that dropped any sub-field — especially
`strict` — would relax upstream's constraints and break
schema-conforming callers.
Both cases are real production paths: every modern agent / RAG
framework that depends on parseable structured output sets
`response_format` on every chat completion. A regression that
dropped this field would leave the model in default text mode
and silently break every JSON-output caller.
References:
- OpenAI structured-outputs guide
<https://platform.openai.com/docs/guides/structured-outputs>
- OpenAI Chat Completions response_format spec
<https://platform.openai.com/docs/api-reference/chat/create#chat-create-response_format>
Refs #151
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
@moonming
moonmingforce-pushed the test/e2e-c8-json-mode branch from fa0c4d0 to 518d5feCompareMay 9, 2026 15:24
@moonming
moonming merged commit 0f5e503 into mainMay 9, 2026
6 checks passed
@moonming
moonming deleted the test/e2e-c8-json-mode branch May 9, 2026 15:32
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

test(e2e): C8 structured outputs / JSON mode passthrough - #172

Merged
moonming merged 2 commits into
mainfrom
test/e2e-c8-json-mode
May 9, 2026
Merged

test(e2e): C8 structured outputs / JSON mode passthrough#172
moonming merged 2 commits into
mainfrom
test/e2e-c8-json-mode

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Third sub-area of #151's C8 row (multimodal/tools). The first two surfaced product gaps:

JSON mode is the simpler request-passthrough field — no new content-block parsing, no cross-provider translation needed, just propagation of the `response_format` field.

What's pinned

CaseUser journeyAsserts
`json_object``response_format: { type: "json_object" }`Upstream's body has the exact `{type: "json_object"}` field; caller-side content (a JSON string) reaches the SDK byte-for-byte and parses as valid JSON
`json_schema``response_format: { type: "json_schema", json_schema: { name, schema, strict: true } }`Upstream-side body preserves `type === "json_schema"`, `json_schema.name`, full schema (every property + `required` array + `additionalProperties`), and `strict: true`

Why this matters

Every modern agent / RAG framework that depends on parseable structured output sets `response_format` on every chat completion. A regression that dropped this field would leave the model in default text mode and silently break every JSON-output caller — the symptom is "I expected JSON but got prose" with no error envelope.

Source-blind discipline

Every assertion derives from external contracts:

No internal Rust paths or struct field names referenced.

Test plan

  • `npm test` (full e2e suite) — 37/37 passing locally (was 35)
  • No mock-data-only paths: each case exercises the real `aisix` binary, real etcd config propagation, real OpenAI Node SDK reverse-call
  • CI green

Refs #151.

CopilotAI review requested due to automatic review settings May 9, 2026 14:44
@coderabbitai

coderabbitaiBot commented May 9, 2026

Copy link
Copy Markdown

Warning

Rate limit exceeded

@moonming has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 10 minutes and 24 seconds before requesting another review.

You’ve run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: b3857ceb-1c54-4744-849f-a1b8d2f3b47d

📥 Commits

Reviewing files that changed from the base of the PR and between 43d5180 and 518d5fe.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/json-mode-e2e.test.ts

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds end-to-end coverage for OpenAI Chat Completions structured outputs (“JSON mode”) to ensure the gateway forwards response_format through to the upstream unchanged, using the real openai Node SDK + the existing OpenAI mock upstream harness.

Changes:

  • Introduces a new e2e suite that exercises /v1/chat/completions with response_format: { type: "json_object" }.
  • Adds a second e2e case for response_format: { type: "json_schema", json_schema: { name, schema, strict: true } } and asserts upstream body preservation.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +178 to +190
test("response_format: { type: 'json_schema', json_schema: {...} } forwarded verbatim", async (ctx) => {
if (!etcdReachable || !app || !upstream) {
ctx.skip();
return;
}

const client = new OpenAI({
apiKey: CALLER_PLAINTEXT,
baseURL: `${app.proxyUrl}/v1`,
maxRetries: 0,
});

// OpenAI's structured-outputs json_schema descriptor per
Comment on lines +211 to +218
const baseline = upstream.receivedRequests.length;
await client.chat.completions.create({
model: "jsonmode-e2e",
messages: [
{ role: "user", content: "Return facts about SF as JSON." },
],
response_format: responseFormat,
});
moonming added a commit that referenced this pull request May 9, 2026
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
moonming added a commit that referenced this pull request May 9, 2026
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
CopilotAI review requested due to automatic review settings May 9, 2026 14:50
@moonming
moonmingforce-pushed the test/e2e-c8-json-mode branch from 263aec4 to fa0c4d0CompareMay 9, 2026 14:50

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment on lines +211 to +218
const baseline = upstream.receivedRequests.length;
await client.chat.completions.create({
model: "jsonmode-e2e",
messages: [
{ role: "user", content: "Return facts about SF as JSON." },
],
response_format: responseFormat,
});
moonming added 2 commits May 9, 2026 23:24
Third C8 sub-area covered. The first two (#170 tools-cross-provider,
#171 vision input parser) surfaced product gaps and are held back
pending fixes. JSON mode is the simpler request-passthrough field —
no new content-block parsing, no cross-provider translation — and
it works.
Two cases pinned:
1. `response_format: { type: "json_object" }` — caller asserts
upstream-side body has `response_format: {type: "json_object"}`
verbatim, and the JSON-shape content reaches the SDK
byte-for-byte and parses as valid JSON.
2. `response_format: { type: "json_schema", json_schema: {name,
schema, strict} }` — caller asserts the full descriptor
(name, full JSON schema, `strict: true`) reaches upstream
verbatim. A regression that dropped any sub-field — especially
`strict` — would relax upstream's constraints and break
schema-conforming callers.
Both cases are real production paths: every modern agent / RAG
framework that depends on parseable structured output sets
`response_format` on every chat completion. A regression that
dropped this field would leave the model in default text mode
and silently break every JSON-output caller.
References:
- OpenAI structured-outputs guide
<https://platform.openai.com/docs/guides/structured-outputs>
- OpenAI Chat Completions response_format spec
<https://platform.openai.com/docs/api-reference/chat/create#chat-create-response_format>
Refs #151
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
@moonming
moonmingforce-pushed the test/e2e-c8-json-mode branch from fa0c4d0 to 518d5feCompareMay 9, 2026 15:24
@moonming
moonming merged commit 0f5e503 into mainMay 9, 2026
6 checks passed
@moonming
moonming deleted the test/e2e-c8-json-mode branch May 9, 2026 15:32
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

test(e2e): C8 structured outputs / JSON mode passthrough - #172

Merged
moonming merged 2 commits into
mainfrom
test/e2e-c8-json-mode
May 9, 2026
Merged

test(e2e): C8 structured outputs / JSON mode passthrough#172
moonming merged 2 commits into
mainfrom
test/e2e-c8-json-mode

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Third sub-area of #151's C8 row (multimodal/tools). The first two surfaced product gaps:

JSON mode is the simpler request-passthrough field — no new content-block parsing, no cross-provider translation needed, just propagation of the `response_format` field.

What's pinned

CaseUser journeyAsserts
`json_object``response_format: { type: "json_object" }`Upstream's body has the exact `{type: "json_object"}` field; caller-side content (a JSON string) reaches the SDK byte-for-byte and parses as valid JSON
`json_schema``response_format: { type: "json_schema", json_schema: { name, schema, strict: true } }`Upstream-side body preserves `type === "json_schema"`, `json_schema.name`, full schema (every property + `required` array + `additionalProperties`), and `strict: true`

Why this matters

Every modern agent / RAG framework that depends on parseable structured output sets `response_format` on every chat completion. A regression that dropped this field would leave the model in default text mode and silently break every JSON-output caller — the symptom is "I expected JSON but got prose" with no error envelope.

Source-blind discipline

Every assertion derives from external contracts:

No internal Rust paths or struct field names referenced.

Test plan

  • `npm test` (full e2e suite) — 37/37 passing locally (was 35)
  • No mock-data-only paths: each case exercises the real `aisix` binary, real etcd config propagation, real OpenAI Node SDK reverse-call
  • CI green

Refs #151.

CopilotAI review requested due to automatic review settings May 9, 2026 14:44
@coderabbitai

coderabbitaiBot commented May 9, 2026

Copy link
Copy Markdown

Warning

Rate limit exceeded

@moonming has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 10 minutes and 24 seconds before requesting another review.

You’ve run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: b3857ceb-1c54-4744-849f-a1b8d2f3b47d

📥 Commits

Reviewing files that changed from the base of the PR and between 43d5180 and 518d5fe.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/json-mode-e2e.test.ts

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds end-to-end coverage for OpenAI Chat Completions structured outputs (“JSON mode”) to ensure the gateway forwards response_format through to the upstream unchanged, using the real openai Node SDK + the existing OpenAI mock upstream harness.

Changes:

  • Introduces a new e2e suite that exercises /v1/chat/completions with response_format: { type: "json_object" }.
  • Adds a second e2e case for response_format: { type: "json_schema", json_schema: { name, schema, strict: true } } and asserts upstream body preservation.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +178 to +190
test("response_format: { type: 'json_schema', json_schema: {...} } forwarded verbatim", async (ctx) => {
if (!etcdReachable || !app || !upstream) {
ctx.skip();
return;
}

const client = new OpenAI({
apiKey: CALLER_PLAINTEXT,
baseURL: `${app.proxyUrl}/v1`,
maxRetries: 0,
});

// OpenAI's structured-outputs json_schema descriptor per
Comment on lines +211 to +218
const baseline = upstream.receivedRequests.length;
await client.chat.completions.create({
model: "jsonmode-e2e",
messages: [
{ role: "user", content: "Return facts about SF as JSON." },
],
response_format: responseFormat,
});
moonming added a commit that referenced this pull request May 9, 2026
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
moonming added a commit that referenced this pull request May 9, 2026
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
CopilotAI review requested due to automatic review settings May 9, 2026 14:50
@moonming
moonmingforce-pushed the test/e2e-c8-json-mode branch from 263aec4 to fa0c4d0CompareMay 9, 2026 14:50

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment on lines +211 to +218
const baseline = upstream.receivedRequests.length;
await client.chat.completions.create({
model: "jsonmode-e2e",
messages: [
{ role: "user", content: "Return facts about SF as JSON." },
],
response_format: responseFormat,
});
moonming added 2 commits May 9, 2026 23:24
Third C8 sub-area covered. The first two (#170 tools-cross-provider,
#171 vision input parser) surfaced product gaps and are held back
pending fixes. JSON mode is the simpler request-passthrough field —
no new content-block parsing, no cross-provider translation — and
it works.
Two cases pinned:
1. `response_format: { type: "json_object" }` — caller asserts
upstream-side body has `response_format: {type: "json_object"}`
verbatim, and the JSON-shape content reaches the SDK
byte-for-byte and parses as valid JSON.
2. `response_format: { type: "json_schema", json_schema: {name,
schema, strict} }` — caller asserts the full descriptor
(name, full JSON schema, `strict: true`) reaches upstream
verbatim. A regression that dropped any sub-field — especially
`strict` — would relax upstream's constraints and break
schema-conforming callers.
Both cases are real production paths: every modern agent / RAG
framework that depends on parseable structured output sets
`response_format` on every chat completion. A regression that
dropped this field would leave the model in default text mode
and silently break every JSON-output caller.
References:
- OpenAI structured-outputs guide
<https://platform.openai.com/docs/guides/structured-outputs>
- OpenAI Chat Completions response_format spec
<https://platform.openai.com/docs/api-reference/chat/create#chat-create-response_format>
Refs #151
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
@moonming
moonmingforce-pushed the test/e2e-c8-json-mode branch from fa0c4d0 to 518d5feCompareMay 9, 2026 15:24
@moonming
moonming merged commit 0f5e503 into mainMay 9, 2026
6 checks passed
@moonming
moonming deleted the test/e2e-c8-json-mode branch May 9, 2026 15:32
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

test(e2e): C8 structured outputs / JSON mode passthrough - #172

Merged
moonming merged 2 commits into
mainfrom
test/e2e-c8-json-mode
May 9, 2026
Merged

test(e2e): C8 structured outputs / JSON mode passthrough#172
moonming merged 2 commits into
mainfrom
test/e2e-c8-json-mode

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Third sub-area of #151's C8 row (multimodal/tools). The first two surfaced product gaps:

JSON mode is the simpler request-passthrough field — no new content-block parsing, no cross-provider translation needed, just propagation of the `response_format` field.

What's pinned

CaseUser journeyAsserts
`json_object``response_format: { type: "json_object" }`Upstream's body has the exact `{type: "json_object"}` field; caller-side content (a JSON string) reaches the SDK byte-for-byte and parses as valid JSON
`json_schema``response_format: { type: "json_schema", json_schema: { name, schema, strict: true } }`Upstream-side body preserves `type === "json_schema"`, `json_schema.name`, full schema (every property + `required` array + `additionalProperties`), and `strict: true`

Why this matters

Every modern agent / RAG framework that depends on parseable structured output sets `response_format` on every chat completion. A regression that dropped this field would leave the model in default text mode and silently break every JSON-output caller — the symptom is "I expected JSON but got prose" with no error envelope.

Source-blind discipline

Every assertion derives from external contracts:

No internal Rust paths or struct field names referenced.

Test plan

  • `npm test` (full e2e suite) — 37/37 passing locally (was 35)
  • No mock-data-only paths: each case exercises the real `aisix` binary, real etcd config propagation, real OpenAI Node SDK reverse-call
  • CI green

Refs #151.

CopilotAI review requested due to automatic review settings May 9, 2026 14:44
@coderabbitai

coderabbitaiBot commented May 9, 2026

Copy link
Copy Markdown

Warning

Rate limit exceeded

@moonming has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 10 minutes and 24 seconds before requesting another review.

You’ve run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: b3857ceb-1c54-4744-849f-a1b8d2f3b47d

📥 Commits

Reviewing files that changed from the base of the PR and between 43d5180 and 518d5fe.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/json-mode-e2e.test.ts

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds end-to-end coverage for OpenAI Chat Completions structured outputs (“JSON mode”) to ensure the gateway forwards response_format through to the upstream unchanged, using the real openai Node SDK + the existing OpenAI mock upstream harness.

Changes:

  • Introduces a new e2e suite that exercises /v1/chat/completions with response_format: { type: "json_object" }.
  • Adds a second e2e case for response_format: { type: "json_schema", json_schema: { name, schema, strict: true } } and asserts upstream body preservation.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +178 to +190
test("response_format: { type: 'json_schema', json_schema: {...} } forwarded verbatim", async (ctx) => {
if (!etcdReachable || !app || !upstream) {
ctx.skip();
return;
}

const client = new OpenAI({
apiKey: CALLER_PLAINTEXT,
baseURL: `${app.proxyUrl}/v1`,
maxRetries: 0,
});

// OpenAI's structured-outputs json_schema descriptor per
Comment on lines +211 to +218
const baseline = upstream.receivedRequests.length;
await client.chat.completions.create({
model: "jsonmode-e2e",
messages: [
{ role: "user", content: "Return facts about SF as JSON." },
],
response_format: responseFormat,
});
moonming added a commit that referenced this pull request May 9, 2026
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
moonming added a commit that referenced this pull request May 9, 2026
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
CopilotAI review requested due to automatic review settings May 9, 2026 14:50
@moonming
moonmingforce-pushed the test/e2e-c8-json-mode branch from 263aec4 to fa0c4d0CompareMay 9, 2026 14:50

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment on lines +211 to +218
const baseline = upstream.receivedRequests.length;
await client.chat.completions.create({
model: "jsonmode-e2e",
messages: [
{ role: "user", content: "Return facts about SF as JSON." },
],
response_format: responseFormat,
});
moonming added 2 commits May 9, 2026 23:24
Third C8 sub-area covered. The first two (#170 tools-cross-provider,
#171 vision input parser) surfaced product gaps and are held back
pending fixes. JSON mode is the simpler request-passthrough field —
no new content-block parsing, no cross-provider translation — and
it works.
Two cases pinned:
1. `response_format: { type: "json_object" }` — caller asserts
upstream-side body has `response_format: {type: "json_object"}`
verbatim, and the JSON-shape content reaches the SDK
byte-for-byte and parses as valid JSON.
2. `response_format: { type: "json_schema", json_schema: {name,
schema, strict} }` — caller asserts the full descriptor
(name, full JSON schema, `strict: true`) reaches upstream
verbatim. A regression that dropped any sub-field — especially
`strict` — would relax upstream's constraints and break
schema-conforming callers.
Both cases are real production paths: every modern agent / RAG
framework that depends on parseable structured output sets
`response_format` on every chat completion. A regression that
dropped this field would leave the model in default text mode
and silently break every JSON-output caller.
References:
- OpenAI structured-outputs guide
<https://platform.openai.com/docs/guides/structured-outputs>
- OpenAI Chat Completions response_format spec
<https://platform.openai.com/docs/api-reference/chat/create#chat-create-response_format>
Refs #151
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
@moonming
moonmingforce-pushed the test/e2e-c8-json-mode branch from fa0c4d0 to 518d5feCompareMay 9, 2026 15:24
@moonming
moonming merged commit 0f5e503 into mainMay 9, 2026
6 checks passed
@moonming
moonming deleted the test/e2e-c8-json-mode branch May 9, 2026 15:32
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

test(e2e): C8 structured outputs / JSON mode passthrough - #172

Merged
moonming merged 2 commits into
mainfrom
test/e2e-c8-json-mode
May 9, 2026
Merged

test(e2e): C8 structured outputs / JSON mode passthrough#172
moonming merged 2 commits into
mainfrom
test/e2e-c8-json-mode

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Third sub-area of #151's C8 row (multimodal/tools). The first two surfaced product gaps:

JSON mode is the simpler request-passthrough field — no new content-block parsing, no cross-provider translation needed, just propagation of the `response_format` field.

What's pinned

CaseUser journeyAsserts
`json_object``response_format: { type: "json_object" }`Upstream's body has the exact `{type: "json_object"}` field; caller-side content (a JSON string) reaches the SDK byte-for-byte and parses as valid JSON
`json_schema``response_format: { type: "json_schema", json_schema: { name, schema, strict: true } }`Upstream-side body preserves `type === "json_schema"`, `json_schema.name`, full schema (every property + `required` array + `additionalProperties`), and `strict: true`

Why this matters

Every modern agent / RAG framework that depends on parseable structured output sets `response_format` on every chat completion. A regression that dropped this field would leave the model in default text mode and silently break every JSON-output caller — the symptom is "I expected JSON but got prose" with no error envelope.

Source-blind discipline

Every assertion derives from external contracts:

No internal Rust paths or struct field names referenced.

Test plan

  • `npm test` (full e2e suite) — 37/37 passing locally (was 35)
  • No mock-data-only paths: each case exercises the real `aisix` binary, real etcd config propagation, real OpenAI Node SDK reverse-call
  • CI green

Refs #151.

CopilotAI review requested due to automatic review settings May 9, 2026 14:44
@coderabbitai

coderabbitaiBot commented May 9, 2026

Copy link
Copy Markdown

Warning

Rate limit exceeded

@moonming has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 10 minutes and 24 seconds before requesting another review.

You’ve run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: b3857ceb-1c54-4744-849f-a1b8d2f3b47d

📥 Commits

Reviewing files that changed from the base of the PR and between 43d5180 and 518d5fe.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/json-mode-e2e.test.ts

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds end-to-end coverage for OpenAI Chat Completions structured outputs (“JSON mode”) to ensure the gateway forwards response_format through to the upstream unchanged, using the real openai Node SDK + the existing OpenAI mock upstream harness.

Changes:

  • Introduces a new e2e suite that exercises /v1/chat/completions with response_format: { type: "json_object" }.
  • Adds a second e2e case for response_format: { type: "json_schema", json_schema: { name, schema, strict: true } } and asserts upstream body preservation.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +178 to +190
test("response_format: { type: 'json_schema', json_schema: {...} } forwarded verbatim", async (ctx) => {
if (!etcdReachable || !app || !upstream) {
ctx.skip();
return;
}

const client = new OpenAI({
apiKey: CALLER_PLAINTEXT,
baseURL: `${app.proxyUrl}/v1`,
maxRetries: 0,
});

// OpenAI's structured-outputs json_schema descriptor per
Comment on lines +211 to +218
const baseline = upstream.receivedRequests.length;
await client.chat.completions.create({
model: "jsonmode-e2e",
messages: [
{ role: "user", content: "Return facts about SF as JSON." },
],
response_format: responseFormat,
});
moonming added a commit that referenced this pull request May 9, 2026
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
moonming added a commit that referenced this pull request May 9, 2026
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
CopilotAI review requested due to automatic review settings May 9, 2026 14:50
@moonming
moonmingforce-pushed the test/e2e-c8-json-mode branch from 263aec4 to fa0c4d0CompareMay 9, 2026 14:50

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment on lines +211 to +218
const baseline = upstream.receivedRequests.length;
await client.chat.completions.create({
model: "jsonmode-e2e",
messages: [
{ role: "user", content: "Return facts about SF as JSON." },
],
response_format: responseFormat,
});
moonming added 2 commits May 9, 2026 23:24
Third C8 sub-area covered. The first two (#170 tools-cross-provider,
#171 vision input parser) surfaced product gaps and are held back
pending fixes. JSON mode is the simpler request-passthrough field —
no new content-block parsing, no cross-provider translation — and
it works.
Two cases pinned:
1. `response_format: { type: "json_object" }` — caller asserts
upstream-side body has `response_format: {type: "json_object"}`
verbatim, and the JSON-shape content reaches the SDK
byte-for-byte and parses as valid JSON.
2. `response_format: { type: "json_schema", json_schema: {name,
schema, strict} }` — caller asserts the full descriptor
(name, full JSON schema, `strict: true`) reaches upstream
verbatim. A regression that dropped any sub-field — especially
`strict` — would relax upstream's constraints and break
schema-conforming callers.
Both cases are real production paths: every modern agent / RAG
framework that depends on parseable structured output sets
`response_format` on every chat completion. A regression that
dropped this field would leave the model in default text mode
and silently break every JSON-output caller.
References:
- OpenAI structured-outputs guide
<https://platform.openai.com/docs/guides/structured-outputs>
- OpenAI Chat Completions response_format spec
<https://platform.openai.com/docs/api-reference/chat/create#chat-create-response_format>
Refs #151
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
@moonming
moonmingforce-pushed the test/e2e-c8-json-mode branch from fa0c4d0 to 518d5feCompareMay 9, 2026 15:24
@moonming
moonming merged commit 0f5e503 into mainMay 9, 2026
6 checks passed
@moonming
moonming deleted the test/e2e-c8-json-mode branch May 9, 2026 15:32
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

test(e2e): C8 structured outputs / JSON mode passthrough - #172

Merged
moonming merged 2 commits into
mainfrom
test/e2e-c8-json-mode
May 9, 2026
Merged

test(e2e): C8 structured outputs / JSON mode passthrough#172
moonming merged 2 commits into
mainfrom
test/e2e-c8-json-mode

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Third sub-area of #151's C8 row (multimodal/tools). The first two surfaced product gaps:

JSON mode is the simpler request-passthrough field — no new content-block parsing, no cross-provider translation needed, just propagation of the `response_format` field.

What's pinned

CaseUser journeyAsserts
`json_object``response_format: { type: "json_object" }`Upstream's body has the exact `{type: "json_object"}` field; caller-side content (a JSON string) reaches the SDK byte-for-byte and parses as valid JSON
`json_schema``response_format: { type: "json_schema", json_schema: { name, schema, strict: true } }`Upstream-side body preserves `type === "json_schema"`, `json_schema.name`, full schema (every property + `required` array + `additionalProperties`), and `strict: true`

Why this matters

Every modern agent / RAG framework that depends on parseable structured output sets `response_format` on every chat completion. A regression that dropped this field would leave the model in default text mode and silently break every JSON-output caller — the symptom is "I expected JSON but got prose" with no error envelope.

Source-blind discipline

Every assertion derives from external contracts:

No internal Rust paths or struct field names referenced.

Test plan

  • `npm test` (full e2e suite) — 37/37 passing locally (was 35)
  • No mock-data-only paths: each case exercises the real `aisix` binary, real etcd config propagation, real OpenAI Node SDK reverse-call
  • CI green

Refs #151.

CopilotAI review requested due to automatic review settings May 9, 2026 14:44
@coderabbitai

coderabbitaiBot commented May 9, 2026

Copy link
Copy Markdown

Warning

Rate limit exceeded

@moonming has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 10 minutes and 24 seconds before requesting another review.

You’ve run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: b3857ceb-1c54-4744-849f-a1b8d2f3b47d

📥 Commits

Reviewing files that changed from the base of the PR and between 43d5180 and 518d5fe.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/json-mode-e2e.test.ts

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds end-to-end coverage for OpenAI Chat Completions structured outputs (“JSON mode”) to ensure the gateway forwards response_format through to the upstream unchanged, using the real openai Node SDK + the existing OpenAI mock upstream harness.

Changes:

  • Introduces a new e2e suite that exercises /v1/chat/completions with response_format: { type: "json_object" }.
  • Adds a second e2e case for response_format: { type: "json_schema", json_schema: { name, schema, strict: true } } and asserts upstream body preservation.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +178 to +190
test("response_format: { type: 'json_schema', json_schema: {...} } forwarded verbatim", async (ctx) => {
if (!etcdReachable || !app || !upstream) {
ctx.skip();
return;
}

const client = new OpenAI({
apiKey: CALLER_PLAINTEXT,
baseURL: `${app.proxyUrl}/v1`,
maxRetries: 0,
});

// OpenAI's structured-outputs json_schema descriptor per
Comment on lines +211 to +218
const baseline = upstream.receivedRequests.length;
await client.chat.completions.create({
model: "jsonmode-e2e",
messages: [
{ role: "user", content: "Return facts about SF as JSON." },
],
response_format: responseFormat,
});
moonming added a commit that referenced this pull request May 9, 2026
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
moonming added a commit that referenced this pull request May 9, 2026
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
CopilotAI review requested due to automatic review settings May 9, 2026 14:50
@moonming
moonmingforce-pushed the test/e2e-c8-json-mode branch from 263aec4 to fa0c4d0CompareMay 9, 2026 14:50

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment on lines +211 to +218
const baseline = upstream.receivedRequests.length;
await client.chat.completions.create({
model: "jsonmode-e2e",
messages: [
{ role: "user", content: "Return facts about SF as JSON." },
],
response_format: responseFormat,
});
moonming added 2 commits May 9, 2026 23:24
Third C8 sub-area covered. The first two (#170 tools-cross-provider,
#171 vision input parser) surfaced product gaps and are held back
pending fixes. JSON mode is the simpler request-passthrough field —
no new content-block parsing, no cross-provider translation — and
it works.
Two cases pinned:
1. `response_format: { type: "json_object" }` — caller asserts
upstream-side body has `response_format: {type: "json_object"}`
verbatim, and the JSON-shape content reaches the SDK
byte-for-byte and parses as valid JSON.
2. `response_format: { type: "json_schema", json_schema: {name,
schema, strict} }` — caller asserts the full descriptor
(name, full JSON schema, `strict: true`) reaches upstream
verbatim. A regression that dropped any sub-field — especially
`strict` — would relax upstream's constraints and break
schema-conforming callers.
Both cases are real production paths: every modern agent / RAG
framework that depends on parseable structured output sets
`response_format` on every chat completion. A regression that
dropped this field would leave the model in default text mode
and silently break every JSON-output caller.
References:
- OpenAI structured-outputs guide
<https://platform.openai.com/docs/guides/structured-outputs>
- OpenAI Chat Completions response_format spec
<https://platform.openai.com/docs/api-reference/chat/create#chat-create-response_format>
Refs #151
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
@moonming
moonmingforce-pushed the test/e2e-c8-json-mode branch from fa0c4d0 to 518d5feCompareMay 9, 2026 15:24
@moonming
moonming merged commit 0f5e503 into mainMay 9, 2026
6 checks passed
@moonming
moonming deleted the test/e2e-c8-json-mode branch May 9, 2026 15:32
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

test(e2e): C8 structured outputs / JSON mode passthrough - #172

Merged
moonming merged 2 commits into
mainfrom
test/e2e-c8-json-mode
May 9, 2026
Merged

test(e2e): C8 structured outputs / JSON mode passthrough#172
moonming merged 2 commits into
mainfrom
test/e2e-c8-json-mode

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Third sub-area of #151's C8 row (multimodal/tools). The first two surfaced product gaps:

JSON mode is the simpler request-passthrough field — no new content-block parsing, no cross-provider translation needed, just propagation of the `response_format` field.

What's pinned

CaseUser journeyAsserts
`json_object``response_format: { type: "json_object" }`Upstream's body has the exact `{type: "json_object"}` field; caller-side content (a JSON string) reaches the SDK byte-for-byte and parses as valid JSON
`json_schema``response_format: { type: "json_schema", json_schema: { name, schema, strict: true } }`Upstream-side body preserves `type === "json_schema"`, `json_schema.name`, full schema (every property + `required` array + `additionalProperties`), and `strict: true`

Why this matters

Every modern agent / RAG framework that depends on parseable structured output sets `response_format` on every chat completion. A regression that dropped this field would leave the model in default text mode and silently break every JSON-output caller — the symptom is "I expected JSON but got prose" with no error envelope.

Source-blind discipline

Every assertion derives from external contracts:

No internal Rust paths or struct field names referenced.

Test plan

  • `npm test` (full e2e suite) — 37/37 passing locally (was 35)
  • No mock-data-only paths: each case exercises the real `aisix` binary, real etcd config propagation, real OpenAI Node SDK reverse-call
  • CI green

Refs #151.

CopilotAI review requested due to automatic review settings May 9, 2026 14:44
@coderabbitai

coderabbitaiBot commented May 9, 2026

Copy link
Copy Markdown

Warning

Rate limit exceeded

@moonming has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 10 minutes and 24 seconds before requesting another review.

You’ve run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: b3857ceb-1c54-4744-849f-a1b8d2f3b47d

📥 Commits

Reviewing files that changed from the base of the PR and between 43d5180 and 518d5fe.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/json-mode-e2e.test.ts

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds end-to-end coverage for OpenAI Chat Completions structured outputs (“JSON mode”) to ensure the gateway forwards response_format through to the upstream unchanged, using the real openai Node SDK + the existing OpenAI mock upstream harness.

Changes:

  • Introduces a new e2e suite that exercises /v1/chat/completions with response_format: { type: "json_object" }.
  • Adds a second e2e case for response_format: { type: "json_schema", json_schema: { name, schema, strict: true } } and asserts upstream body preservation.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +178 to +190
test("response_format: { type: 'json_schema', json_schema: {...} } forwarded verbatim", async (ctx) => {
if (!etcdReachable || !app || !upstream) {
ctx.skip();
return;
}

const client = new OpenAI({
apiKey: CALLER_PLAINTEXT,
baseURL: `${app.proxyUrl}/v1`,
maxRetries: 0,
});

// OpenAI's structured-outputs json_schema descriptor per
Comment on lines +211 to +218
const baseline = upstream.receivedRequests.length;
await client.chat.completions.create({
model: "jsonmode-e2e",
messages: [
{ role: "user", content: "Return facts about SF as JSON." },
],
response_format: responseFormat,
});
moonming added a commit that referenced this pull request May 9, 2026
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
moonming added a commit that referenced this pull request May 9, 2026
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
CopilotAI review requested due to automatic review settings May 9, 2026 14:50
@moonming
moonmingforce-pushed the test/e2e-c8-json-mode branch from 263aec4 to fa0c4d0CompareMay 9, 2026 14:50

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment on lines +211 to +218
const baseline = upstream.receivedRequests.length;
await client.chat.completions.create({
model: "jsonmode-e2e",
messages: [
{ role: "user", content: "Return facts about SF as JSON." },
],
response_format: responseFormat,
});
moonming added 2 commits May 9, 2026 23:24
Third C8 sub-area covered. The first two (#170 tools-cross-provider,
#171 vision input parser) surfaced product gaps and are held back
pending fixes. JSON mode is the simpler request-passthrough field —
no new content-block parsing, no cross-provider translation — and
it works.
Two cases pinned:
1. `response_format: { type: "json_object" }` — caller asserts
upstream-side body has `response_format: {type: "json_object"}`
verbatim, and the JSON-shape content reaches the SDK
byte-for-byte and parses as valid JSON.
2. `response_format: { type: "json_schema", json_schema: {name,
schema, strict} }` — caller asserts the full descriptor
(name, full JSON schema, `strict: true`) reaches upstream
verbatim. A regression that dropped any sub-field — especially
`strict` — would relax upstream's constraints and break
schema-conforming callers.
Both cases are real production paths: every modern agent / RAG
framework that depends on parseable structured output sets
`response_format` on every chat completion. A regression that
dropped this field would leave the model in default text mode
and silently break every JSON-output caller.
References:
- OpenAI structured-outputs guide
<https://platform.openai.com/docs/guides/structured-outputs>
- OpenAI Chat Completions response_format spec
<https://platform.openai.com/docs/api-reference/chat/create#chat-create-response_format>
Refs #151
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
@moonming
moonmingforce-pushed the test/e2e-c8-json-mode branch from fa0c4d0 to 518d5feCompareMay 9, 2026 15:24
@moonming
moonming merged commit 0f5e503 into mainMay 9, 2026
6 checks passed
@moonming
moonming deleted the test/e2e-c8-json-mode branch May 9, 2026 15:32
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

test(e2e): C8 structured outputs / JSON mode passthrough - #172

Merged
moonming merged 2 commits into
mainfrom
test/e2e-c8-json-mode
May 9, 2026
Merged

test(e2e): C8 structured outputs / JSON mode passthrough#172
moonming merged 2 commits into
mainfrom
test/e2e-c8-json-mode

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

Third sub-area of #151's C8 row (multimodal/tools). The first two surfaced product gaps:

JSON mode is the simpler request-passthrough field — no new content-block parsing, no cross-provider translation needed, just propagation of the `response_format` field.

What's pinned

CaseUser journeyAsserts
`json_object``response_format: { type: "json_object" }`Upstream's body has the exact `{type: "json_object"}` field; caller-side content (a JSON string) reaches the SDK byte-for-byte and parses as valid JSON
`json_schema``response_format: { type: "json_schema", json_schema: { name, schema, strict: true } }`Upstream-side body preserves `type === "json_schema"`, `json_schema.name`, full schema (every property + `required` array + `additionalProperties`), and `strict: true`

Why this matters

Every modern agent / RAG framework that depends on parseable structured output sets `response_format` on every chat completion. A regression that dropped this field would leave the model in default text mode and silently break every JSON-output caller — the symptom is "I expected JSON but got prose" with no error envelope.

Source-blind discipline

Every assertion derives from external contracts:

No internal Rust paths or struct field names referenced.

Test plan

  • `npm test` (full e2e suite) — 37/37 passing locally (was 35)
  • No mock-data-only paths: each case exercises the real `aisix` binary, real etcd config propagation, real OpenAI Node SDK reverse-call
  • CI green

Refs #151.

CopilotAI review requested due to automatic review settings May 9, 2026 14:44
@coderabbitai

coderabbitaiBot commented May 9, 2026

Copy link
Copy Markdown

Warning

Rate limit exceeded

@moonming has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 10 minutes and 24 seconds before requesting another review.

You’ve run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: b3857ceb-1c54-4744-849f-a1b8d2f3b47d

📥 Commits

Reviewing files that changed from the base of the PR and between 43d5180 and 518d5fe.

📒 Files selected for processing (1)
  • tests/e2e/src/cases/json-mode-e2e.test.ts

Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds end-to-end coverage for OpenAI Chat Completions structured outputs (“JSON mode”) to ensure the gateway forwards response_format through to the upstream unchanged, using the real openai Node SDK + the existing OpenAI mock upstream harness.

Changes:

  • Introduces a new e2e suite that exercises /v1/chat/completions with response_format: { type: "json_object" }.
  • Adds a second e2e case for response_format: { type: "json_schema", json_schema: { name, schema, strict: true } } and asserts upstream body preservation.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +178 to +190
test("response_format: { type: 'json_schema', json_schema: {...} } forwarded verbatim", async (ctx) => {
if (!etcdReachable || !app || !upstream) {
ctx.skip();
return;
}

const client = new OpenAI({
apiKey: CALLER_PLAINTEXT,
baseURL: `${app.proxyUrl}/v1`,
maxRetries: 0,
});

// OpenAI's structured-outputs json_schema descriptor per
Comment on lines +211 to +218
const baseline = upstream.receivedRequests.length;
await client.chat.completions.create({
model: "jsonmode-e2e",
messages: [
{ role: "user", content: "Return facts about SF as JSON." },
],
response_format: responseFormat,
});
moonming added a commit that referenced this pull request May 9, 2026
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
moonming added a commit that referenced this pull request May 9, 2026
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
CopilotAI review requested due to automatic review settings May 9, 2026 14:50
@moonming
moonmingforce-pushed the test/e2e-c8-json-mode branch from 263aec4 to fa0c4d0CompareMay 9, 2026 14:50

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment on lines +211 to +218
const baseline = upstream.receivedRequests.length;
await client.chat.completions.create({
model: "jsonmode-e2e",
messages: [
{ role: "user", content: "Return facts about SF as JSON." },
],
response_format: responseFormat,
});
moonming added 2 commits May 9, 2026 23:24
Third C8 sub-area covered. The first two (#170 tools-cross-provider,
#171 vision input parser) surfaced product gaps and are held back
pending fixes. JSON mode is the simpler request-passthrough field —
no new content-block parsing, no cross-provider translation — and
it works.
Two cases pinned:
1. `response_format: { type: "json_object" }` — caller asserts
upstream-side body has `response_format: {type: "json_object"}`
verbatim, and the JSON-shape content reaches the SDK
byte-for-byte and parses as valid JSON.
2. `response_format: { type: "json_schema", json_schema: {name,
schema, strict} }` — caller asserts the full descriptor
(name, full JSON schema, `strict: true`) reaches upstream
verbatim. A regression that dropped any sub-field — especially
`strict` — would relax upstream's constraints and break
schema-conforming callers.
Both cases are real production paths: every modern agent / RAG
framework that depends on parseable structured output sets
`response_format` on every chat completion. A regression that
dropped this field would leave the model in default text mode
and silently break every JSON-output caller.
References:
- OpenAI structured-outputs guide
<https://platform.openai.com/docs/guides/structured-outputs>
- OpenAI Chat Completions response_format spec
<https://platform.openai.com/docs/api-reference/chat/create#chat-create-response_format>
Refs #151
Audit on commit d04abe8 returned 0 HIGH, 1 MEDIUM (rebase-after-#169
infra fix), 3 LOW. Applying LOW-1 here:
LOW-1: case 2 (json_schema) asserted only response_format fields
on the upstream side. A regression that mangled the rest of the
request body while preserving response_format would not surface.
Added cross-check: pin `model` translation (display name → upstream
model_name) and the user message reaches upstream verbatim — same
defense case 1's byte-for-byte content assertion provides on the
response side.
LOW-2 (streaming + response_format coverage gap) and LOW-3 (case
1 comment polish) deferred.
MEDIUM-1 (rebase on #169 once it merges to inherit waitConfig
budget bump + maxForks reduction) addressed by waiting for #169
to land and rebasing before merge.
@moonming
moonmingforce-pushed the test/e2e-c8-json-mode branch from fa0c4d0 to 518d5feCompareMay 9, 2026 15:24
@moonming
moonming merged commit 0f5e503 into mainMay 9, 2026
6 checks passed
@moonming
moonming deleted the test/e2e-c8-json-mode branch May 9, 2026 15:32
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming