fix(server): install rustls default CryptoProvider at startup - #34

Merged
moonming merged 1 commit into
mainfrom
fix/rustls-default-provider
Apr 24, 2026
Merged

fix(server): install rustls default CryptoProvider at startup#34
moonming merged 1 commit into
mainfrom
fix/rustls-default-provider

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

aisix panics on first TLS handshake (etcd connect, reqwest /dp/register, etc.) with:

```
thread 'main' panicked at rustls/.../crypto/mod.rs:249:14:
Could not automatically determine the process-level CryptoProvider from Rustls crate features.
```

rustls 0.23 dropped implicit provider selection — when both `aws-lc-rs` and `ring` are reachable through transitive deps (reqwest `rustls-tls`, etcd-client, tokio-rustls…), the runtime can't pick. Fix: call `aws_lc_rs::default_provider().install_default()` at the very top of `main`, before anything else.

`let _ =` because install is idempotent and the return value only signals "another crate got there first" — which doesn't affect correctness.

Why aws-lc-rs and not ring

It's the upstream rustls default as of 0.23, FIPS-capable out of the box, and already in the transitive dep graph through reqwest + etcd-client. Picking it avoids adding another RSA/EC crypto implementation to the binary.

How we caught it

AISIX-Cloud e2e stack exposed this: the DP container is `docker run` via the test harness, hits this panic within 100ms of startup, exits, and `docker run --rm` removes the container before the harness can query `docker port`. Container logs now dumped on failure.

Test plan

  • `cargo build --bin aisix`
  • `cargo fmt --check`
  • `cargo clippy --all-targets --locked -- -D warnings`
  • CI green
  • Re-publish `ghcr.io/moonming/ai-gateway:aisix-e2e` from main so AISIX-Cloud e2e picks it up

aisix panics on first TLS handshake (etcd connect, reqwest register
call, etc.) with:
thread 'main' panicked at rustls/.../crypto/mod.rs:249:14:
Could not automatically determine the process-level CryptoProvider
from Rustls crate features.
rustls 0.23 dropped implicit provider selection — when both
`aws-lc-rs` and `ring` are reachable through transitive deps (we have
reqwest `rustls-tls`, etcd-client, tokio-rustls… all enabling one or
the other), the runtime can't pick. The panic fires the first time
any TLS operation touches the crypto layer.
Call `aws_lc_rs::default_provider().install_default()` at the very
top of main, before anything else loads. `let _ =` because install
is idempotent and we don't care if another crate beat us to it.
CopilotAI review requested due to automatic review settings April 24, 2026 02:47

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Ensures aisix doesn’t panic on first TLS usage under rustls 0.23 by explicitly installing a process-wide rustls CryptoProvider at startup.

Changes:

  • Install the aws-lc-rs rustls CryptoProvider at the start of main() to avoid runtime provider auto-detection panics when multiple crypto backends are present via transitive dependencies.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +56 to +57
// already depends on transitively. Falls back to ring only if
// the process somehow has a provider installed already (idempotent).

CopilotAIApr 24, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The comment says this "falls back to ring" if a provider is already installed, but the code doesn't install ring or perform any fallback logic; it simply keeps whatever default provider was installed first. Consider rewording to avoid implying ring is involved (e.g., "if another provider is already installed, keep it").

Suggested change
// already depends on transitively. Falls back to ring only if
// the process somehow has a provider installed already (idempotent).
// already depends on transitively. If another provider is already
// installed for the process, `install_default()` keeps it unchanged.

Copilot uses AI. Check for mistakes.
@moonming
moonming merged commit bc06934 into mainApr 24, 2026
8 of 10 checks passed
moonming added a commit that referenced this pull request Apr 26, 2026
…t tests
CI on main has been red since #34 because aisix-core's Config crate
loader merges every AISIX_-prefixed env var into the root Config
struct (config-rs Environment::with_prefix("AISIX")), and Config has
#[serde(deny_unknown_fields)]. The redis integration test sets
AISIX_REDIS_URL on the rust-unit job, which leaks into every
Config::load_from_path call as `redis_url` and panics:
Config("deserialize: unknown field `redis_url`,
expected one of `etcd`, `proxy`, `admin`,
`observability`, `cache`, `managed`")
8 of 9 aisix-core::config::tests fail (the one that doesn't is
rejects_unknown_fields, which intentionally swallows the error).
Rename the env var so it doesn't sit under the AISIX_ prefix at
all. crates/aisix-cache/tests/redis_integration.rs reads
CACHE_TEST_REDIS_URL; CI sets the same. docs/testing.md +
crates/aisix-cache/src/redis.rs comment updated to match.
Verified: with the rename, all 9 config tests pass even with
CACHE_TEST_REDIS_URL set; reproducing with the old AISIX_REDIS_URL
still fails as expected (so the loader behaviour is unchanged for
real AISIX_-prefixed env overrides).
moonming added a commit that referenced this pull request Apr 26, 2026
…40)
* fix(aisix-etcd): make supervisor cache-write tests deterministic
Two supervisor tests waited on the spawned cache write via
`tokio::time::sleep(50ms)`. Under heavy CI load the spawn lost the
race against the disk read that followed, surfacing as:
resync_writes_to_disk_cache_then_restore_replays_it FAILED
put_and_delete_keep_cache_in_sync FAILED
Track the JoinHandle for each spawned write in a `pending_writes`
Mutex<Vec<_>> on the Supervisor, and expose a test-only async
`await_pending_cache_writes` that drains and awaits them. Both tests
now wait on real completion instead of a wall clock.
The new field is `#[cfg(test)]`-friendly via the awaiter — production
code never reads it. If a handle is dropped during shutdown the
underlying write either completed or was cancelled; the on-disk
cache is best-effort, and the next live cycle re-publishes from
etcd anyway.
* fix(ci): rename AISIX_REDIS_URL → CACHE_TEST_REDIS_URL to unblock unit tests
CI on main has been red since #34 because aisix-core's Config crate
loader merges every AISIX_-prefixed env var into the root Config
struct (config-rs Environment::with_prefix("AISIX")), and Config has
#[serde(deny_unknown_fields)]. The redis integration test sets
AISIX_REDIS_URL on the rust-unit job, which leaks into every
Config::load_from_path call as `redis_url` and panics:
Config("deserialize: unknown field `redis_url`,
expected one of `etcd`, `proxy`, `admin`,
`observability`, `cache`, `managed`")
8 of 9 aisix-core::config::tests fail (the one that doesn't is
rejects_unknown_fields, which intentionally swallows the error).
Rename the env var so it doesn't sit under the AISIX_ prefix at
all. crates/aisix-cache/tests/redis_integration.rs reads
CACHE_TEST_REDIS_URL; CI sets the same. docs/testing.md +
crates/aisix-cache/src/redis.rs comment updated to match.
Verified: with the rename, all 9 config tests pass even with
CACHE_TEST_REDIS_URL set; reproducing with the old AISIX_REDIS_URL
still fails as expected (so the loader behaviour is unchanged for
real AISIX_-prefixed env overrides).
* ci: tolerate artifact upload quota errors temporarily
Both `rust unit + coverage` and `build ui` jobs are currently failing
on the upload-artifact step with:
Failed to CreateArtifact: Artifact storage quota has been hit.
Unable to upload any new artifacts.
Usage is recalculated every 6-12 hours.
Tests + clippy pass; only the artifact upload is blocked. Add
`continue-on-error: true` to those two upload-artifact steps so the
test-passing signal isn't masked by the quota issue.
Downstream jobs that need ui-dist (build-bin → e2e) will fail at
download-artifact when the upload was skipped; e2e is already
`continue-on-error: true` at the job level, and coverage-gate is
advisory.
Revert this once the org-level storage usage refreshes (within
6-12h) or the quota is raised.
* ci: soft-fail build-aisix while artifact storage quota persists
build-aisix downloads ui-dist from build-ui. With build-ui's
upload-artifact set to continue-on-error during the storage quota
outage, the download fails and build-aisix errors. Since build-aisix
only feeds the advisory e2e job, mark it continue-on-error too so
the PR doesn't go red on a transitive dependency. Revert with the
other two when storage usage refreshes.
@jarvis9443
jarvis9443 deleted the fix/rustls-default-provider branch June 25, 2026 06:26
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

fix(server): install rustls default CryptoProvider at startup - #34

Merged
moonming merged 1 commit into
mainfrom
fix/rustls-default-provider
Apr 24, 2026
Merged

fix(server): install rustls default CryptoProvider at startup#34
moonming merged 1 commit into
mainfrom
fix/rustls-default-provider

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

aisix panics on first TLS handshake (etcd connect, reqwest /dp/register, etc.) with:

```
thread 'main' panicked at rustls/.../crypto/mod.rs:249:14:
Could not automatically determine the process-level CryptoProvider from Rustls crate features.
```

rustls 0.23 dropped implicit provider selection — when both `aws-lc-rs` and `ring` are reachable through transitive deps (reqwest `rustls-tls`, etcd-client, tokio-rustls…), the runtime can't pick. Fix: call `aws_lc_rs::default_provider().install_default()` at the very top of `main`, before anything else.

`let _ =` because install is idempotent and the return value only signals "another crate got there first" — which doesn't affect correctness.

Why aws-lc-rs and not ring

It's the upstream rustls default as of 0.23, FIPS-capable out of the box, and already in the transitive dep graph through reqwest + etcd-client. Picking it avoids adding another RSA/EC crypto implementation to the binary.

How we caught it

AISIX-Cloud e2e stack exposed this: the DP container is `docker run` via the test harness, hits this panic within 100ms of startup, exits, and `docker run --rm` removes the container before the harness can query `docker port`. Container logs now dumped on failure.

Test plan

  • `cargo build --bin aisix`
  • `cargo fmt --check`
  • `cargo clippy --all-targets --locked -- -D warnings`
  • CI green
  • Re-publish `ghcr.io/moonming/ai-gateway:aisix-e2e` from main so AISIX-Cloud e2e picks it up

aisix panics on first TLS handshake (etcd connect, reqwest register
call, etc.) with:
thread 'main' panicked at rustls/.../crypto/mod.rs:249:14:
Could not automatically determine the process-level CryptoProvider
from Rustls crate features.
rustls 0.23 dropped implicit provider selection — when both
`aws-lc-rs` and `ring` are reachable through transitive deps (we have
reqwest `rustls-tls`, etcd-client, tokio-rustls… all enabling one or
the other), the runtime can't pick. The panic fires the first time
any TLS operation touches the crypto layer.
Call `aws_lc_rs::default_provider().install_default()` at the very
top of main, before anything else loads. `let _ =` because install
is idempotent and we don't care if another crate beat us to it.
CopilotAI review requested due to automatic review settings April 24, 2026 02:47

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Ensures aisix doesn’t panic on first TLS usage under rustls 0.23 by explicitly installing a process-wide rustls CryptoProvider at startup.

Changes:

  • Install the aws-lc-rs rustls CryptoProvider at the start of main() to avoid runtime provider auto-detection panics when multiple crypto backends are present via transitive dependencies.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +56 to +57
// already depends on transitively. Falls back to ring only if
// the process somehow has a provider installed already (idempotent).

CopilotAIApr 24, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The comment says this "falls back to ring" if a provider is already installed, but the code doesn't install ring or perform any fallback logic; it simply keeps whatever default provider was installed first. Consider rewording to avoid implying ring is involved (e.g., "if another provider is already installed, keep it").

Suggested change
// already depends on transitively. Falls back to ring only if
// the process somehow has a provider installed already (idempotent).
// already depends on transitively. If another provider is already
// installed for the process, `install_default()` keeps it unchanged.

Copilot uses AI. Check for mistakes.
@moonming
moonming merged commit bc06934 into mainApr 24, 2026
8 of 10 checks passed
moonming added a commit that referenced this pull request Apr 26, 2026
…t tests
CI on main has been red since #34 because aisix-core's Config crate
loader merges every AISIX_-prefixed env var into the root Config
struct (config-rs Environment::with_prefix("AISIX")), and Config has
#[serde(deny_unknown_fields)]. The redis integration test sets
AISIX_REDIS_URL on the rust-unit job, which leaks into every
Config::load_from_path call as `redis_url` and panics:
Config("deserialize: unknown field `redis_url`,
expected one of `etcd`, `proxy`, `admin`,
`observability`, `cache`, `managed`")
8 of 9 aisix-core::config::tests fail (the one that doesn't is
rejects_unknown_fields, which intentionally swallows the error).
Rename the env var so it doesn't sit under the AISIX_ prefix at
all. crates/aisix-cache/tests/redis_integration.rs reads
CACHE_TEST_REDIS_URL; CI sets the same. docs/testing.md +
crates/aisix-cache/src/redis.rs comment updated to match.
Verified: with the rename, all 9 config tests pass even with
CACHE_TEST_REDIS_URL set; reproducing with the old AISIX_REDIS_URL
still fails as expected (so the loader behaviour is unchanged for
real AISIX_-prefixed env overrides).
moonming added a commit that referenced this pull request Apr 26, 2026
…40)
* fix(aisix-etcd): make supervisor cache-write tests deterministic
Two supervisor tests waited on the spawned cache write via
`tokio::time::sleep(50ms)`. Under heavy CI load the spawn lost the
race against the disk read that followed, surfacing as:
resync_writes_to_disk_cache_then_restore_replays_it FAILED
put_and_delete_keep_cache_in_sync FAILED
Track the JoinHandle for each spawned write in a `pending_writes`
Mutex<Vec<_>> on the Supervisor, and expose a test-only async
`await_pending_cache_writes` that drains and awaits them. Both tests
now wait on real completion instead of a wall clock.
The new field is `#[cfg(test)]`-friendly via the awaiter — production
code never reads it. If a handle is dropped during shutdown the
underlying write either completed or was cancelled; the on-disk
cache is best-effort, and the next live cycle re-publishes from
etcd anyway.
* fix(ci): rename AISIX_REDIS_URL → CACHE_TEST_REDIS_URL to unblock unit tests
CI on main has been red since #34 because aisix-core's Config crate
loader merges every AISIX_-prefixed env var into the root Config
struct (config-rs Environment::with_prefix("AISIX")), and Config has
#[serde(deny_unknown_fields)]. The redis integration test sets
AISIX_REDIS_URL on the rust-unit job, which leaks into every
Config::load_from_path call as `redis_url` and panics:
Config("deserialize: unknown field `redis_url`,
expected one of `etcd`, `proxy`, `admin`,
`observability`, `cache`, `managed`")
8 of 9 aisix-core::config::tests fail (the one that doesn't is
rejects_unknown_fields, which intentionally swallows the error).
Rename the env var so it doesn't sit under the AISIX_ prefix at
all. crates/aisix-cache/tests/redis_integration.rs reads
CACHE_TEST_REDIS_URL; CI sets the same. docs/testing.md +
crates/aisix-cache/src/redis.rs comment updated to match.
Verified: with the rename, all 9 config tests pass even with
CACHE_TEST_REDIS_URL set; reproducing with the old AISIX_REDIS_URL
still fails as expected (so the loader behaviour is unchanged for
real AISIX_-prefixed env overrides).
* ci: tolerate artifact upload quota errors temporarily
Both `rust unit + coverage` and `build ui` jobs are currently failing
on the upload-artifact step with:
Failed to CreateArtifact: Artifact storage quota has been hit.
Unable to upload any new artifacts.
Usage is recalculated every 6-12 hours.
Tests + clippy pass; only the artifact upload is blocked. Add
`continue-on-error: true` to those two upload-artifact steps so the
test-passing signal isn't masked by the quota issue.
Downstream jobs that need ui-dist (build-bin → e2e) will fail at
download-artifact when the upload was skipped; e2e is already
`continue-on-error: true` at the job level, and coverage-gate is
advisory.
Revert this once the org-level storage usage refreshes (within
6-12h) or the quota is raised.
* ci: soft-fail build-aisix while artifact storage quota persists
build-aisix downloads ui-dist from build-ui. With build-ui's
upload-artifact set to continue-on-error during the storage quota
outage, the download fails and build-aisix errors. Since build-aisix
only feeds the advisory e2e job, mark it continue-on-error too so
the PR doesn't go red on a transitive dependency. Revert with the
other two when storage usage refreshes.
@jarvis9443
jarvis9443 deleted the fix/rustls-default-provider branch June 25, 2026 06:26
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): install rustls default CryptoProvider at startup - #34

Merged
moonming merged 1 commit into
mainfrom
fix/rustls-default-provider
Apr 24, 2026
Merged

fix(server): install rustls default CryptoProvider at startup#34
moonming merged 1 commit into
mainfrom
fix/rustls-default-provider

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

aisix panics on first TLS handshake (etcd connect, reqwest /dp/register, etc.) with:

```
thread 'main' panicked at rustls/.../crypto/mod.rs:249:14:
Could not automatically determine the process-level CryptoProvider from Rustls crate features.
```

rustls 0.23 dropped implicit provider selection — when both `aws-lc-rs` and `ring` are reachable through transitive deps (reqwest `rustls-tls`, etcd-client, tokio-rustls…), the runtime can't pick. Fix: call `aws_lc_rs::default_provider().install_default()` at the very top of `main`, before anything else.

`let _ =` because install is idempotent and the return value only signals "another crate got there first" — which doesn't affect correctness.

Why aws-lc-rs and not ring

It's the upstream rustls default as of 0.23, FIPS-capable out of the box, and already in the transitive dep graph through reqwest + etcd-client. Picking it avoids adding another RSA/EC crypto implementation to the binary.

How we caught it

AISIX-Cloud e2e stack exposed this: the DP container is `docker run` via the test harness, hits this panic within 100ms of startup, exits, and `docker run --rm` removes the container before the harness can query `docker port`. Container logs now dumped on failure.

Test plan

  • `cargo build --bin aisix`
  • `cargo fmt --check`
  • `cargo clippy --all-targets --locked -- -D warnings`
  • CI green
  • Re-publish `ghcr.io/moonming/ai-gateway:aisix-e2e` from main so AISIX-Cloud e2e picks it up

aisix panics on first TLS handshake (etcd connect, reqwest register
call, etc.) with:
thread 'main' panicked at rustls/.../crypto/mod.rs:249:14:
Could not automatically determine the process-level CryptoProvider
from Rustls crate features.
rustls 0.23 dropped implicit provider selection — when both
`aws-lc-rs` and `ring` are reachable through transitive deps (we have
reqwest `rustls-tls`, etcd-client, tokio-rustls… all enabling one or
the other), the runtime can't pick. The panic fires the first time
any TLS operation touches the crypto layer.
Call `aws_lc_rs::default_provider().install_default()` at the very
top of main, before anything else loads. `let _ =` because install
is idempotent and we don't care if another crate beat us to it.
CopilotAI review requested due to automatic review settings April 24, 2026 02:47

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Ensures aisix doesn’t panic on first TLS usage under rustls 0.23 by explicitly installing a process-wide rustls CryptoProvider at startup.

Changes:

  • Install the aws-lc-rs rustls CryptoProvider at the start of main() to avoid runtime provider auto-detection panics when multiple crypto backends are present via transitive dependencies.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +56 to +57
// already depends on transitively. Falls back to ring only if
// the process somehow has a provider installed already (idempotent).

CopilotAIApr 24, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The comment says this "falls back to ring" if a provider is already installed, but the code doesn't install ring or perform any fallback logic; it simply keeps whatever default provider was installed first. Consider rewording to avoid implying ring is involved (e.g., "if another provider is already installed, keep it").

Suggested change
// already depends on transitively. Falls back to ring only if
// the process somehow has a provider installed already (idempotent).
// already depends on transitively. If another provider is already
// installed for the process, `install_default()` keeps it unchanged.

Copilot uses AI. Check for mistakes.
@moonming
moonming merged commit bc06934 into mainApr 24, 2026
8 of 10 checks passed
moonming added a commit that referenced this pull request Apr 26, 2026
…t tests
CI on main has been red since #34 because aisix-core's Config crate
loader merges every AISIX_-prefixed env var into the root Config
struct (config-rs Environment::with_prefix("AISIX")), and Config has
#[serde(deny_unknown_fields)]. The redis integration test sets
AISIX_REDIS_URL on the rust-unit job, which leaks into every
Config::load_from_path call as `redis_url` and panics:
Config("deserialize: unknown field `redis_url`,
expected one of `etcd`, `proxy`, `admin`,
`observability`, `cache`, `managed`")
8 of 9 aisix-core::config::tests fail (the one that doesn't is
rejects_unknown_fields, which intentionally swallows the error).
Rename the env var so it doesn't sit under the AISIX_ prefix at
all. crates/aisix-cache/tests/redis_integration.rs reads
CACHE_TEST_REDIS_URL; CI sets the same. docs/testing.md +
crates/aisix-cache/src/redis.rs comment updated to match.
Verified: with the rename, all 9 config tests pass even with
CACHE_TEST_REDIS_URL set; reproducing with the old AISIX_REDIS_URL
still fails as expected (so the loader behaviour is unchanged for
real AISIX_-prefixed env overrides).
moonming added a commit that referenced this pull request Apr 26, 2026
…40)
* fix(aisix-etcd): make supervisor cache-write tests deterministic
Two supervisor tests waited on the spawned cache write via
`tokio::time::sleep(50ms)`. Under heavy CI load the spawn lost the
race against the disk read that followed, surfacing as:
resync_writes_to_disk_cache_then_restore_replays_it FAILED
put_and_delete_keep_cache_in_sync FAILED
Track the JoinHandle for each spawned write in a `pending_writes`
Mutex<Vec<_>> on the Supervisor, and expose a test-only async
`await_pending_cache_writes` that drains and awaits them. Both tests
now wait on real completion instead of a wall clock.
The new field is `#[cfg(test)]`-friendly via the awaiter — production
code never reads it. If a handle is dropped during shutdown the
underlying write either completed or was cancelled; the on-disk
cache is best-effort, and the next live cycle re-publishes from
etcd anyway.
* fix(ci): rename AISIX_REDIS_URL → CACHE_TEST_REDIS_URL to unblock unit tests
CI on main has been red since #34 because aisix-core's Config crate
loader merges every AISIX_-prefixed env var into the root Config
struct (config-rs Environment::with_prefix("AISIX")), and Config has
#[serde(deny_unknown_fields)]. The redis integration test sets
AISIX_REDIS_URL on the rust-unit job, which leaks into every
Config::load_from_path call as `redis_url` and panics:
Config("deserialize: unknown field `redis_url`,
expected one of `etcd`, `proxy`, `admin`,
`observability`, `cache`, `managed`")
8 of 9 aisix-core::config::tests fail (the one that doesn't is
rejects_unknown_fields, which intentionally swallows the error).
Rename the env var so it doesn't sit under the AISIX_ prefix at
all. crates/aisix-cache/tests/redis_integration.rs reads
CACHE_TEST_REDIS_URL; CI sets the same. docs/testing.md +
crates/aisix-cache/src/redis.rs comment updated to match.
Verified: with the rename, all 9 config tests pass even with
CACHE_TEST_REDIS_URL set; reproducing with the old AISIX_REDIS_URL
still fails as expected (so the loader behaviour is unchanged for
real AISIX_-prefixed env overrides).
* ci: tolerate artifact upload quota errors temporarily
Both `rust unit + coverage` and `build ui` jobs are currently failing
on the upload-artifact step with:
Failed to CreateArtifact: Artifact storage quota has been hit.
Unable to upload any new artifacts.
Usage is recalculated every 6-12 hours.
Tests + clippy pass; only the artifact upload is blocked. Add
`continue-on-error: true` to those two upload-artifact steps so the
test-passing signal isn't masked by the quota issue.
Downstream jobs that need ui-dist (build-bin → e2e) will fail at
download-artifact when the upload was skipped; e2e is already
`continue-on-error: true` at the job level, and coverage-gate is
advisory.
Revert this once the org-level storage usage refreshes (within
6-12h) or the quota is raised.
* ci: soft-fail build-aisix while artifact storage quota persists
build-aisix downloads ui-dist from build-ui. With build-ui's
upload-artifact set to continue-on-error during the storage quota
outage, the download fails and build-aisix errors. Since build-aisix
only feeds the advisory e2e job, mark it continue-on-error too so
the PR doesn't go red on a transitive dependency. Revert with the
other two when storage usage refreshes.
@jarvis9443
jarvis9443 deleted the fix/rustls-default-provider branch June 25, 2026 06:26
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): install rustls default CryptoProvider at startup - #34

Merged
moonming merged 1 commit into
mainfrom
fix/rustls-default-provider
Apr 24, 2026
Merged

fix(server): install rustls default CryptoProvider at startup#34
moonming merged 1 commit into
mainfrom
fix/rustls-default-provider

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

aisix panics on first TLS handshake (etcd connect, reqwest /dp/register, etc.) with:

```
thread 'main' panicked at rustls/.../crypto/mod.rs:249:14:
Could not automatically determine the process-level CryptoProvider from Rustls crate features.
```

rustls 0.23 dropped implicit provider selection — when both `aws-lc-rs` and `ring` are reachable through transitive deps (reqwest `rustls-tls`, etcd-client, tokio-rustls…), the runtime can't pick. Fix: call `aws_lc_rs::default_provider().install_default()` at the very top of `main`, before anything else.

`let _ =` because install is idempotent and the return value only signals "another crate got there first" — which doesn't affect correctness.

Why aws-lc-rs and not ring

It's the upstream rustls default as of 0.23, FIPS-capable out of the box, and already in the transitive dep graph through reqwest + etcd-client. Picking it avoids adding another RSA/EC crypto implementation to the binary.

How we caught it

AISIX-Cloud e2e stack exposed this: the DP container is `docker run` via the test harness, hits this panic within 100ms of startup, exits, and `docker run --rm` removes the container before the harness can query `docker port`. Container logs now dumped on failure.

Test plan

  • `cargo build --bin aisix`
  • `cargo fmt --check`
  • `cargo clippy --all-targets --locked -- -D warnings`
  • CI green
  • Re-publish `ghcr.io/moonming/ai-gateway:aisix-e2e` from main so AISIX-Cloud e2e picks it up

aisix panics on first TLS handshake (etcd connect, reqwest register
call, etc.) with:
thread 'main' panicked at rustls/.../crypto/mod.rs:249:14:
Could not automatically determine the process-level CryptoProvider
from Rustls crate features.
rustls 0.23 dropped implicit provider selection — when both
`aws-lc-rs` and `ring` are reachable through transitive deps (we have
reqwest `rustls-tls`, etcd-client, tokio-rustls… all enabling one or
the other), the runtime can't pick. The panic fires the first time
any TLS operation touches the crypto layer.
Call `aws_lc_rs::default_provider().install_default()` at the very
top of main, before anything else loads. `let _ =` because install
is idempotent and we don't care if another crate beat us to it.
CopilotAI review requested due to automatic review settings April 24, 2026 02:47

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Ensures aisix doesn’t panic on first TLS usage under rustls 0.23 by explicitly installing a process-wide rustls CryptoProvider at startup.

Changes:

  • Install the aws-lc-rs rustls CryptoProvider at the start of main() to avoid runtime provider auto-detection panics when multiple crypto backends are present via transitive dependencies.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +56 to +57
// already depends on transitively. Falls back to ring only if
// the process somehow has a provider installed already (idempotent).

CopilotAIApr 24, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The comment says this "falls back to ring" if a provider is already installed, but the code doesn't install ring or perform any fallback logic; it simply keeps whatever default provider was installed first. Consider rewording to avoid implying ring is involved (e.g., "if another provider is already installed, keep it").

Suggested change
// already depends on transitively. Falls back to ring only if
// the process somehow has a provider installed already (idempotent).
// already depends on transitively. If another provider is already
// installed for the process, `install_default()` keeps it unchanged.

Copilot uses AI. Check for mistakes.
@moonming
moonming merged commit bc06934 into mainApr 24, 2026
8 of 10 checks passed
moonming added a commit that referenced this pull request Apr 26, 2026
…t tests
CI on main has been red since #34 because aisix-core's Config crate
loader merges every AISIX_-prefixed env var into the root Config
struct (config-rs Environment::with_prefix("AISIX")), and Config has
#[serde(deny_unknown_fields)]. The redis integration test sets
AISIX_REDIS_URL on the rust-unit job, which leaks into every
Config::load_from_path call as `redis_url` and panics:
Config("deserialize: unknown field `redis_url`,
expected one of `etcd`, `proxy`, `admin`,
`observability`, `cache`, `managed`")
8 of 9 aisix-core::config::tests fail (the one that doesn't is
rejects_unknown_fields, which intentionally swallows the error).
Rename the env var so it doesn't sit under the AISIX_ prefix at
all. crates/aisix-cache/tests/redis_integration.rs reads
CACHE_TEST_REDIS_URL; CI sets the same. docs/testing.md +
crates/aisix-cache/src/redis.rs comment updated to match.
Verified: with the rename, all 9 config tests pass even with
CACHE_TEST_REDIS_URL set; reproducing with the old AISIX_REDIS_URL
still fails as expected (so the loader behaviour is unchanged for
real AISIX_-prefixed env overrides).
moonming added a commit that referenced this pull request Apr 26, 2026
…40)
* fix(aisix-etcd): make supervisor cache-write tests deterministic
Two supervisor tests waited on the spawned cache write via
`tokio::time::sleep(50ms)`. Under heavy CI load the spawn lost the
race against the disk read that followed, surfacing as:
resync_writes_to_disk_cache_then_restore_replays_it FAILED
put_and_delete_keep_cache_in_sync FAILED
Track the JoinHandle for each spawned write in a `pending_writes`
Mutex<Vec<_>> on the Supervisor, and expose a test-only async
`await_pending_cache_writes` that drains and awaits them. Both tests
now wait on real completion instead of a wall clock.
The new field is `#[cfg(test)]`-friendly via the awaiter — production
code never reads it. If a handle is dropped during shutdown the
underlying write either completed or was cancelled; the on-disk
cache is best-effort, and the next live cycle re-publishes from
etcd anyway.
* fix(ci): rename AISIX_REDIS_URL → CACHE_TEST_REDIS_URL to unblock unit tests
CI on main has been red since #34 because aisix-core's Config crate
loader merges every AISIX_-prefixed env var into the root Config
struct (config-rs Environment::with_prefix("AISIX")), and Config has
#[serde(deny_unknown_fields)]. The redis integration test sets
AISIX_REDIS_URL on the rust-unit job, which leaks into every
Config::load_from_path call as `redis_url` and panics:
Config("deserialize: unknown field `redis_url`,
expected one of `etcd`, `proxy`, `admin`,
`observability`, `cache`, `managed`")
8 of 9 aisix-core::config::tests fail (the one that doesn't is
rejects_unknown_fields, which intentionally swallows the error).
Rename the env var so it doesn't sit under the AISIX_ prefix at
all. crates/aisix-cache/tests/redis_integration.rs reads
CACHE_TEST_REDIS_URL; CI sets the same. docs/testing.md +
crates/aisix-cache/src/redis.rs comment updated to match.
Verified: with the rename, all 9 config tests pass even with
CACHE_TEST_REDIS_URL set; reproducing with the old AISIX_REDIS_URL
still fails as expected (so the loader behaviour is unchanged for
real AISIX_-prefixed env overrides).
* ci: tolerate artifact upload quota errors temporarily
Both `rust unit + coverage` and `build ui` jobs are currently failing
on the upload-artifact step with:
Failed to CreateArtifact: Artifact storage quota has been hit.
Unable to upload any new artifacts.
Usage is recalculated every 6-12 hours.
Tests + clippy pass; only the artifact upload is blocked. Add
`continue-on-error: true` to those two upload-artifact steps so the
test-passing signal isn't masked by the quota issue.
Downstream jobs that need ui-dist (build-bin → e2e) will fail at
download-artifact when the upload was skipped; e2e is already
`continue-on-error: true` at the job level, and coverage-gate is
advisory.
Revert this once the org-level storage usage refreshes (within
6-12h) or the quota is raised.
* ci: soft-fail build-aisix while artifact storage quota persists
build-aisix downloads ui-dist from build-ui. With build-ui's
upload-artifact set to continue-on-error during the storage quota
outage, the download fails and build-aisix errors. Since build-aisix
only feeds the advisory e2e job, mark it continue-on-error too so
the PR doesn't go red on a transitive dependency. Revert with the
other two when storage usage refreshes.
@jarvis9443
jarvis9443 deleted the fix/rustls-default-provider branch June 25, 2026 06:26
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

fix(server): install rustls default CryptoProvider at startup - #34

Merged
moonming merged 1 commit into
mainfrom
fix/rustls-default-provider
Apr 24, 2026
Merged

fix(server): install rustls default CryptoProvider at startup#34
moonming merged 1 commit into
mainfrom
fix/rustls-default-provider

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

aisix panics on first TLS handshake (etcd connect, reqwest /dp/register, etc.) with:

```
thread 'main' panicked at rustls/.../crypto/mod.rs:249:14:
Could not automatically determine the process-level CryptoProvider from Rustls crate features.
```

rustls 0.23 dropped implicit provider selection — when both `aws-lc-rs` and `ring` are reachable through transitive deps (reqwest `rustls-tls`, etcd-client, tokio-rustls…), the runtime can't pick. Fix: call `aws_lc_rs::default_provider().install_default()` at the very top of `main`, before anything else.

`let _ =` because install is idempotent and the return value only signals "another crate got there first" — which doesn't affect correctness.

Why aws-lc-rs and not ring

It's the upstream rustls default as of 0.23, FIPS-capable out of the box, and already in the transitive dep graph through reqwest + etcd-client. Picking it avoids adding another RSA/EC crypto implementation to the binary.

How we caught it

AISIX-Cloud e2e stack exposed this: the DP container is `docker run` via the test harness, hits this panic within 100ms of startup, exits, and `docker run --rm` removes the container before the harness can query `docker port`. Container logs now dumped on failure.

Test plan

  • `cargo build --bin aisix`
  • `cargo fmt --check`
  • `cargo clippy --all-targets --locked -- -D warnings`
  • CI green
  • Re-publish `ghcr.io/moonming/ai-gateway:aisix-e2e` from main so AISIX-Cloud e2e picks it up

aisix panics on first TLS handshake (etcd connect, reqwest register
call, etc.) with:
thread 'main' panicked at rustls/.../crypto/mod.rs:249:14:
Could not automatically determine the process-level CryptoProvider
from Rustls crate features.
rustls 0.23 dropped implicit provider selection — when both
`aws-lc-rs` and `ring` are reachable through transitive deps (we have
reqwest `rustls-tls`, etcd-client, tokio-rustls… all enabling one or
the other), the runtime can't pick. The panic fires the first time
any TLS operation touches the crypto layer.
Call `aws_lc_rs::default_provider().install_default()` at the very
top of main, before anything else loads. `let _ =` because install
is idempotent and we don't care if another crate beat us to it.
CopilotAI review requested due to automatic review settings April 24, 2026 02:47

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Ensures aisix doesn’t panic on first TLS usage under rustls 0.23 by explicitly installing a process-wide rustls CryptoProvider at startup.

Changes:

  • Install the aws-lc-rs rustls CryptoProvider at the start of main() to avoid runtime provider auto-detection panics when multiple crypto backends are present via transitive dependencies.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +56 to +57
// already depends on transitively. Falls back to ring only if
// the process somehow has a provider installed already (idempotent).

CopilotAIApr 24, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The comment says this "falls back to ring" if a provider is already installed, but the code doesn't install ring or perform any fallback logic; it simply keeps whatever default provider was installed first. Consider rewording to avoid implying ring is involved (e.g., "if another provider is already installed, keep it").

Suggested change
// already depends on transitively. Falls back to ring only if
// the process somehow has a provider installed already (idempotent).
// already depends on transitively. If another provider is already
// installed for the process, `install_default()` keeps it unchanged.

Copilot uses AI. Check for mistakes.
@moonming
moonming merged commit bc06934 into mainApr 24, 2026
8 of 10 checks passed
moonming added a commit that referenced this pull request Apr 26, 2026
…t tests
CI on main has been red since #34 because aisix-core's Config crate
loader merges every AISIX_-prefixed env var into the root Config
struct (config-rs Environment::with_prefix("AISIX")), and Config has
#[serde(deny_unknown_fields)]. The redis integration test sets
AISIX_REDIS_URL on the rust-unit job, which leaks into every
Config::load_from_path call as `redis_url` and panics:
Config("deserialize: unknown field `redis_url`,
expected one of `etcd`, `proxy`, `admin`,
`observability`, `cache`, `managed`")
8 of 9 aisix-core::config::tests fail (the one that doesn't is
rejects_unknown_fields, which intentionally swallows the error).
Rename the env var so it doesn't sit under the AISIX_ prefix at
all. crates/aisix-cache/tests/redis_integration.rs reads
CACHE_TEST_REDIS_URL; CI sets the same. docs/testing.md +
crates/aisix-cache/src/redis.rs comment updated to match.
Verified: with the rename, all 9 config tests pass even with
CACHE_TEST_REDIS_URL set; reproducing with the old AISIX_REDIS_URL
still fails as expected (so the loader behaviour is unchanged for
real AISIX_-prefixed env overrides).
moonming added a commit that referenced this pull request Apr 26, 2026
…40)
* fix(aisix-etcd): make supervisor cache-write tests deterministic
Two supervisor tests waited on the spawned cache write via
`tokio::time::sleep(50ms)`. Under heavy CI load the spawn lost the
race against the disk read that followed, surfacing as:
resync_writes_to_disk_cache_then_restore_replays_it FAILED
put_and_delete_keep_cache_in_sync FAILED
Track the JoinHandle for each spawned write in a `pending_writes`
Mutex<Vec<_>> on the Supervisor, and expose a test-only async
`await_pending_cache_writes` that drains and awaits them. Both tests
now wait on real completion instead of a wall clock.
The new field is `#[cfg(test)]`-friendly via the awaiter — production
code never reads it. If a handle is dropped during shutdown the
underlying write either completed or was cancelled; the on-disk
cache is best-effort, and the next live cycle re-publishes from
etcd anyway.
* fix(ci): rename AISIX_REDIS_URL → CACHE_TEST_REDIS_URL to unblock unit tests
CI on main has been red since #34 because aisix-core's Config crate
loader merges every AISIX_-prefixed env var into the root Config
struct (config-rs Environment::with_prefix("AISIX")), and Config has
#[serde(deny_unknown_fields)]. The redis integration test sets
AISIX_REDIS_URL on the rust-unit job, which leaks into every
Config::load_from_path call as `redis_url` and panics:
Config("deserialize: unknown field `redis_url`,
expected one of `etcd`, `proxy`, `admin`,
`observability`, `cache`, `managed`")
8 of 9 aisix-core::config::tests fail (the one that doesn't is
rejects_unknown_fields, which intentionally swallows the error).
Rename the env var so it doesn't sit under the AISIX_ prefix at
all. crates/aisix-cache/tests/redis_integration.rs reads
CACHE_TEST_REDIS_URL; CI sets the same. docs/testing.md +
crates/aisix-cache/src/redis.rs comment updated to match.
Verified: with the rename, all 9 config tests pass even with
CACHE_TEST_REDIS_URL set; reproducing with the old AISIX_REDIS_URL
still fails as expected (so the loader behaviour is unchanged for
real AISIX_-prefixed env overrides).
* ci: tolerate artifact upload quota errors temporarily
Both `rust unit + coverage` and `build ui` jobs are currently failing
on the upload-artifact step with:
Failed to CreateArtifact: Artifact storage quota has been hit.
Unable to upload any new artifacts.
Usage is recalculated every 6-12 hours.
Tests + clippy pass; only the artifact upload is blocked. Add
`continue-on-error: true` to those two upload-artifact steps so the
test-passing signal isn't masked by the quota issue.
Downstream jobs that need ui-dist (build-bin → e2e) will fail at
download-artifact when the upload was skipped; e2e is already
`continue-on-error: true` at the job level, and coverage-gate is
advisory.
Revert this once the org-level storage usage refreshes (within
6-12h) or the quota is raised.
* ci: soft-fail build-aisix while artifact storage quota persists
build-aisix downloads ui-dist from build-ui. With build-ui's
upload-artifact set to continue-on-error during the storage quota
outage, the download fails and build-aisix errors. Since build-aisix
only feeds the advisory e2e job, mark it continue-on-error too so
the PR doesn't go red on a transitive dependency. Revert with the
other two when storage usage refreshes.
@jarvis9443
jarvis9443 deleted the fix/rustls-default-provider branch June 25, 2026 06:26
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): install rustls default CryptoProvider at startup - #34

Merged
moonming merged 1 commit into
mainfrom
fix/rustls-default-provider
Apr 24, 2026
Merged

fix(server): install rustls default CryptoProvider at startup#34
moonming merged 1 commit into
mainfrom
fix/rustls-default-provider

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

aisix panics on first TLS handshake (etcd connect, reqwest /dp/register, etc.) with:

```
thread 'main' panicked at rustls/.../crypto/mod.rs:249:14:
Could not automatically determine the process-level CryptoProvider from Rustls crate features.
```

rustls 0.23 dropped implicit provider selection — when both `aws-lc-rs` and `ring` are reachable through transitive deps (reqwest `rustls-tls`, etcd-client, tokio-rustls…), the runtime can't pick. Fix: call `aws_lc_rs::default_provider().install_default()` at the very top of `main`, before anything else.

`let _ =` because install is idempotent and the return value only signals "another crate got there first" — which doesn't affect correctness.

Why aws-lc-rs and not ring

It's the upstream rustls default as of 0.23, FIPS-capable out of the box, and already in the transitive dep graph through reqwest + etcd-client. Picking it avoids adding another RSA/EC crypto implementation to the binary.

How we caught it

AISIX-Cloud e2e stack exposed this: the DP container is `docker run` via the test harness, hits this panic within 100ms of startup, exits, and `docker run --rm` removes the container before the harness can query `docker port`. Container logs now dumped on failure.

Test plan

  • `cargo build --bin aisix`
  • `cargo fmt --check`
  • `cargo clippy --all-targets --locked -- -D warnings`
  • CI green
  • Re-publish `ghcr.io/moonming/ai-gateway:aisix-e2e` from main so AISIX-Cloud e2e picks it up

aisix panics on first TLS handshake (etcd connect, reqwest register
call, etc.) with:
thread 'main' panicked at rustls/.../crypto/mod.rs:249:14:
Could not automatically determine the process-level CryptoProvider
from Rustls crate features.
rustls 0.23 dropped implicit provider selection — when both
`aws-lc-rs` and `ring` are reachable through transitive deps (we have
reqwest `rustls-tls`, etcd-client, tokio-rustls… all enabling one or
the other), the runtime can't pick. The panic fires the first time
any TLS operation touches the crypto layer.
Call `aws_lc_rs::default_provider().install_default()` at the very
top of main, before anything else loads. `let _ =` because install
is idempotent and we don't care if another crate beat us to it.
CopilotAI review requested due to automatic review settings April 24, 2026 02:47

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Ensures aisix doesn’t panic on first TLS usage under rustls 0.23 by explicitly installing a process-wide rustls CryptoProvider at startup.

Changes:

  • Install the aws-lc-rs rustls CryptoProvider at the start of main() to avoid runtime provider auto-detection panics when multiple crypto backends are present via transitive dependencies.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +56 to +57
// already depends on transitively. Falls back to ring only if
// the process somehow has a provider installed already (idempotent).

CopilotAIApr 24, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The comment says this "falls back to ring" if a provider is already installed, but the code doesn't install ring or perform any fallback logic; it simply keeps whatever default provider was installed first. Consider rewording to avoid implying ring is involved (e.g., "if another provider is already installed, keep it").

Suggested change
// already depends on transitively. Falls back to ring only if
// the process somehow has a provider installed already (idempotent).
// already depends on transitively. If another provider is already
// installed for the process, `install_default()` keeps it unchanged.

Copilot uses AI. Check for mistakes.
@moonming
moonming merged commit bc06934 into mainApr 24, 2026
8 of 10 checks passed
moonming added a commit that referenced this pull request Apr 26, 2026
…t tests
CI on main has been red since #34 because aisix-core's Config crate
loader merges every AISIX_-prefixed env var into the root Config
struct (config-rs Environment::with_prefix("AISIX")), and Config has
#[serde(deny_unknown_fields)]. The redis integration test sets
AISIX_REDIS_URL on the rust-unit job, which leaks into every
Config::load_from_path call as `redis_url` and panics:
Config("deserialize: unknown field `redis_url`,
expected one of `etcd`, `proxy`, `admin`,
`observability`, `cache`, `managed`")
8 of 9 aisix-core::config::tests fail (the one that doesn't is
rejects_unknown_fields, which intentionally swallows the error).
Rename the env var so it doesn't sit under the AISIX_ prefix at
all. crates/aisix-cache/tests/redis_integration.rs reads
CACHE_TEST_REDIS_URL; CI sets the same. docs/testing.md +
crates/aisix-cache/src/redis.rs comment updated to match.
Verified: with the rename, all 9 config tests pass even with
CACHE_TEST_REDIS_URL set; reproducing with the old AISIX_REDIS_URL
still fails as expected (so the loader behaviour is unchanged for
real AISIX_-prefixed env overrides).
moonming added a commit that referenced this pull request Apr 26, 2026
…40)
* fix(aisix-etcd): make supervisor cache-write tests deterministic
Two supervisor tests waited on the spawned cache write via
`tokio::time::sleep(50ms)`. Under heavy CI load the spawn lost the
race against the disk read that followed, surfacing as:
resync_writes_to_disk_cache_then_restore_replays_it FAILED
put_and_delete_keep_cache_in_sync FAILED
Track the JoinHandle for each spawned write in a `pending_writes`
Mutex<Vec<_>> on the Supervisor, and expose a test-only async
`await_pending_cache_writes` that drains and awaits them. Both tests
now wait on real completion instead of a wall clock.
The new field is `#[cfg(test)]`-friendly via the awaiter — production
code never reads it. If a handle is dropped during shutdown the
underlying write either completed or was cancelled; the on-disk
cache is best-effort, and the next live cycle re-publishes from
etcd anyway.
* fix(ci): rename AISIX_REDIS_URL → CACHE_TEST_REDIS_URL to unblock unit tests
CI on main has been red since #34 because aisix-core's Config crate
loader merges every AISIX_-prefixed env var into the root Config
struct (config-rs Environment::with_prefix("AISIX")), and Config has
#[serde(deny_unknown_fields)]. The redis integration test sets
AISIX_REDIS_URL on the rust-unit job, which leaks into every
Config::load_from_path call as `redis_url` and panics:
Config("deserialize: unknown field `redis_url`,
expected one of `etcd`, `proxy`, `admin`,
`observability`, `cache`, `managed`")
8 of 9 aisix-core::config::tests fail (the one that doesn't is
rejects_unknown_fields, which intentionally swallows the error).
Rename the env var so it doesn't sit under the AISIX_ prefix at
all. crates/aisix-cache/tests/redis_integration.rs reads
CACHE_TEST_REDIS_URL; CI sets the same. docs/testing.md +
crates/aisix-cache/src/redis.rs comment updated to match.
Verified: with the rename, all 9 config tests pass even with
CACHE_TEST_REDIS_URL set; reproducing with the old AISIX_REDIS_URL
still fails as expected (so the loader behaviour is unchanged for
real AISIX_-prefixed env overrides).
* ci: tolerate artifact upload quota errors temporarily
Both `rust unit + coverage` and `build ui` jobs are currently failing
on the upload-artifact step with:
Failed to CreateArtifact: Artifact storage quota has been hit.
Unable to upload any new artifacts.
Usage is recalculated every 6-12 hours.
Tests + clippy pass; only the artifact upload is blocked. Add
`continue-on-error: true` to those two upload-artifact steps so the
test-passing signal isn't masked by the quota issue.
Downstream jobs that need ui-dist (build-bin → e2e) will fail at
download-artifact when the upload was skipped; e2e is already
`continue-on-error: true` at the job level, and coverage-gate is
advisory.
Revert this once the org-level storage usage refreshes (within
6-12h) or the quota is raised.
* ci: soft-fail build-aisix while artifact storage quota persists
build-aisix downloads ui-dist from build-ui. With build-ui's
upload-artifact set to continue-on-error during the storage quota
outage, the download fails and build-aisix errors. Since build-aisix
only feeds the advisory e2e job, mark it continue-on-error too so
the PR doesn't go red on a transitive dependency. Revert with the
other two when storage usage refreshes.
@jarvis9443
jarvis9443 deleted the fix/rustls-default-provider branch June 25, 2026 06:26
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

fix(server): install rustls default CryptoProvider at startup - #34

Merged
moonming merged 1 commit into
mainfrom
fix/rustls-default-provider
Apr 24, 2026
Merged

fix(server): install rustls default CryptoProvider at startup#34
moonming merged 1 commit into
mainfrom
fix/rustls-default-provider

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

aisix panics on first TLS handshake (etcd connect, reqwest /dp/register, etc.) with:

```
thread 'main' panicked at rustls/.../crypto/mod.rs:249:14:
Could not automatically determine the process-level CryptoProvider from Rustls crate features.
```

rustls 0.23 dropped implicit provider selection — when both `aws-lc-rs` and `ring` are reachable through transitive deps (reqwest `rustls-tls`, etcd-client, tokio-rustls…), the runtime can't pick. Fix: call `aws_lc_rs::default_provider().install_default()` at the very top of `main`, before anything else.

`let _ =` because install is idempotent and the return value only signals "another crate got there first" — which doesn't affect correctness.

Why aws-lc-rs and not ring

It's the upstream rustls default as of 0.23, FIPS-capable out of the box, and already in the transitive dep graph through reqwest + etcd-client. Picking it avoids adding another RSA/EC crypto implementation to the binary.

How we caught it

AISIX-Cloud e2e stack exposed this: the DP container is `docker run` via the test harness, hits this panic within 100ms of startup, exits, and `docker run --rm` removes the container before the harness can query `docker port`. Container logs now dumped on failure.

Test plan

  • `cargo build --bin aisix`
  • `cargo fmt --check`
  • `cargo clippy --all-targets --locked -- -D warnings`
  • CI green
  • Re-publish `ghcr.io/moonming/ai-gateway:aisix-e2e` from main so AISIX-Cloud e2e picks it up

aisix panics on first TLS handshake (etcd connect, reqwest register
call, etc.) with:
thread 'main' panicked at rustls/.../crypto/mod.rs:249:14:
Could not automatically determine the process-level CryptoProvider
from Rustls crate features.
rustls 0.23 dropped implicit provider selection — when both
`aws-lc-rs` and `ring` are reachable through transitive deps (we have
reqwest `rustls-tls`, etcd-client, tokio-rustls… all enabling one or
the other), the runtime can't pick. The panic fires the first time
any TLS operation touches the crypto layer.
Call `aws_lc_rs::default_provider().install_default()` at the very
top of main, before anything else loads. `let _ =` because install
is idempotent and we don't care if another crate beat us to it.
CopilotAI review requested due to automatic review settings April 24, 2026 02:47

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Ensures aisix doesn’t panic on first TLS usage under rustls 0.23 by explicitly installing a process-wide rustls CryptoProvider at startup.

Changes:

  • Install the aws-lc-rs rustls CryptoProvider at the start of main() to avoid runtime provider auto-detection panics when multiple crypto backends are present via transitive dependencies.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +56 to +57
// already depends on transitively. Falls back to ring only if
// the process somehow has a provider installed already (idempotent).

CopilotAIApr 24, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The comment says this "falls back to ring" if a provider is already installed, but the code doesn't install ring or perform any fallback logic; it simply keeps whatever default provider was installed first. Consider rewording to avoid implying ring is involved (e.g., "if another provider is already installed, keep it").

Suggested change
// already depends on transitively. Falls back to ring only if
// the process somehow has a provider installed already (idempotent).
// already depends on transitively. If another provider is already
// installed for the process, `install_default()` keeps it unchanged.

Copilot uses AI. Check for mistakes.
@moonming
moonming merged commit bc06934 into mainApr 24, 2026
8 of 10 checks passed
moonming added a commit that referenced this pull request Apr 26, 2026
…t tests
CI on main has been red since #34 because aisix-core's Config crate
loader merges every AISIX_-prefixed env var into the root Config
struct (config-rs Environment::with_prefix("AISIX")), and Config has
#[serde(deny_unknown_fields)]. The redis integration test sets
AISIX_REDIS_URL on the rust-unit job, which leaks into every
Config::load_from_path call as `redis_url` and panics:
Config("deserialize: unknown field `redis_url`,
expected one of `etcd`, `proxy`, `admin`,
`observability`, `cache`, `managed`")
8 of 9 aisix-core::config::tests fail (the one that doesn't is
rejects_unknown_fields, which intentionally swallows the error).
Rename the env var so it doesn't sit under the AISIX_ prefix at
all. crates/aisix-cache/tests/redis_integration.rs reads
CACHE_TEST_REDIS_URL; CI sets the same. docs/testing.md +
crates/aisix-cache/src/redis.rs comment updated to match.
Verified: with the rename, all 9 config tests pass even with
CACHE_TEST_REDIS_URL set; reproducing with the old AISIX_REDIS_URL
still fails as expected (so the loader behaviour is unchanged for
real AISIX_-prefixed env overrides).
moonming added a commit that referenced this pull request Apr 26, 2026
…40)
* fix(aisix-etcd): make supervisor cache-write tests deterministic
Two supervisor tests waited on the spawned cache write via
`tokio::time::sleep(50ms)`. Under heavy CI load the spawn lost the
race against the disk read that followed, surfacing as:
resync_writes_to_disk_cache_then_restore_replays_it FAILED
put_and_delete_keep_cache_in_sync FAILED
Track the JoinHandle for each spawned write in a `pending_writes`
Mutex<Vec<_>> on the Supervisor, and expose a test-only async
`await_pending_cache_writes` that drains and awaits them. Both tests
now wait on real completion instead of a wall clock.
The new field is `#[cfg(test)]`-friendly via the awaiter — production
code never reads it. If a handle is dropped during shutdown the
underlying write either completed or was cancelled; the on-disk
cache is best-effort, and the next live cycle re-publishes from
etcd anyway.
* fix(ci): rename AISIX_REDIS_URL → CACHE_TEST_REDIS_URL to unblock unit tests
CI on main has been red since #34 because aisix-core's Config crate
loader merges every AISIX_-prefixed env var into the root Config
struct (config-rs Environment::with_prefix("AISIX")), and Config has
#[serde(deny_unknown_fields)]. The redis integration test sets
AISIX_REDIS_URL on the rust-unit job, which leaks into every
Config::load_from_path call as `redis_url` and panics:
Config("deserialize: unknown field `redis_url`,
expected one of `etcd`, `proxy`, `admin`,
`observability`, `cache`, `managed`")
8 of 9 aisix-core::config::tests fail (the one that doesn't is
rejects_unknown_fields, which intentionally swallows the error).
Rename the env var so it doesn't sit under the AISIX_ prefix at
all. crates/aisix-cache/tests/redis_integration.rs reads
CACHE_TEST_REDIS_URL; CI sets the same. docs/testing.md +
crates/aisix-cache/src/redis.rs comment updated to match.
Verified: with the rename, all 9 config tests pass even with
CACHE_TEST_REDIS_URL set; reproducing with the old AISIX_REDIS_URL
still fails as expected (so the loader behaviour is unchanged for
real AISIX_-prefixed env overrides).
* ci: tolerate artifact upload quota errors temporarily
Both `rust unit + coverage` and `build ui` jobs are currently failing
on the upload-artifact step with:
Failed to CreateArtifact: Artifact storage quota has been hit.
Unable to upload any new artifacts.
Usage is recalculated every 6-12 hours.
Tests + clippy pass; only the artifact upload is blocked. Add
`continue-on-error: true` to those two upload-artifact steps so the
test-passing signal isn't masked by the quota issue.
Downstream jobs that need ui-dist (build-bin → e2e) will fail at
download-artifact when the upload was skipped; e2e is already
`continue-on-error: true` at the job level, and coverage-gate is
advisory.
Revert this once the org-level storage usage refreshes (within
6-12h) or the quota is raised.
* ci: soft-fail build-aisix while artifact storage quota persists
build-aisix downloads ui-dist from build-ui. With build-ui's
upload-artifact set to continue-on-error during the storage quota
outage, the download fails and build-aisix errors. Since build-aisix
only feeds the advisory e2e job, mark it continue-on-error too so
the PR doesn't go red on a transitive dependency. Revert with the
other two when storage usage refreshes.
@jarvis9443
jarvis9443 deleted the fix/rustls-default-provider branch June 25, 2026 06:26
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

fix(server): install rustls default CryptoProvider at startup - #34

Merged
moonming merged 1 commit into
mainfrom
fix/rustls-default-provider
Apr 24, 2026
Merged

fix(server): install rustls default CryptoProvider at startup#34
moonming merged 1 commit into
mainfrom
fix/rustls-default-provider

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

aisix panics on first TLS handshake (etcd connect, reqwest /dp/register, etc.) with:

```
thread 'main' panicked at rustls/.../crypto/mod.rs:249:14:
Could not automatically determine the process-level CryptoProvider from Rustls crate features.
```

rustls 0.23 dropped implicit provider selection — when both `aws-lc-rs` and `ring` are reachable through transitive deps (reqwest `rustls-tls`, etcd-client, tokio-rustls…), the runtime can't pick. Fix: call `aws_lc_rs::default_provider().install_default()` at the very top of `main`, before anything else.

`let _ =` because install is idempotent and the return value only signals "another crate got there first" — which doesn't affect correctness.

Why aws-lc-rs and not ring

It's the upstream rustls default as of 0.23, FIPS-capable out of the box, and already in the transitive dep graph through reqwest + etcd-client. Picking it avoids adding another RSA/EC crypto implementation to the binary.

How we caught it

AISIX-Cloud e2e stack exposed this: the DP container is `docker run` via the test harness, hits this panic within 100ms of startup, exits, and `docker run --rm` removes the container before the harness can query `docker port`. Container logs now dumped on failure.

Test plan

  • `cargo build --bin aisix`
  • `cargo fmt --check`
  • `cargo clippy --all-targets --locked -- -D warnings`
  • CI green
  • Re-publish `ghcr.io/moonming/ai-gateway:aisix-e2e` from main so AISIX-Cloud e2e picks it up

aisix panics on first TLS handshake (etcd connect, reqwest register
call, etc.) with:
thread 'main' panicked at rustls/.../crypto/mod.rs:249:14:
Could not automatically determine the process-level CryptoProvider
from Rustls crate features.
rustls 0.23 dropped implicit provider selection — when both
`aws-lc-rs` and `ring` are reachable through transitive deps (we have
reqwest `rustls-tls`, etcd-client, tokio-rustls… all enabling one or
the other), the runtime can't pick. The panic fires the first time
any TLS operation touches the crypto layer.
Call `aws_lc_rs::default_provider().install_default()` at the very
top of main, before anything else loads. `let _ =` because install
is idempotent and we don't care if another crate beat us to it.
CopilotAI review requested due to automatic review settings April 24, 2026 02:47

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Ensures aisix doesn’t panic on first TLS usage under rustls 0.23 by explicitly installing a process-wide rustls CryptoProvider at startup.

Changes:

  • Install the aws-lc-rs rustls CryptoProvider at the start of main() to avoid runtime provider auto-detection panics when multiple crypto backends are present via transitive dependencies.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +56 to +57
// already depends on transitively. Falls back to ring only if
// the process somehow has a provider installed already (idempotent).

CopilotAIApr 24, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The comment says this "falls back to ring" if a provider is already installed, but the code doesn't install ring or perform any fallback logic; it simply keeps whatever default provider was installed first. Consider rewording to avoid implying ring is involved (e.g., "if another provider is already installed, keep it").

Suggested change
// already depends on transitively. Falls back to ring only if
// the process somehow has a provider installed already (idempotent).
// already depends on transitively. If another provider is already
// installed for the process, `install_default()` keeps it unchanged.

Copilot uses AI. Check for mistakes.
@moonming
moonming merged commit bc06934 into mainApr 24, 2026
8 of 10 checks passed
moonming added a commit that referenced this pull request Apr 26, 2026
…t tests
CI on main has been red since #34 because aisix-core's Config crate
loader merges every AISIX_-prefixed env var into the root Config
struct (config-rs Environment::with_prefix("AISIX")), and Config has
#[serde(deny_unknown_fields)]. The redis integration test sets
AISIX_REDIS_URL on the rust-unit job, which leaks into every
Config::load_from_path call as `redis_url` and panics:
Config("deserialize: unknown field `redis_url`,
expected one of `etcd`, `proxy`, `admin`,
`observability`, `cache`, `managed`")
8 of 9 aisix-core::config::tests fail (the one that doesn't is
rejects_unknown_fields, which intentionally swallows the error).
Rename the env var so it doesn't sit under the AISIX_ prefix at
all. crates/aisix-cache/tests/redis_integration.rs reads
CACHE_TEST_REDIS_URL; CI sets the same. docs/testing.md +
crates/aisix-cache/src/redis.rs comment updated to match.
Verified: with the rename, all 9 config tests pass even with
CACHE_TEST_REDIS_URL set; reproducing with the old AISIX_REDIS_URL
still fails as expected (so the loader behaviour is unchanged for
real AISIX_-prefixed env overrides).
moonming added a commit that referenced this pull request Apr 26, 2026
…40)
* fix(aisix-etcd): make supervisor cache-write tests deterministic
Two supervisor tests waited on the spawned cache write via
`tokio::time::sleep(50ms)`. Under heavy CI load the spawn lost the
race against the disk read that followed, surfacing as:
resync_writes_to_disk_cache_then_restore_replays_it FAILED
put_and_delete_keep_cache_in_sync FAILED
Track the JoinHandle for each spawned write in a `pending_writes`
Mutex<Vec<_>> on the Supervisor, and expose a test-only async
`await_pending_cache_writes` that drains and awaits them. Both tests
now wait on real completion instead of a wall clock.
The new field is `#[cfg(test)]`-friendly via the awaiter — production
code never reads it. If a handle is dropped during shutdown the
underlying write either completed or was cancelled; the on-disk
cache is best-effort, and the next live cycle re-publishes from
etcd anyway.
* fix(ci): rename AISIX_REDIS_URL → CACHE_TEST_REDIS_URL to unblock unit tests
CI on main has been red since #34 because aisix-core's Config crate
loader merges every AISIX_-prefixed env var into the root Config
struct (config-rs Environment::with_prefix("AISIX")), and Config has
#[serde(deny_unknown_fields)]. The redis integration test sets
AISIX_REDIS_URL on the rust-unit job, which leaks into every
Config::load_from_path call as `redis_url` and panics:
Config("deserialize: unknown field `redis_url`,
expected one of `etcd`, `proxy`, `admin`,
`observability`, `cache`, `managed`")
8 of 9 aisix-core::config::tests fail (the one that doesn't is
rejects_unknown_fields, which intentionally swallows the error).
Rename the env var so it doesn't sit under the AISIX_ prefix at
all. crates/aisix-cache/tests/redis_integration.rs reads
CACHE_TEST_REDIS_URL; CI sets the same. docs/testing.md +
crates/aisix-cache/src/redis.rs comment updated to match.
Verified: with the rename, all 9 config tests pass even with
CACHE_TEST_REDIS_URL set; reproducing with the old AISIX_REDIS_URL
still fails as expected (so the loader behaviour is unchanged for
real AISIX_-prefixed env overrides).
* ci: tolerate artifact upload quota errors temporarily
Both `rust unit + coverage` and `build ui` jobs are currently failing
on the upload-artifact step with:
Failed to CreateArtifact: Artifact storage quota has been hit.
Unable to upload any new artifacts.
Usage is recalculated every 6-12 hours.
Tests + clippy pass; only the artifact upload is blocked. Add
`continue-on-error: true` to those two upload-artifact steps so the
test-passing signal isn't masked by the quota issue.
Downstream jobs that need ui-dist (build-bin → e2e) will fail at
download-artifact when the upload was skipped; e2e is already
`continue-on-error: true` at the job level, and coverage-gate is
advisory.
Revert this once the org-level storage usage refreshes (within
6-12h) or the quota is raised.
* ci: soft-fail build-aisix while artifact storage quota persists
build-aisix downloads ui-dist from build-ui. With build-ui's
upload-artifact set to continue-on-error during the storage quota
outage, the download fails and build-aisix errors. Since build-aisix
only feeds the advisory e2e job, mark it continue-on-error too so
the PR doesn't go red on a transitive dependency. Revert with the
other two when storage usage refreshes.
@jarvis9443
jarvis9443 deleted the fix/rustls-default-provider branch June 25, 2026 06:26
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@moonming