Repository files navigation

Activescale

Activescale enables fast Kubernetes autoscaling with push-based, per-pod active request and connection metrics from Envoy.

Active means currently in flight: requests not yet completed and connections still open. By Little's Law (L = λW), active requests reflect both request rate and latency.

Unlike CPU and memory metrics collected through metrics-server, Envoy pushes active signals directly to Activescale, avoiding extra scrape and polling delays.

Documentation

  • Guide: 한국어 | English — metric selection, HPA configuration, and verification
  • Motivation: 한국어 | English — rationale, latency comparison, and architecture
  • Test results: 한국어 | English — scaling responsiveness, Scale-out Coverage, and estimated resource utilization summary
  • Kubernetes manifest example: k8s-manifest-example.yaml — deployment and HPA example requiring Redis, ServiceAccount/RBAC, serving certificates, and a target workload

Features

  • Active request concurrency as an autoscaling signal that reflects both request rate and request latency
  • Push-based Envoy StreamMetrics ingestion without intermediary Prometheus scrape and KEDA polling delays
  • Pod-scoped Custom Metrics API values that HPA averages across workload replicas
  • Direction-aware inbound_active_requests and outbound_active_requests, plus aggregate active_connections
  • Stateless high availability with shared Redis/Valkey state and TTL-based stale metric expiry
  • Redis standalone and Cluster modes with optional TLS
  • Configurable gRPC receive limit and Klog-based summary and debug logging

Metrics

  • inbound_active_requests: inbound in-flight HTTP requests from http.inbound_*.downstream_rq_active
  • outbound_active_requests: outbound in-flight HTTP requests from http.outbound_*.downstream_rq_active
  • active_connections: open connections across aggregate Envoy service listeners from listener.<address>.downstream_cx_active

All metrics are exposed as pod-scoped custom metrics for HPA.

Choosing a request metric

Choose the metric based on the role of the workload being scaled.

WorkloadMetricWhen to use
Application or API serviceinbound_active_requestsUse for a workload that receives and processes client requests. This is the default choice for typical services.
Dedicated gateway or proxyoutbound_active_requestsUse when the workload primarily forwards client requests to upstream services and inbound_active_requests is unavailable or does not reflect the traffic.

Do not configure both metrics as a fallback. For a normal application, outbound requests usually represent calls to dependencies and can trigger scaling for unrelated downstream traffic.

If the workload role is unclear, generate a small amount of client traffic and query both metrics. Select the one that returns per-pod values and increases with the client request concurrency being tested.

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/inbound_active_requests?labelSelector=<selector>'
kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/outbound_active_requests?labelSelector=<selector>'

Architecture

graph LR
PodA["📦 service pod A<br/>Istio/Envoy"]
PodB["📦 service pod B<br/>Istio/Envoy"]
PodC["📦 service pod C<br/>Istio/Envoy"]
Agg(("⚙️ <b>activescale<b/><br/>(Stateless, HA)"))
Redis[("🗄️ shared memory<br/>(e.g., Redis)")]
KEDA{{"🚀 HPA<br/>(average computed by HPA)"}}
PodA -."1. Push<br/>(5s delay)".-> Agg
PodB -."1. Push<br/>(5s delay)".-> Agg
PodC -."1. Push<br/>(5s delay)".-> Agg
Agg ==>|2. update + TTL| Redis
KEDA --"3. Query"--> Agg
Redis -."4. metric value (per pod)".-> Agg
Agg --"5. metric value (per pod)"--> KEDA
style Redis fill:#E3ECF8,stroke:#6E8FB3,stroke-width:2px,color:#000
style KEDA fill:#DDEFD8,stroke:#7DA67D,stroke-width:2px,color:#000
style Agg fill:#E3ECFF,stroke:#6E8FB3,color:#000
style PodA fill:#FAFAFA,stroke:#999
style PodB fill:#FAFAFA,stroke:#999
style PodC fill:#FAFAFA,stroke:#999
Loading

Envoy Metrics Message Shape

Activescale receives Envoy gRPC StreamMetrics messages. Each message contains a node identity and a list of metric families. A simplified shape:

{
"identifier": {
"node": {
"id": "sidecar~10.0.0.1~my-pod.my-ns~my-ns.svc.cluster.local",
"metadata": {
"NAME": "my-pod",
"NAMESPACE": "my-ns"
}
}
},
"envoy_metrics": [
{
"name": "http.inbound_0.0.0.0_8080;.downstream_rq_active",
"metric": [
{ "gauge": { "value": 3 } }
]
},
{
"name": "http.outbound_0.0.0.0_8080;.downstream_rq_active",
"metric": [
{ "gauge": { "value": 4 } }
]
},
{
"name": "listener.0.0.0.0_8080.downstream_cx_active",
"metric": [
{ "gauge": { "value": 12 } }
]
},
{
"name": "cluster.xds-grpc;.circuit_breakers.default.cx_pool_open",
"metric": [
{ "gauge": { "value": 1 } }
]
}
]
}

Pod identity extraction prefers node.metadata.NAME and node.metadata.NAMESPACE. If either metadata field is missing, activescale falls back to parsing the Istio-style node.id (sidecar~<ip>~<pod>.<namespace>~<namespace>.svc.cluster.local).

Summary Counters Meaning

  • messages: number of StreamMetrics messages received (one Recv() call).
  • stored_metrics: number of metric writes stored in Redis for accepted request and connection samples.
  • dropped_by_ids: messages dropped because pod identity could not be extracted.
  • dropped_by_names: metric families skipped because their name did not match the inbound, outbound, or aggregate service listener rules.

stored does not necessarily equal messages because a message can contain multiple metric families or multiple samples, and dropped_by_names is counted per metric family, not per message.

Debugging

Check Custom Metrics API is registered:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2'

Query inbound active requests with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/inbound_active_requests?labelSelector=app=<app>'

Query outbound active requests with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/outbound_active_requests?labelSelector=app=<app>'

Query active connections with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/active_connections?labelSelector=app=<app>'

Check activescale ingest logs:

kubectl logs -n ns-observability deploy/activescale | rg -n "stored (inbound_active_requests|outbound_active_requests|active_connections)|skipping metric name|missing pod identity"

Enable debug logs and restart:

kubectl -n ns-observability set env deploy/activescale LOG_VERBOSITY=4
kubectl -n ns-observability rollout restart deploy/activescale

Confirm Envoy bootstrap includes the metrics service:

istioctl proxy-config bootstrap <pod> -n <ns>| rg -n "envoyMetricsService|metrics_service|envoy_grpc|cluster_name|activescale|9000"

A minimal Kubernetes reference example lives in k8s-manifest-example.yaml.

Notes

Activescale reads both Envoy HTTP connection manager and listener stats, depending on the metric.

  • Activescale exposes inbound and outbound HTTP connection-manager request metrics separately.
  • http.admin.*, http.agent.*, and Prometheus-style envoy_http_downstream_rq_active aliases are ignored.
  • active_connections accepts aggregate listener.<address>.downstream_cx_active families for service listener ports.
  • Worker breakdowns such as listener.<address>.worker_0.downstream_cx_active and Istio-reserved proxy ports for admin, failure detection, debug, telemetry, health, and DNS are ignored.
  • Traffic-path listeners for outbound (15001), inbound (15006), and HBONE (15008) remain eligible.

Appendix

Missing metric values

Envoy can omit an active stat family for a new or idle pod. Activescale refreshes a per-pod heartbeat on every StreamMetrics message and applies the metric TTL to that heartbeat. The default TTL is 20s and can be changed with METRIC_TTL or --ttl.

  • A fresh heartbeat with no metric key is returned as 0 because collection is healthy and the pod has no observed active work.
  • A missing or expired heartbeat remains missing and can produce HTTP 404 NotFound because returning 0 would hide an Envoy collection outage and could cause an unsafe HPA scale-in.

Heartbeat-gated zeros prevent healthy idle pods from causing FailedGetPodsMetric or conservative scale-in delays without treating stale telemetry as zero. Provider summary logs count these values as synthesized_zeros.

Handling HTTP 404 NotFound

HTTP 404 NotFound means telemetry is unavailable, not that the metric value is zero.

Kubernetes HPA retries automatically and handles missing metrics conservatively, as described by the HPA algorithm.

Metric stateHPA action
All configured metrics return HTTP 404 NotFoundKeep the current replicas; neither scale up nor scale down
One metric returns HTTP 404 NotFound, another valid metric requests scale-upScale up using the valid metric
One metric returns HTTP 404 NotFound, another valid metric requests scale-downSkip scale-down and keep the current replicas
All configured metrics are availableUse the largest desired replica count

A direct Custom Metrics API client should apply the same policy: allow scale-up from another valid metric, block scale-down while any metric is unavailable, keep the current replicas when all metrics are unavailable, and retry on the next collection cycle. It must not substitute zero and should alert if HTTP 404 NotFound persists.

About

Push-based Kubernetes HPA adapter using active request/connection metrics to detect load spikes far earlier than CPU/Memory signals while improving resource utilization.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Activescale

Activescale enables fast Kubernetes autoscaling with push-based, per-pod active request and connection metrics from Envoy.

Active means currently in flight: requests not yet completed and connections still open. By Little's Law (L = λW), active requests reflect both request rate and latency.

Unlike CPU and memory metrics collected through metrics-server, Envoy pushes active signals directly to Activescale, avoiding extra scrape and polling delays.

Documentation

  • Guide: 한국어 | English — metric selection, HPA configuration, and verification
  • Motivation: 한국어 | English — rationale, latency comparison, and architecture
  • Test results: 한국어 | English — scaling responsiveness, Scale-out Coverage, and estimated resource utilization summary
  • Kubernetes manifest example: k8s-manifest-example.yaml — deployment and HPA example requiring Redis, ServiceAccount/RBAC, serving certificates, and a target workload

Features

  • Active request concurrency as an autoscaling signal that reflects both request rate and request latency
  • Push-based Envoy StreamMetrics ingestion without intermediary Prometheus scrape and KEDA polling delays
  • Pod-scoped Custom Metrics API values that HPA averages across workload replicas
  • Direction-aware inbound_active_requests and outbound_active_requests, plus aggregate active_connections
  • Stateless high availability with shared Redis/Valkey state and TTL-based stale metric expiry
  • Redis standalone and Cluster modes with optional TLS
  • Configurable gRPC receive limit and Klog-based summary and debug logging

Metrics

  • inbound_active_requests: inbound in-flight HTTP requests from http.inbound_*.downstream_rq_active
  • outbound_active_requests: outbound in-flight HTTP requests from http.outbound_*.downstream_rq_active
  • active_connections: open connections across aggregate Envoy service listeners from listener.<address>.downstream_cx_active

All metrics are exposed as pod-scoped custom metrics for HPA.

Choosing a request metric

Choose the metric based on the role of the workload being scaled.

WorkloadMetricWhen to use
Application or API serviceinbound_active_requestsUse for a workload that receives and processes client requests. This is the default choice for typical services.
Dedicated gateway or proxyoutbound_active_requestsUse when the workload primarily forwards client requests to upstream services and inbound_active_requests is unavailable or does not reflect the traffic.

Do not configure both metrics as a fallback. For a normal application, outbound requests usually represent calls to dependencies and can trigger scaling for unrelated downstream traffic.

If the workload role is unclear, generate a small amount of client traffic and query both metrics. Select the one that returns per-pod values and increases with the client request concurrency being tested.

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/inbound_active_requests?labelSelector=<selector>'
kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/outbound_active_requests?labelSelector=<selector>'

Architecture

graph LR
PodA["📦 service pod A<br/>Istio/Envoy"]
PodB["📦 service pod B<br/>Istio/Envoy"]
PodC["📦 service pod C<br/>Istio/Envoy"]
Agg(("⚙️ <b>activescale<b/><br/>(Stateless, HA)"))
Redis[("🗄️ shared memory<br/>(e.g., Redis)")]
KEDA{{"🚀 HPA<br/>(average computed by HPA)"}}
PodA -."1. Push<br/>(5s delay)".-> Agg
PodB -."1. Push<br/>(5s delay)".-> Agg
PodC -."1. Push<br/>(5s delay)".-> Agg
Agg ==>|2. update + TTL| Redis
KEDA --"3. Query"--> Agg
Redis -."4. metric value (per pod)".-> Agg
Agg --"5. metric value (per pod)"--> KEDA
style Redis fill:#E3ECF8,stroke:#6E8FB3,stroke-width:2px,color:#000
style KEDA fill:#DDEFD8,stroke:#7DA67D,stroke-width:2px,color:#000
style Agg fill:#E3ECFF,stroke:#6E8FB3,color:#000
style PodA fill:#FAFAFA,stroke:#999
style PodB fill:#FAFAFA,stroke:#999
style PodC fill:#FAFAFA,stroke:#999
Loading

Envoy Metrics Message Shape

Activescale receives Envoy gRPC StreamMetrics messages. Each message contains a node identity and a list of metric families. A simplified shape:

{
"identifier": {
"node": {
"id": "sidecar~10.0.0.1~my-pod.my-ns~my-ns.svc.cluster.local",
"metadata": {
"NAME": "my-pod",
"NAMESPACE": "my-ns"
}
}
},
"envoy_metrics": [
{
"name": "http.inbound_0.0.0.0_8080;.downstream_rq_active",
"metric": [
{ "gauge": { "value": 3 } }
]
},
{
"name": "http.outbound_0.0.0.0_8080;.downstream_rq_active",
"metric": [
{ "gauge": { "value": 4 } }
]
},
{
"name": "listener.0.0.0.0_8080.downstream_cx_active",
"metric": [
{ "gauge": { "value": 12 } }
]
},
{
"name": "cluster.xds-grpc;.circuit_breakers.default.cx_pool_open",
"metric": [
{ "gauge": { "value": 1 } }
]
}
]
}

Pod identity extraction prefers node.metadata.NAME and node.metadata.NAMESPACE. If either metadata field is missing, activescale falls back to parsing the Istio-style node.id (sidecar~<ip>~<pod>.<namespace>~<namespace>.svc.cluster.local).

Summary Counters Meaning

  • messages: number of StreamMetrics messages received (one Recv() call).
  • stored_metrics: number of metric writes stored in Redis for accepted request and connection samples.
  • dropped_by_ids: messages dropped because pod identity could not be extracted.
  • dropped_by_names: metric families skipped because their name did not match the inbound, outbound, or aggregate service listener rules.

stored does not necessarily equal messages because a message can contain multiple metric families or multiple samples, and dropped_by_names is counted per metric family, not per message.

Debugging

Check Custom Metrics API is registered:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2'

Query inbound active requests with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/inbound_active_requests?labelSelector=app=<app>'

Query outbound active requests with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/outbound_active_requests?labelSelector=app=<app>'

Query active connections with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/active_connections?labelSelector=app=<app>'

Check activescale ingest logs:

kubectl logs -n ns-observability deploy/activescale | rg -n "stored (inbound_active_requests|outbound_active_requests|active_connections)|skipping metric name|missing pod identity"

Enable debug logs and restart:

kubectl -n ns-observability set env deploy/activescale LOG_VERBOSITY=4
kubectl -n ns-observability rollout restart deploy/activescale

Confirm Envoy bootstrap includes the metrics service:

istioctl proxy-config bootstrap <pod> -n <ns>| rg -n "envoyMetricsService|metrics_service|envoy_grpc|cluster_name|activescale|9000"

A minimal Kubernetes reference example lives in k8s-manifest-example.yaml.

Notes

Activescale reads both Envoy HTTP connection manager and listener stats, depending on the metric.

  • Activescale exposes inbound and outbound HTTP connection-manager request metrics separately.
  • http.admin.*, http.agent.*, and Prometheus-style envoy_http_downstream_rq_active aliases are ignored.
  • active_connections accepts aggregate listener.<address>.downstream_cx_active families for service listener ports.
  • Worker breakdowns such as listener.<address>.worker_0.downstream_cx_active and Istio-reserved proxy ports for admin, failure detection, debug, telemetry, health, and DNS are ignored.
  • Traffic-path listeners for outbound (15001), inbound (15006), and HBONE (15008) remain eligible.

Appendix

Missing metric values

Envoy can omit an active stat family for a new or idle pod. Activescale refreshes a per-pod heartbeat on every StreamMetrics message and applies the metric TTL to that heartbeat. The default TTL is 20s and can be changed with METRIC_TTL or --ttl.

  • A fresh heartbeat with no metric key is returned as 0 because collection is healthy and the pod has no observed active work.
  • A missing or expired heartbeat remains missing and can produce HTTP 404 NotFound because returning 0 would hide an Envoy collection outage and could cause an unsafe HPA scale-in.

Heartbeat-gated zeros prevent healthy idle pods from causing FailedGetPodsMetric or conservative scale-in delays without treating stale telemetry as zero. Provider summary logs count these values as synthesized_zeros.

Handling HTTP 404 NotFound

HTTP 404 NotFound means telemetry is unavailable, not that the metric value is zero.

Kubernetes HPA retries automatically and handles missing metrics conservatively, as described by the HPA algorithm.

Metric stateHPA action
All configured metrics return HTTP 404 NotFoundKeep the current replicas; neither scale up nor scale down
One metric returns HTTP 404 NotFound, another valid metric requests scale-upScale up using the valid metric
One metric returns HTTP 404 NotFound, another valid metric requests scale-downSkip scale-down and keep the current replicas
All configured metrics are availableUse the largest desired replica count

A direct Custom Metrics API client should apply the same policy: allow scale-up from another valid metric, block scale-down while any metric is unavailable, keep the current replicas when all metrics are unavailable, and retry on the next collection cycle. It must not substitute zero and should alert if HTTP 404 NotFound persists.

About

Push-based Kubernetes HPA adapter using active request/connection metrics to detect load spikes far earlier than CPU/Memory signals while improving resource utilization.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Activescale

Activescale enables fast Kubernetes autoscaling with push-based, per-pod active request and connection metrics from Envoy.

Active means currently in flight: requests not yet completed and connections still open. By Little's Law (L = λW), active requests reflect both request rate and latency.

Unlike CPU and memory metrics collected through metrics-server, Envoy pushes active signals directly to Activescale, avoiding extra scrape and polling delays.

Documentation

  • Guide: 한국어 | English — metric selection, HPA configuration, and verification
  • Motivation: 한국어 | English — rationale, latency comparison, and architecture
  • Test results: 한국어 | English — scaling responsiveness, Scale-out Coverage, and estimated resource utilization summary
  • Kubernetes manifest example: k8s-manifest-example.yaml — deployment and HPA example requiring Redis, ServiceAccount/RBAC, serving certificates, and a target workload

Features

  • Active request concurrency as an autoscaling signal that reflects both request rate and request latency
  • Push-based Envoy StreamMetrics ingestion without intermediary Prometheus scrape and KEDA polling delays
  • Pod-scoped Custom Metrics API values that HPA averages across workload replicas
  • Direction-aware inbound_active_requests and outbound_active_requests, plus aggregate active_connections
  • Stateless high availability with shared Redis/Valkey state and TTL-based stale metric expiry
  • Redis standalone and Cluster modes with optional TLS
  • Configurable gRPC receive limit and Klog-based summary and debug logging

Metrics

  • inbound_active_requests: inbound in-flight HTTP requests from http.inbound_*.downstream_rq_active
  • outbound_active_requests: outbound in-flight HTTP requests from http.outbound_*.downstream_rq_active
  • active_connections: open connections across aggregate Envoy service listeners from listener.<address>.downstream_cx_active

All metrics are exposed as pod-scoped custom metrics for HPA.

Choosing a request metric

Choose the metric based on the role of the workload being scaled.

WorkloadMetricWhen to use
Application or API serviceinbound_active_requestsUse for a workload that receives and processes client requests. This is the default choice for typical services.
Dedicated gateway or proxyoutbound_active_requestsUse when the workload primarily forwards client requests to upstream services and inbound_active_requests is unavailable or does not reflect the traffic.

Do not configure both metrics as a fallback. For a normal application, outbound requests usually represent calls to dependencies and can trigger scaling for unrelated downstream traffic.

If the workload role is unclear, generate a small amount of client traffic and query both metrics. Select the one that returns per-pod values and increases with the client request concurrency being tested.

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/inbound_active_requests?labelSelector=<selector>'
kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/outbound_active_requests?labelSelector=<selector>'

Architecture

graph LR
PodA["📦 service pod A<br/>Istio/Envoy"]
PodB["📦 service pod B<br/>Istio/Envoy"]
PodC["📦 service pod C<br/>Istio/Envoy"]
Agg(("⚙️ <b>activescale<b/><br/>(Stateless, HA)"))
Redis[("🗄️ shared memory<br/>(e.g., Redis)")]
KEDA{{"🚀 HPA<br/>(average computed by HPA)"}}
PodA -."1. Push<br/>(5s delay)".-> Agg
PodB -."1. Push<br/>(5s delay)".-> Agg
PodC -."1. Push<br/>(5s delay)".-> Agg
Agg ==>|2. update + TTL| Redis
KEDA --"3. Query"--> Agg
Redis -."4. metric value (per pod)".-> Agg
Agg --"5. metric value (per pod)"--> KEDA
style Redis fill:#E3ECF8,stroke:#6E8FB3,stroke-width:2px,color:#000
style KEDA fill:#DDEFD8,stroke:#7DA67D,stroke-width:2px,color:#000
style Agg fill:#E3ECFF,stroke:#6E8FB3,color:#000
style PodA fill:#FAFAFA,stroke:#999
style PodB fill:#FAFAFA,stroke:#999
style PodC fill:#FAFAFA,stroke:#999
Loading

Envoy Metrics Message Shape

Activescale receives Envoy gRPC StreamMetrics messages. Each message contains a node identity and a list of metric families. A simplified shape:

{
"identifier": {
"node": {
"id": "sidecar~10.0.0.1~my-pod.my-ns~my-ns.svc.cluster.local",
"metadata": {
"NAME": "my-pod",
"NAMESPACE": "my-ns"
}
}
},
"envoy_metrics": [
{
"name": "http.inbound_0.0.0.0_8080;.downstream_rq_active",
"metric": [
{ "gauge": { "value": 3 } }
]
},
{
"name": "http.outbound_0.0.0.0_8080;.downstream_rq_active",
"metric": [
{ "gauge": { "value": 4 } }
]
},
{
"name": "listener.0.0.0.0_8080.downstream_cx_active",
"metric": [
{ "gauge": { "value": 12 } }
]
},
{
"name": "cluster.xds-grpc;.circuit_breakers.default.cx_pool_open",
"metric": [
{ "gauge": { "value": 1 } }
]
}
]
}

Pod identity extraction prefers node.metadata.NAME and node.metadata.NAMESPACE. If either metadata field is missing, activescale falls back to parsing the Istio-style node.id (sidecar~<ip>~<pod>.<namespace>~<namespace>.svc.cluster.local).

Summary Counters Meaning

  • messages: number of StreamMetrics messages received (one Recv() call).
  • stored_metrics: number of metric writes stored in Redis for accepted request and connection samples.
  • dropped_by_ids: messages dropped because pod identity could not be extracted.
  • dropped_by_names: metric families skipped because their name did not match the inbound, outbound, or aggregate service listener rules.

stored does not necessarily equal messages because a message can contain multiple metric families or multiple samples, and dropped_by_names is counted per metric family, not per message.

Debugging

Check Custom Metrics API is registered:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2'

Query inbound active requests with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/inbound_active_requests?labelSelector=app=<app>'

Query outbound active requests with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/outbound_active_requests?labelSelector=app=<app>'

Query active connections with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/active_connections?labelSelector=app=<app>'

Check activescale ingest logs:

kubectl logs -n ns-observability deploy/activescale | rg -n "stored (inbound_active_requests|outbound_active_requests|active_connections)|skipping metric name|missing pod identity"

Enable debug logs and restart:

kubectl -n ns-observability set env deploy/activescale LOG_VERBOSITY=4
kubectl -n ns-observability rollout restart deploy/activescale

Confirm Envoy bootstrap includes the metrics service:

istioctl proxy-config bootstrap <pod> -n <ns>| rg -n "envoyMetricsService|metrics_service|envoy_grpc|cluster_name|activescale|9000"

A minimal Kubernetes reference example lives in k8s-manifest-example.yaml.

Notes

Activescale reads both Envoy HTTP connection manager and listener stats, depending on the metric.

  • Activescale exposes inbound and outbound HTTP connection-manager request metrics separately.
  • http.admin.*, http.agent.*, and Prometheus-style envoy_http_downstream_rq_active aliases are ignored.
  • active_connections accepts aggregate listener.<address>.downstream_cx_active families for service listener ports.
  • Worker breakdowns such as listener.<address>.worker_0.downstream_cx_active and Istio-reserved proxy ports for admin, failure detection, debug, telemetry, health, and DNS are ignored.
  • Traffic-path listeners for outbound (15001), inbound (15006), and HBONE (15008) remain eligible.

Appendix

Missing metric values

Envoy can omit an active stat family for a new or idle pod. Activescale refreshes a per-pod heartbeat on every StreamMetrics message and applies the metric TTL to that heartbeat. The default TTL is 20s and can be changed with METRIC_TTL or --ttl.

  • A fresh heartbeat with no metric key is returned as 0 because collection is healthy and the pod has no observed active work.
  • A missing or expired heartbeat remains missing and can produce HTTP 404 NotFound because returning 0 would hide an Envoy collection outage and could cause an unsafe HPA scale-in.

Heartbeat-gated zeros prevent healthy idle pods from causing FailedGetPodsMetric or conservative scale-in delays without treating stale telemetry as zero. Provider summary logs count these values as synthesized_zeros.

Handling HTTP 404 NotFound

HTTP 404 NotFound means telemetry is unavailable, not that the metric value is zero.

Kubernetes HPA retries automatically and handles missing metrics conservatively, as described by the HPA algorithm.

Metric stateHPA action
All configured metrics return HTTP 404 NotFoundKeep the current replicas; neither scale up nor scale down
One metric returns HTTP 404 NotFound, another valid metric requests scale-upScale up using the valid metric
One metric returns HTTP 404 NotFound, another valid metric requests scale-downSkip scale-down and keep the current replicas
All configured metrics are availableUse the largest desired replica count

A direct Custom Metrics API client should apply the same policy: allow scale-up from another valid metric, block scale-down while any metric is unavailable, keep the current replicas when all metrics are unavailable, and retry on the next collection cycle. It must not substitute zero and should alert if HTTP 404 NotFound persists.

About

Push-based Kubernetes HPA adapter using active request/connection metrics to detect load spikes far earlier than CPU/Memory signals while improving resource utilization.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Activescale

Activescale enables fast Kubernetes autoscaling with push-based, per-pod active request and connection metrics from Envoy.

Active means currently in flight: requests not yet completed and connections still open. By Little's Law (L = λW), active requests reflect both request rate and latency.

Unlike CPU and memory metrics collected through metrics-server, Envoy pushes active signals directly to Activescale, avoiding extra scrape and polling delays.

Documentation

  • Guide: 한국어 | English — metric selection, HPA configuration, and verification
  • Motivation: 한국어 | English — rationale, latency comparison, and architecture
  • Test results: 한국어 | English — scaling responsiveness, Scale-out Coverage, and estimated resource utilization summary
  • Kubernetes manifest example: k8s-manifest-example.yaml — deployment and HPA example requiring Redis, ServiceAccount/RBAC, serving certificates, and a target workload

Features

  • Active request concurrency as an autoscaling signal that reflects both request rate and request latency
  • Push-based Envoy StreamMetrics ingestion without intermediary Prometheus scrape and KEDA polling delays
  • Pod-scoped Custom Metrics API values that HPA averages across workload replicas
  • Direction-aware inbound_active_requests and outbound_active_requests, plus aggregate active_connections
  • Stateless high availability with shared Redis/Valkey state and TTL-based stale metric expiry
  • Redis standalone and Cluster modes with optional TLS
  • Configurable gRPC receive limit and Klog-based summary and debug logging

Metrics

  • inbound_active_requests: inbound in-flight HTTP requests from http.inbound_*.downstream_rq_active
  • outbound_active_requests: outbound in-flight HTTP requests from http.outbound_*.downstream_rq_active
  • active_connections: open connections across aggregate Envoy service listeners from listener.<address>.downstream_cx_active

All metrics are exposed as pod-scoped custom metrics for HPA.

Choosing a request metric

Choose the metric based on the role of the workload being scaled.

WorkloadMetricWhen to use
Application or API serviceinbound_active_requestsUse for a workload that receives and processes client requests. This is the default choice for typical services.
Dedicated gateway or proxyoutbound_active_requestsUse when the workload primarily forwards client requests to upstream services and inbound_active_requests is unavailable or does not reflect the traffic.

Do not configure both metrics as a fallback. For a normal application, outbound requests usually represent calls to dependencies and can trigger scaling for unrelated downstream traffic.

If the workload role is unclear, generate a small amount of client traffic and query both metrics. Select the one that returns per-pod values and increases with the client request concurrency being tested.

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/inbound_active_requests?labelSelector=<selector>'
kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/outbound_active_requests?labelSelector=<selector>'

Architecture

graph LR
PodA["📦 service pod A<br/>Istio/Envoy"]
PodB["📦 service pod B<br/>Istio/Envoy"]
PodC["📦 service pod C<br/>Istio/Envoy"]
Agg(("⚙️ <b>activescale<b/><br/>(Stateless, HA)"))
Redis[("🗄️ shared memory<br/>(e.g., Redis)")]
KEDA{{"🚀 HPA<br/>(average computed by HPA)"}}
PodA -."1. Push<br/>(5s delay)".-> Agg
PodB -."1. Push<br/>(5s delay)".-> Agg
PodC -."1. Push<br/>(5s delay)".-> Agg
Agg ==>|2. update + TTL| Redis
KEDA --"3. Query"--> Agg
Redis -."4. metric value (per pod)".-> Agg
Agg --"5. metric value (per pod)"--> KEDA
style Redis fill:#E3ECF8,stroke:#6E8FB3,stroke-width:2px,color:#000
style KEDA fill:#DDEFD8,stroke:#7DA67D,stroke-width:2px,color:#000
style Agg fill:#E3ECFF,stroke:#6E8FB3,color:#000
style PodA fill:#FAFAFA,stroke:#999
style PodB fill:#FAFAFA,stroke:#999
style PodC fill:#FAFAFA,stroke:#999
Loading

Envoy Metrics Message Shape

Activescale receives Envoy gRPC StreamMetrics messages. Each message contains a node identity and a list of metric families. A simplified shape:

{
"identifier": {
"node": {
"id": "sidecar~10.0.0.1~my-pod.my-ns~my-ns.svc.cluster.local",
"metadata": {
"NAME": "my-pod",
"NAMESPACE": "my-ns"
}
}
},
"envoy_metrics": [
{
"name": "http.inbound_0.0.0.0_8080;.downstream_rq_active",
"metric": [
{ "gauge": { "value": 3 } }
]
},
{
"name": "http.outbound_0.0.0.0_8080;.downstream_rq_active",
"metric": [
{ "gauge": { "value": 4 } }
]
},
{
"name": "listener.0.0.0.0_8080.downstream_cx_active",
"metric": [
{ "gauge": { "value": 12 } }
]
},
{
"name": "cluster.xds-grpc;.circuit_breakers.default.cx_pool_open",
"metric": [
{ "gauge": { "value": 1 } }
]
}
]
}

Pod identity extraction prefers node.metadata.NAME and node.metadata.NAMESPACE. If either metadata field is missing, activescale falls back to parsing the Istio-style node.id (sidecar~<ip>~<pod>.<namespace>~<namespace>.svc.cluster.local).

Summary Counters Meaning

  • messages: number of StreamMetrics messages received (one Recv() call).
  • stored_metrics: number of metric writes stored in Redis for accepted request and connection samples.
  • dropped_by_ids: messages dropped because pod identity could not be extracted.
  • dropped_by_names: metric families skipped because their name did not match the inbound, outbound, or aggregate service listener rules.

stored does not necessarily equal messages because a message can contain multiple metric families or multiple samples, and dropped_by_names is counted per metric family, not per message.

Debugging

Check Custom Metrics API is registered:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2'

Query inbound active requests with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/inbound_active_requests?labelSelector=app=<app>'

Query outbound active requests with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/outbound_active_requests?labelSelector=app=<app>'

Query active connections with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/active_connections?labelSelector=app=<app>'

Check activescale ingest logs:

kubectl logs -n ns-observability deploy/activescale | rg -n "stored (inbound_active_requests|outbound_active_requests|active_connections)|skipping metric name|missing pod identity"

Enable debug logs and restart:

kubectl -n ns-observability set env deploy/activescale LOG_VERBOSITY=4
kubectl -n ns-observability rollout restart deploy/activescale

Confirm Envoy bootstrap includes the metrics service:

istioctl proxy-config bootstrap <pod> -n <ns>| rg -n "envoyMetricsService|metrics_service|envoy_grpc|cluster_name|activescale|9000"

A minimal Kubernetes reference example lives in k8s-manifest-example.yaml.

Notes

Activescale reads both Envoy HTTP connection manager and listener stats, depending on the metric.

  • Activescale exposes inbound and outbound HTTP connection-manager request metrics separately.
  • http.admin.*, http.agent.*, and Prometheus-style envoy_http_downstream_rq_active aliases are ignored.
  • active_connections accepts aggregate listener.<address>.downstream_cx_active families for service listener ports.
  • Worker breakdowns such as listener.<address>.worker_0.downstream_cx_active and Istio-reserved proxy ports for admin, failure detection, debug, telemetry, health, and DNS are ignored.
  • Traffic-path listeners for outbound (15001), inbound (15006), and HBONE (15008) remain eligible.

Appendix

Missing metric values

Envoy can omit an active stat family for a new or idle pod. Activescale refreshes a per-pod heartbeat on every StreamMetrics message and applies the metric TTL to that heartbeat. The default TTL is 20s and can be changed with METRIC_TTL or --ttl.

  • A fresh heartbeat with no metric key is returned as 0 because collection is healthy and the pod has no observed active work.
  • A missing or expired heartbeat remains missing and can produce HTTP 404 NotFound because returning 0 would hide an Envoy collection outage and could cause an unsafe HPA scale-in.

Heartbeat-gated zeros prevent healthy idle pods from causing FailedGetPodsMetric or conservative scale-in delays without treating stale telemetry as zero. Provider summary logs count these values as synthesized_zeros.

Handling HTTP 404 NotFound

HTTP 404 NotFound means telemetry is unavailable, not that the metric value is zero.

Kubernetes HPA retries automatically and handles missing metrics conservatively, as described by the HPA algorithm.

Metric stateHPA action
All configured metrics return HTTP 404 NotFoundKeep the current replicas; neither scale up nor scale down
One metric returns HTTP 404 NotFound, another valid metric requests scale-upScale up using the valid metric
One metric returns HTTP 404 NotFound, another valid metric requests scale-downSkip scale-down and keep the current replicas
All configured metrics are availableUse the largest desired replica count

A direct Custom Metrics API client should apply the same policy: allow scale-up from another valid metric, block scale-down while any metric is unavailable, keep the current replicas when all metrics are unavailable, and retry on the next collection cycle. It must not substitute zero and should alert if HTTP 404 NotFound persists.

About

Push-based Kubernetes HPA adapter using active request/connection metrics to detect load spikes far earlier than CPU/Memory signals while improving resource utilization.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Activescale

Activescale enables fast Kubernetes autoscaling with push-based, per-pod active request and connection metrics from Envoy.

Active means currently in flight: requests not yet completed and connections still open. By Little's Law (L = λW), active requests reflect both request rate and latency.

Unlike CPU and memory metrics collected through metrics-server, Envoy pushes active signals directly to Activescale, avoiding extra scrape and polling delays.

Documentation

  • Guide: 한국어 | English — metric selection, HPA configuration, and verification
  • Motivation: 한국어 | English — rationale, latency comparison, and architecture
  • Test results: 한국어 | English — scaling responsiveness, Scale-out Coverage, and estimated resource utilization summary
  • Kubernetes manifest example: k8s-manifest-example.yaml — deployment and HPA example requiring Redis, ServiceAccount/RBAC, serving certificates, and a target workload

Features

  • Active request concurrency as an autoscaling signal that reflects both request rate and request latency
  • Push-based Envoy StreamMetrics ingestion without intermediary Prometheus scrape and KEDA polling delays
  • Pod-scoped Custom Metrics API values that HPA averages across workload replicas
  • Direction-aware inbound_active_requests and outbound_active_requests, plus aggregate active_connections
  • Stateless high availability with shared Redis/Valkey state and TTL-based stale metric expiry
  • Redis standalone and Cluster modes with optional TLS
  • Configurable gRPC receive limit and Klog-based summary and debug logging

Metrics

  • inbound_active_requests: inbound in-flight HTTP requests from http.inbound_*.downstream_rq_active
  • outbound_active_requests: outbound in-flight HTTP requests from http.outbound_*.downstream_rq_active
  • active_connections: open connections across aggregate Envoy service listeners from listener.<address>.downstream_cx_active

All metrics are exposed as pod-scoped custom metrics for HPA.

Choosing a request metric

Choose the metric based on the role of the workload being scaled.

WorkloadMetricWhen to use
Application or API serviceinbound_active_requestsUse for a workload that receives and processes client requests. This is the default choice for typical services.
Dedicated gateway or proxyoutbound_active_requestsUse when the workload primarily forwards client requests to upstream services and inbound_active_requests is unavailable or does not reflect the traffic.

Do not configure both metrics as a fallback. For a normal application, outbound requests usually represent calls to dependencies and can trigger scaling for unrelated downstream traffic.

If the workload role is unclear, generate a small amount of client traffic and query both metrics. Select the one that returns per-pod values and increases with the client request concurrency being tested.

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/inbound_active_requests?labelSelector=<selector>'
kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/outbound_active_requests?labelSelector=<selector>'

Architecture

graph LR
PodA["📦 service pod A<br/>Istio/Envoy"]
PodB["📦 service pod B<br/>Istio/Envoy"]
PodC["📦 service pod C<br/>Istio/Envoy"]
Agg(("⚙️ <b>activescale<b/><br/>(Stateless, HA)"))
Redis[("🗄️ shared memory<br/>(e.g., Redis)")]
KEDA{{"🚀 HPA<br/>(average computed by HPA)"}}
PodA -."1. Push<br/>(5s delay)".-> Agg
PodB -."1. Push<br/>(5s delay)".-> Agg
PodC -."1. Push<br/>(5s delay)".-> Agg
Agg ==>|2. update + TTL| Redis
KEDA --"3. Query"--> Agg
Redis -."4. metric value (per pod)".-> Agg
Agg --"5. metric value (per pod)"--> KEDA
style Redis fill:#E3ECF8,stroke:#6E8FB3,stroke-width:2px,color:#000
style KEDA fill:#DDEFD8,stroke:#7DA67D,stroke-width:2px,color:#000
style Agg fill:#E3ECFF,stroke:#6E8FB3,color:#000
style PodA fill:#FAFAFA,stroke:#999
style PodB fill:#FAFAFA,stroke:#999
style PodC fill:#FAFAFA,stroke:#999
Loading

Envoy Metrics Message Shape

Activescale receives Envoy gRPC StreamMetrics messages. Each message contains a node identity and a list of metric families. A simplified shape:

{
"identifier": {
"node": {
"id": "sidecar~10.0.0.1~my-pod.my-ns~my-ns.svc.cluster.local",
"metadata": {
"NAME": "my-pod",
"NAMESPACE": "my-ns"
}
}
},
"envoy_metrics": [
{
"name": "http.inbound_0.0.0.0_8080;.downstream_rq_active",
"metric": [
{ "gauge": { "value": 3 } }
]
},
{
"name": "http.outbound_0.0.0.0_8080;.downstream_rq_active",
"metric": [
{ "gauge": { "value": 4 } }
]
},
{
"name": "listener.0.0.0.0_8080.downstream_cx_active",
"metric": [
{ "gauge": { "value": 12 } }
]
},
{
"name": "cluster.xds-grpc;.circuit_breakers.default.cx_pool_open",
"metric": [
{ "gauge": { "value": 1 } }
]
}
]
}

Pod identity extraction prefers node.metadata.NAME and node.metadata.NAMESPACE. If either metadata field is missing, activescale falls back to parsing the Istio-style node.id (sidecar~<ip>~<pod>.<namespace>~<namespace>.svc.cluster.local).

Summary Counters Meaning

  • messages: number of StreamMetrics messages received (one Recv() call).
  • stored_metrics: number of metric writes stored in Redis for accepted request and connection samples.
  • dropped_by_ids: messages dropped because pod identity could not be extracted.
  • dropped_by_names: metric families skipped because their name did not match the inbound, outbound, or aggregate service listener rules.

stored does not necessarily equal messages because a message can contain multiple metric families or multiple samples, and dropped_by_names is counted per metric family, not per message.

Debugging

Check Custom Metrics API is registered:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2'

Query inbound active requests with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/inbound_active_requests?labelSelector=app=<app>'

Query outbound active requests with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/outbound_active_requests?labelSelector=app=<app>'

Query active connections with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/active_connections?labelSelector=app=<app>'

Check activescale ingest logs:

kubectl logs -n ns-observability deploy/activescale | rg -n "stored (inbound_active_requests|outbound_active_requests|active_connections)|skipping metric name|missing pod identity"

Enable debug logs and restart:

kubectl -n ns-observability set env deploy/activescale LOG_VERBOSITY=4
kubectl -n ns-observability rollout restart deploy/activescale

Confirm Envoy bootstrap includes the metrics service:

istioctl proxy-config bootstrap <pod> -n <ns>| rg -n "envoyMetricsService|metrics_service|envoy_grpc|cluster_name|activescale|9000"

A minimal Kubernetes reference example lives in k8s-manifest-example.yaml.

Notes

Activescale reads both Envoy HTTP connection manager and listener stats, depending on the metric.

  • Activescale exposes inbound and outbound HTTP connection-manager request metrics separately.
  • http.admin.*, http.agent.*, and Prometheus-style envoy_http_downstream_rq_active aliases are ignored.
  • active_connections accepts aggregate listener.<address>.downstream_cx_active families for service listener ports.
  • Worker breakdowns such as listener.<address>.worker_0.downstream_cx_active and Istio-reserved proxy ports for admin, failure detection, debug, telemetry, health, and DNS are ignored.
  • Traffic-path listeners for outbound (15001), inbound (15006), and HBONE (15008) remain eligible.

Appendix

Missing metric values

Envoy can omit an active stat family for a new or idle pod. Activescale refreshes a per-pod heartbeat on every StreamMetrics message and applies the metric TTL to that heartbeat. The default TTL is 20s and can be changed with METRIC_TTL or --ttl.

  • A fresh heartbeat with no metric key is returned as 0 because collection is healthy and the pod has no observed active work.
  • A missing or expired heartbeat remains missing and can produce HTTP 404 NotFound because returning 0 would hide an Envoy collection outage and could cause an unsafe HPA scale-in.

Heartbeat-gated zeros prevent healthy idle pods from causing FailedGetPodsMetric or conservative scale-in delays without treating stale telemetry as zero. Provider summary logs count these values as synthesized_zeros.

Handling HTTP 404 NotFound

HTTP 404 NotFound means telemetry is unavailable, not that the metric value is zero.

Kubernetes HPA retries automatically and handles missing metrics conservatively, as described by the HPA algorithm.

Metric stateHPA action
All configured metrics return HTTP 404 NotFoundKeep the current replicas; neither scale up nor scale down
One metric returns HTTP 404 NotFound, another valid metric requests scale-upScale up using the valid metric
One metric returns HTTP 404 NotFound, another valid metric requests scale-downSkip scale-down and keep the current replicas
All configured metrics are availableUse the largest desired replica count

A direct Custom Metrics API client should apply the same policy: allow scale-up from another valid metric, block scale-down while any metric is unavailable, keep the current replicas when all metrics are unavailable, and retry on the next collection cycle. It must not substitute zero and should alert if HTTP 404 NotFound persists.

About

Push-based Kubernetes HPA adapter using active request/connection metrics to detect load spikes far earlier than CPU/Memory signals while improving resource utilization.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Activescale

Activescale enables fast Kubernetes autoscaling with push-based, per-pod active request and connection metrics from Envoy.

Active means currently in flight: requests not yet completed and connections still open. By Little's Law (L = λW), active requests reflect both request rate and latency.

Unlike CPU and memory metrics collected through metrics-server, Envoy pushes active signals directly to Activescale, avoiding extra scrape and polling delays.

Documentation

  • Guide: 한국어 | English — metric selection, HPA configuration, and verification
  • Motivation: 한국어 | English — rationale, latency comparison, and architecture
  • Test results: 한국어 | English — scaling responsiveness, Scale-out Coverage, and estimated resource utilization summary
  • Kubernetes manifest example: k8s-manifest-example.yaml — deployment and HPA example requiring Redis, ServiceAccount/RBAC, serving certificates, and a target workload

Features

  • Active request concurrency as an autoscaling signal that reflects both request rate and request latency
  • Push-based Envoy StreamMetrics ingestion without intermediary Prometheus scrape and KEDA polling delays
  • Pod-scoped Custom Metrics API values that HPA averages across workload replicas
  • Direction-aware inbound_active_requests and outbound_active_requests, plus aggregate active_connections
  • Stateless high availability with shared Redis/Valkey state and TTL-based stale metric expiry
  • Redis standalone and Cluster modes with optional TLS
  • Configurable gRPC receive limit and Klog-based summary and debug logging

Metrics

  • inbound_active_requests: inbound in-flight HTTP requests from http.inbound_*.downstream_rq_active
  • outbound_active_requests: outbound in-flight HTTP requests from http.outbound_*.downstream_rq_active
  • active_connections: open connections across aggregate Envoy service listeners from listener.<address>.downstream_cx_active

All metrics are exposed as pod-scoped custom metrics for HPA.

Choosing a request metric

Choose the metric based on the role of the workload being scaled.

WorkloadMetricWhen to use
Application or API serviceinbound_active_requestsUse for a workload that receives and processes client requests. This is the default choice for typical services.
Dedicated gateway or proxyoutbound_active_requestsUse when the workload primarily forwards client requests to upstream services and inbound_active_requests is unavailable or does not reflect the traffic.

Do not configure both metrics as a fallback. For a normal application, outbound requests usually represent calls to dependencies and can trigger scaling for unrelated downstream traffic.

If the workload role is unclear, generate a small amount of client traffic and query both metrics. Select the one that returns per-pod values and increases with the client request concurrency being tested.

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/inbound_active_requests?labelSelector=<selector>'
kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/outbound_active_requests?labelSelector=<selector>'

Architecture

graph LR
PodA["📦 service pod A<br/>Istio/Envoy"]
PodB["📦 service pod B<br/>Istio/Envoy"]
PodC["📦 service pod C<br/>Istio/Envoy"]
Agg(("⚙️ <b>activescale<b/><br/>(Stateless, HA)"))
Redis[("🗄️ shared memory<br/>(e.g., Redis)")]
KEDA{{"🚀 HPA<br/>(average computed by HPA)"}}
PodA -."1. Push<br/>(5s delay)".-> Agg
PodB -."1. Push<br/>(5s delay)".-> Agg
PodC -."1. Push<br/>(5s delay)".-> Agg
Agg ==>|2. update + TTL| Redis
KEDA --"3. Query"--> Agg
Redis -."4. metric value (per pod)".-> Agg
Agg --"5. metric value (per pod)"--> KEDA
style Redis fill:#E3ECF8,stroke:#6E8FB3,stroke-width:2px,color:#000
style KEDA fill:#DDEFD8,stroke:#7DA67D,stroke-width:2px,color:#000
style Agg fill:#E3ECFF,stroke:#6E8FB3,color:#000
style PodA fill:#FAFAFA,stroke:#999
style PodB fill:#FAFAFA,stroke:#999
style PodC fill:#FAFAFA,stroke:#999
Loading

Envoy Metrics Message Shape

Activescale receives Envoy gRPC StreamMetrics messages. Each message contains a node identity and a list of metric families. A simplified shape:

{
"identifier": {
"node": {
"id": "sidecar~10.0.0.1~my-pod.my-ns~my-ns.svc.cluster.local",
"metadata": {
"NAME": "my-pod",
"NAMESPACE": "my-ns"
}
}
},
"envoy_metrics": [
{
"name": "http.inbound_0.0.0.0_8080;.downstream_rq_active",
"metric": [
{ "gauge": { "value": 3 } }
]
},
{
"name": "http.outbound_0.0.0.0_8080;.downstream_rq_active",
"metric": [
{ "gauge": { "value": 4 } }
]
},
{
"name": "listener.0.0.0.0_8080.downstream_cx_active",
"metric": [
{ "gauge": { "value": 12 } }
]
},
{
"name": "cluster.xds-grpc;.circuit_breakers.default.cx_pool_open",
"metric": [
{ "gauge": { "value": 1 } }
]
}
]
}

Pod identity extraction prefers node.metadata.NAME and node.metadata.NAMESPACE. If either metadata field is missing, activescale falls back to parsing the Istio-style node.id (sidecar~<ip>~<pod>.<namespace>~<namespace>.svc.cluster.local).

Summary Counters Meaning

  • messages: number of StreamMetrics messages received (one Recv() call).
  • stored_metrics: number of metric writes stored in Redis for accepted request and connection samples.
  • dropped_by_ids: messages dropped because pod identity could not be extracted.
  • dropped_by_names: metric families skipped because their name did not match the inbound, outbound, or aggregate service listener rules.

stored does not necessarily equal messages because a message can contain multiple metric families or multiple samples, and dropped_by_names is counted per metric family, not per message.

Debugging

Check Custom Metrics API is registered:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2'

Query inbound active requests with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/inbound_active_requests?labelSelector=app=<app>'

Query outbound active requests with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/outbound_active_requests?labelSelector=app=<app>'

Query active connections with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/active_connections?labelSelector=app=<app>'

Check activescale ingest logs:

kubectl logs -n ns-observability deploy/activescale | rg -n "stored (inbound_active_requests|outbound_active_requests|active_connections)|skipping metric name|missing pod identity"

Enable debug logs and restart:

kubectl -n ns-observability set env deploy/activescale LOG_VERBOSITY=4
kubectl -n ns-observability rollout restart deploy/activescale

Confirm Envoy bootstrap includes the metrics service:

istioctl proxy-config bootstrap <pod> -n <ns>| rg -n "envoyMetricsService|metrics_service|envoy_grpc|cluster_name|activescale|9000"

A minimal Kubernetes reference example lives in k8s-manifest-example.yaml.

Notes

Activescale reads both Envoy HTTP connection manager and listener stats, depending on the metric.

  • Activescale exposes inbound and outbound HTTP connection-manager request metrics separately.
  • http.admin.*, http.agent.*, and Prometheus-style envoy_http_downstream_rq_active aliases are ignored.
  • active_connections accepts aggregate listener.<address>.downstream_cx_active families for service listener ports.
  • Worker breakdowns such as listener.<address>.worker_0.downstream_cx_active and Istio-reserved proxy ports for admin, failure detection, debug, telemetry, health, and DNS are ignored.
  • Traffic-path listeners for outbound (15001), inbound (15006), and HBONE (15008) remain eligible.

Appendix

Missing metric values

Envoy can omit an active stat family for a new or idle pod. Activescale refreshes a per-pod heartbeat on every StreamMetrics message and applies the metric TTL to that heartbeat. The default TTL is 20s and can be changed with METRIC_TTL or --ttl.

  • A fresh heartbeat with no metric key is returned as 0 because collection is healthy and the pod has no observed active work.
  • A missing or expired heartbeat remains missing and can produce HTTP 404 NotFound because returning 0 would hide an Envoy collection outage and could cause an unsafe HPA scale-in.

Heartbeat-gated zeros prevent healthy idle pods from causing FailedGetPodsMetric or conservative scale-in delays without treating stale telemetry as zero. Provider summary logs count these values as synthesized_zeros.

Handling HTTP 404 NotFound

HTTP 404 NotFound means telemetry is unavailable, not that the metric value is zero.

Kubernetes HPA retries automatically and handles missing metrics conservatively, as described by the HPA algorithm.

Metric stateHPA action
All configured metrics return HTTP 404 NotFoundKeep the current replicas; neither scale up nor scale down
One metric returns HTTP 404 NotFound, another valid metric requests scale-upScale up using the valid metric
One metric returns HTTP 404 NotFound, another valid metric requests scale-downSkip scale-down and keep the current replicas
All configured metrics are availableUse the largest desired replica count

A direct Custom Metrics API client should apply the same policy: allow scale-up from another valid metric, block scale-down while any metric is unavailable, keep the current replicas when all metrics are unavailable, and retry on the next collection cycle. It must not substitute zero and should alert if HTTP 404 NotFound persists.

About

Push-based Kubernetes HPA adapter using active request/connection metrics to detect load spikes far earlier than CPU/Memory signals while improving resource utilization.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Activescale

Activescale enables fast Kubernetes autoscaling with push-based, per-pod active request and connection metrics from Envoy.

Active means currently in flight: requests not yet completed and connections still open. By Little's Law (L = λW), active requests reflect both request rate and latency.

Unlike CPU and memory metrics collected through metrics-server, Envoy pushes active signals directly to Activescale, avoiding extra scrape and polling delays.

Documentation

  • Guide: 한국어 | English — metric selection, HPA configuration, and verification
  • Motivation: 한국어 | English — rationale, latency comparison, and architecture
  • Test results: 한국어 | English — scaling responsiveness, Scale-out Coverage, and estimated resource utilization summary
  • Kubernetes manifest example: k8s-manifest-example.yaml — deployment and HPA example requiring Redis, ServiceAccount/RBAC, serving certificates, and a target workload

Features

  • Active request concurrency as an autoscaling signal that reflects both request rate and request latency
  • Push-based Envoy StreamMetrics ingestion without intermediary Prometheus scrape and KEDA polling delays
  • Pod-scoped Custom Metrics API values that HPA averages across workload replicas
  • Direction-aware inbound_active_requests and outbound_active_requests, plus aggregate active_connections
  • Stateless high availability with shared Redis/Valkey state and TTL-based stale metric expiry
  • Redis standalone and Cluster modes with optional TLS
  • Configurable gRPC receive limit and Klog-based summary and debug logging

Metrics

  • inbound_active_requests: inbound in-flight HTTP requests from http.inbound_*.downstream_rq_active
  • outbound_active_requests: outbound in-flight HTTP requests from http.outbound_*.downstream_rq_active
  • active_connections: open connections across aggregate Envoy service listeners from listener.<address>.downstream_cx_active

All metrics are exposed as pod-scoped custom metrics for HPA.

Choosing a request metric

Choose the metric based on the role of the workload being scaled.

WorkloadMetricWhen to use
Application or API serviceinbound_active_requestsUse for a workload that receives and processes client requests. This is the default choice for typical services.
Dedicated gateway or proxyoutbound_active_requestsUse when the workload primarily forwards client requests to upstream services and inbound_active_requests is unavailable or does not reflect the traffic.

Do not configure both metrics as a fallback. For a normal application, outbound requests usually represent calls to dependencies and can trigger scaling for unrelated downstream traffic.

If the workload role is unclear, generate a small amount of client traffic and query both metrics. Select the one that returns per-pod values and increases with the client request concurrency being tested.

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/inbound_active_requests?labelSelector=<selector>'
kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/outbound_active_requests?labelSelector=<selector>'

Architecture

graph LR
PodA["📦 service pod A<br/>Istio/Envoy"]
PodB["📦 service pod B<br/>Istio/Envoy"]
PodC["📦 service pod C<br/>Istio/Envoy"]
Agg(("⚙️ <b>activescale<b/><br/>(Stateless, HA)"))
Redis[("🗄️ shared memory<br/>(e.g., Redis)")]
KEDA{{"🚀 HPA<br/>(average computed by HPA)"}}
PodA -."1. Push<br/>(5s delay)".-> Agg
PodB -."1. Push<br/>(5s delay)".-> Agg
PodC -."1. Push<br/>(5s delay)".-> Agg
Agg ==>|2. update + TTL| Redis
KEDA --"3. Query"--> Agg
Redis -."4. metric value (per pod)".-> Agg
Agg --"5. metric value (per pod)"--> KEDA
style Redis fill:#E3ECF8,stroke:#6E8FB3,stroke-width:2px,color:#000
style KEDA fill:#DDEFD8,stroke:#7DA67D,stroke-width:2px,color:#000
style Agg fill:#E3ECFF,stroke:#6E8FB3,color:#000
style PodA fill:#FAFAFA,stroke:#999
style PodB fill:#FAFAFA,stroke:#999
style PodC fill:#FAFAFA,stroke:#999
Loading

Envoy Metrics Message Shape

Activescale receives Envoy gRPC StreamMetrics messages. Each message contains a node identity and a list of metric families. A simplified shape:

{
"identifier": {
"node": {
"id": "sidecar~10.0.0.1~my-pod.my-ns~my-ns.svc.cluster.local",
"metadata": {
"NAME": "my-pod",
"NAMESPACE": "my-ns"
}
}
},
"envoy_metrics": [
{
"name": "http.inbound_0.0.0.0_8080;.downstream_rq_active",
"metric": [
{ "gauge": { "value": 3 } }
]
},
{
"name": "http.outbound_0.0.0.0_8080;.downstream_rq_active",
"metric": [
{ "gauge": { "value": 4 } }
]
},
{
"name": "listener.0.0.0.0_8080.downstream_cx_active",
"metric": [
{ "gauge": { "value": 12 } }
]
},
{
"name": "cluster.xds-grpc;.circuit_breakers.default.cx_pool_open",
"metric": [
{ "gauge": { "value": 1 } }
]
}
]
}

Pod identity extraction prefers node.metadata.NAME and node.metadata.NAMESPACE. If either metadata field is missing, activescale falls back to parsing the Istio-style node.id (sidecar~<ip>~<pod>.<namespace>~<namespace>.svc.cluster.local).

Summary Counters Meaning

  • messages: number of StreamMetrics messages received (one Recv() call).
  • stored_metrics: number of metric writes stored in Redis for accepted request and connection samples.
  • dropped_by_ids: messages dropped because pod identity could not be extracted.
  • dropped_by_names: metric families skipped because their name did not match the inbound, outbound, or aggregate service listener rules.

stored does not necessarily equal messages because a message can contain multiple metric families or multiple samples, and dropped_by_names is counted per metric family, not per message.

Debugging

Check Custom Metrics API is registered:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2'

Query inbound active requests with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/inbound_active_requests?labelSelector=app=<app>'

Query outbound active requests with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/outbound_active_requests?labelSelector=app=<app>'

Query active connections with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/active_connections?labelSelector=app=<app>'

Check activescale ingest logs:

kubectl logs -n ns-observability deploy/activescale | rg -n "stored (inbound_active_requests|outbound_active_requests|active_connections)|skipping metric name|missing pod identity"

Enable debug logs and restart:

kubectl -n ns-observability set env deploy/activescale LOG_VERBOSITY=4
kubectl -n ns-observability rollout restart deploy/activescale

Confirm Envoy bootstrap includes the metrics service:

istioctl proxy-config bootstrap <pod> -n <ns>| rg -n "envoyMetricsService|metrics_service|envoy_grpc|cluster_name|activescale|9000"

A minimal Kubernetes reference example lives in k8s-manifest-example.yaml.

Notes

Activescale reads both Envoy HTTP connection manager and listener stats, depending on the metric.

  • Activescale exposes inbound and outbound HTTP connection-manager request metrics separately.
  • http.admin.*, http.agent.*, and Prometheus-style envoy_http_downstream_rq_active aliases are ignored.
  • active_connections accepts aggregate listener.<address>.downstream_cx_active families for service listener ports.
  • Worker breakdowns such as listener.<address>.worker_0.downstream_cx_active and Istio-reserved proxy ports for admin, failure detection, debug, telemetry, health, and DNS are ignored.
  • Traffic-path listeners for outbound (15001), inbound (15006), and HBONE (15008) remain eligible.

Appendix

Missing metric values

Envoy can omit an active stat family for a new or idle pod. Activescale refreshes a per-pod heartbeat on every StreamMetrics message and applies the metric TTL to that heartbeat. The default TTL is 20s and can be changed with METRIC_TTL or --ttl.

  • A fresh heartbeat with no metric key is returned as 0 because collection is healthy and the pod has no observed active work.
  • A missing or expired heartbeat remains missing and can produce HTTP 404 NotFound because returning 0 would hide an Envoy collection outage and could cause an unsafe HPA scale-in.

Heartbeat-gated zeros prevent healthy idle pods from causing FailedGetPodsMetric or conservative scale-in delays without treating stale telemetry as zero. Provider summary logs count these values as synthesized_zeros.

Handling HTTP 404 NotFound

HTTP 404 NotFound means telemetry is unavailable, not that the metric value is zero.

Kubernetes HPA retries automatically and handles missing metrics conservatively, as described by the HPA algorithm.

Metric stateHPA action
All configured metrics return HTTP 404 NotFoundKeep the current replicas; neither scale up nor scale down
One metric returns HTTP 404 NotFound, another valid metric requests scale-upScale up using the valid metric
One metric returns HTTP 404 NotFound, another valid metric requests scale-downSkip scale-down and keep the current replicas
All configured metrics are availableUse the largest desired replica count

A direct Custom Metrics API client should apply the same policy: allow scale-up from another valid metric, block scale-down while any metric is unavailable, keep the current replicas when all metrics are unavailable, and retry on the next collection cycle. It must not substitute zero and should alert if HTTP 404 NotFound persists.

About

Push-based Kubernetes HPA adapter using active request/connection metrics to detect load spikes far earlier than CPU/Memory signals while improving resource utilization.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Activescale

Activescale enables fast Kubernetes autoscaling with push-based, per-pod active request and connection metrics from Envoy.

Active means currently in flight: requests not yet completed and connections still open. By Little's Law (L = λW), active requests reflect both request rate and latency.

Unlike CPU and memory metrics collected through metrics-server, Envoy pushes active signals directly to Activescale, avoiding extra scrape and polling delays.

Documentation

  • Guide: 한국어 | English — metric selection, HPA configuration, and verification
  • Motivation: 한국어 | English — rationale, latency comparison, and architecture
  • Test results: 한국어 | English — scaling responsiveness, Scale-out Coverage, and estimated resource utilization summary
  • Kubernetes manifest example: k8s-manifest-example.yaml — deployment and HPA example requiring Redis, ServiceAccount/RBAC, serving certificates, and a target workload

Features

  • Active request concurrency as an autoscaling signal that reflects both request rate and request latency
  • Push-based Envoy StreamMetrics ingestion without intermediary Prometheus scrape and KEDA polling delays
  • Pod-scoped Custom Metrics API values that HPA averages across workload replicas
  • Direction-aware inbound_active_requests and outbound_active_requests, plus aggregate active_connections
  • Stateless high availability with shared Redis/Valkey state and TTL-based stale metric expiry
  • Redis standalone and Cluster modes with optional TLS
  • Configurable gRPC receive limit and Klog-based summary and debug logging

Metrics

  • inbound_active_requests: inbound in-flight HTTP requests from http.inbound_*.downstream_rq_active
  • outbound_active_requests: outbound in-flight HTTP requests from http.outbound_*.downstream_rq_active
  • active_connections: open connections across aggregate Envoy service listeners from listener.<address>.downstream_cx_active

All metrics are exposed as pod-scoped custom metrics for HPA.

Choosing a request metric

Choose the metric based on the role of the workload being scaled.

WorkloadMetricWhen to use
Application or API serviceinbound_active_requestsUse for a workload that receives and processes client requests. This is the default choice for typical services.
Dedicated gateway or proxyoutbound_active_requestsUse when the workload primarily forwards client requests to upstream services and inbound_active_requests is unavailable or does not reflect the traffic.

Do not configure both metrics as a fallback. For a normal application, outbound requests usually represent calls to dependencies and can trigger scaling for unrelated downstream traffic.

If the workload role is unclear, generate a small amount of client traffic and query both metrics. Select the one that returns per-pod values and increases with the client request concurrency being tested.

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/inbound_active_requests?labelSelector=<selector>'
kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/outbound_active_requests?labelSelector=<selector>'

Architecture

graph LR
PodA["📦 service pod A<br/>Istio/Envoy"]
PodB["📦 service pod B<br/>Istio/Envoy"]
PodC["📦 service pod C<br/>Istio/Envoy"]
Agg(("⚙️ <b>activescale<b/><br/>(Stateless, HA)"))
Redis[("🗄️ shared memory<br/>(e.g., Redis)")]
KEDA{{"🚀 HPA<br/>(average computed by HPA)"}}
PodA -."1. Push<br/>(5s delay)".-> Agg
PodB -."1. Push<br/>(5s delay)".-> Agg
PodC -."1. Push<br/>(5s delay)".-> Agg
Agg ==>|2. update + TTL| Redis
KEDA --"3. Query"--> Agg
Redis -."4. metric value (per pod)".-> Agg
Agg --"5. metric value (per pod)"--> KEDA
style Redis fill:#E3ECF8,stroke:#6E8FB3,stroke-width:2px,color:#000
style KEDA fill:#DDEFD8,stroke:#7DA67D,stroke-width:2px,color:#000
style Agg fill:#E3ECFF,stroke:#6E8FB3,color:#000
style PodA fill:#FAFAFA,stroke:#999
style PodB fill:#FAFAFA,stroke:#999
style PodC fill:#FAFAFA,stroke:#999
Loading

Envoy Metrics Message Shape

Activescale receives Envoy gRPC StreamMetrics messages. Each message contains a node identity and a list of metric families. A simplified shape:

{
"identifier": {
"node": {
"id": "sidecar~10.0.0.1~my-pod.my-ns~my-ns.svc.cluster.local",
"metadata": {
"NAME": "my-pod",
"NAMESPACE": "my-ns"
}
}
},
"envoy_metrics": [
{
"name": "http.inbound_0.0.0.0_8080;.downstream_rq_active",
"metric": [
{ "gauge": { "value": 3 } }
]
},
{
"name": "http.outbound_0.0.0.0_8080;.downstream_rq_active",
"metric": [
{ "gauge": { "value": 4 } }
]
},
{
"name": "listener.0.0.0.0_8080.downstream_cx_active",
"metric": [
{ "gauge": { "value": 12 } }
]
},
{
"name": "cluster.xds-grpc;.circuit_breakers.default.cx_pool_open",
"metric": [
{ "gauge": { "value": 1 } }
]
}
]
}

Pod identity extraction prefers node.metadata.NAME and node.metadata.NAMESPACE. If either metadata field is missing, activescale falls back to parsing the Istio-style node.id (sidecar~<ip>~<pod>.<namespace>~<namespace>.svc.cluster.local).

Summary Counters Meaning

  • messages: number of StreamMetrics messages received (one Recv() call).
  • stored_metrics: number of metric writes stored in Redis for accepted request and connection samples.
  • dropped_by_ids: messages dropped because pod identity could not be extracted.
  • dropped_by_names: metric families skipped because their name did not match the inbound, outbound, or aggregate service listener rules.

stored does not necessarily equal messages because a message can contain multiple metric families or multiple samples, and dropped_by_names is counted per metric family, not per message.

Debugging

Check Custom Metrics API is registered:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2'

Query inbound active requests with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/inbound_active_requests?labelSelector=app=<app>'

Query outbound active requests with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/outbound_active_requests?labelSelector=app=<app>'

Query active connections with a selector:

kubectl get --raw '/apis/custom.metrics.k8s.io/v1beta2/namespaces/<ns>/pods/*/active_connections?labelSelector=app=<app>'

Check activescale ingest logs:

kubectl logs -n ns-observability deploy/activescale | rg -n "stored (inbound_active_requests|outbound_active_requests|active_connections)|skipping metric name|missing pod identity"

Enable debug logs and restart:

kubectl -n ns-observability set env deploy/activescale LOG_VERBOSITY=4
kubectl -n ns-observability rollout restart deploy/activescale

Confirm Envoy bootstrap includes the metrics service:

istioctl proxy-config bootstrap <pod> -n <ns>| rg -n "envoyMetricsService|metrics_service|envoy_grpc|cluster_name|activescale|9000"

A minimal Kubernetes reference example lives in k8s-manifest-example.yaml.

Notes

Activescale reads both Envoy HTTP connection manager and listener stats, depending on the metric.

  • Activescale exposes inbound and outbound HTTP connection-manager request metrics separately.
  • http.admin.*, http.agent.*, and Prometheus-style envoy_http_downstream_rq_active aliases are ignored.
  • active_connections accepts aggregate listener.<address>.downstream_cx_active families for service listener ports.
  • Worker breakdowns such as listener.<address>.worker_0.downstream_cx_active and Istio-reserved proxy ports for admin, failure detection, debug, telemetry, health, and DNS are ignored.
  • Traffic-path listeners for outbound (15001), inbound (15006), and HBONE (15008) remain eligible.

Appendix

Missing metric values

Envoy can omit an active stat family for a new or idle pod. Activescale refreshes a per-pod heartbeat on every StreamMetrics message and applies the metric TTL to that heartbeat. The default TTL is 20s and can be changed with METRIC_TTL or --ttl.

  • A fresh heartbeat with no metric key is returned as 0 because collection is healthy and the pod has no observed active work.
  • A missing or expired heartbeat remains missing and can produce HTTP 404 NotFound because returning 0 would hide an Envoy collection outage and could cause an unsafe HPA scale-in.

Heartbeat-gated zeros prevent healthy idle pods from causing FailedGetPodsMetric or conservative scale-in delays without treating stale telemetry as zero. Provider summary logs count these values as synthesized_zeros.

Handling HTTP 404 NotFound

HTTP 404 NotFound means telemetry is unavailable, not that the metric value is zero.

Kubernetes HPA retries automatically and handles missing metrics conservatively, as described by the HPA algorithm.

Metric stateHPA action
All configured metrics return HTTP 404 NotFoundKeep the current replicas; neither scale up nor scale down
One metric returns HTTP 404 NotFound, another valid metric requests scale-upScale up using the valid metric
One metric returns HTTP 404 NotFound, another valid metric requests scale-downSkip scale-down and keep the current replicas
All configured metrics are availableUse the largest desired replica count

A direct Custom Metrics API client should apply the same policy: allow scale-up from another valid metric, block scale-down while any metric is unavailable, keep the current replicas when all metrics are unavailable, and retry on the next collection cycle. It must not substitute zero and should alert if HTTP 404 NotFound persists.

About

Push-based Kubernetes HPA adapter using active request/connection metrics to detect load spikes far earlier than CPU/Memory signals while improving resource utilization.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages