端到端 SLO (服务级别目标) 治理平台,集成 MeterSphere 拨测数据,自动计算误差预算,并由 AI 辅助输出结构化诊断报告
-
Updated
Dec 11, 2025 - Python
端到端 SLO (服务级别目标) 治理平台,集成 MeterSphere 拨测数据,自动计算误差预算,并由 AI 辅助输出结构化诊断报告
Azure-native SLO/SLI engine with error budget tracking, burn-rate alerts, and CLI for Azure Monitor, Application Insights & Log Analytics
Azure SLO Dashboard with AI-powered error budget explainer — Application Insights + Container Apps + Claude
SLO compliance, error budgets, and multi-window burn rate from Prometheus — markdown reports + CI-friendly check command
Multi-window multi-burn-rate SLO alerting over spans, with an honest 3-regime benchmark: a naive threshold detects real regressions faster, but raises 3 false pages on healthy spiky traffic where multi-window raises 0. The value is precision, not speed, and the benchmark proves it both ways.
SLO error-budget toolkit. Budget accounting + projection and Google SRE Workbook multi-window burn-rate alerting. CLI, library, and CI gate. Zero dependencies.
Conformidade de observabilidade como controle executavel: 30 controles mapeados para praticas ITIL 4 e objetivos COBIT 2019, com regra de teto no score de maturidade e dashboard read-only.
SLO tracking, error budget calculation & burn-rate alerting for Kubernetes | Google SRE model | Prometheus | Slack | PagerDuty
SRE lab for multi-window SLO burn rates, alert ownership and route coverage, incident reports, FastAPI, Prometheus, Kubernetes, and Terraform.
SRE error-budget and burn-rate calculator: SLI, budget consumed, time to exhaustion, and multi-window burn-rate alerting.
Real postmortems, runbooks and SLOs from incidents I actually hit — a public metrics disclosure, a security control that failed open, and an OIDC trust mismatch. Blameless, with open action items left open.
Multi-window burn-rate alerts and Grafana dashboards from a YAML SLO spec
SLO-based alerting reference: objectives as code, multi-window burn-rate rules, fault injection to prove alerts fire
SLO + error-budget tracker for Python services. FastAPI middleware, Prometheus exporter, multi-window burn-rate alerts. Part of the Platform Reliability Stack.
SLO error-budget and multi-window multi-burn-rate tracking (Google SRE workbook approach) from plain SLI CSVs — no Prometheus stack needed to try it
Python service for allocating error-budget burn across services, dependencies, deploy windows, and operational ownership lanes.
Executive-grade platform SLO dashboard aggregating PagerDuty and Dynatrace reliability signals
Blocks unattended remediation once the SLO error budget is spent. Automation may propose the fix; the error budget decides whether it acts. Pure Python, zero runtime deps.
Kubernetes operator (kopf) that compiles ServiceLevelObjective CRDs into Prometheus SLI recording rules and multi-window multi-burn-rate alerts, with live error-budget reporting.
Add a description, image, and links to the error-budget topic page so that developers can more easily learn about it.
To associate your repository with the error-budget topic, visit your repo's landing page and select "manage topics."